AI labs face a data crunch as websites block scrapers and sue over unlicensed training. Internal warnings called it theft on a historic scale. New licensing deals and opt-in standards emerge while synthetic and proprietary sources gain ground. The slurp juice era marks a messy transition.
The AI industry has a new metaphor. “Slurp juice” once described an absurd NFT mechanic for minting cartoon apes. Now it captures something raw about how the biggest labs chase fresh material to feed their models. They slurp. They consume. They move on when the source runs dry.
Gizmodo coined the phrase this week. The piece nails the frenzy. New models drop weekly. Names multiply into tongue twisters. Executives race to differentiate products while the underlying fuel — high quality text, code, images, conversations — grows scarce. The open web that powered early breakthroughs no longer delivers the same returns.
But the metaphor runs deeper than product overload. It speaks to the aggressive data acquisition strategies that have defined frontier AI for years. Scraping at massive scale. Circumventing paywalls. Training on news articles, books, social posts and proprietary records without explicit permission. The approach delivered impressive capabilities fast. It also triggered lawsuits, blocked crawlers and internal warnings that read like confessions.
Documents unsealed in the New York Times copyright case against OpenAI and Microsoft paint a stark picture. Microsoft Director of Applied Science Brent Hecht described the practice as “an astonishing theft of unprecedented proportions” and potentially “the largest theft of labor in human history.” He argued it made “a complete mockery of the idea of ‘fair use.'” OpenAI’s Nick Turley warned publishers faced an “existential threat.” Traffic data cited in the filings showed click-through rates plunging 51 to 94 percent for some news organizations after AI products launched.
The scale shocked even skeptics. One Common Crawl-derived dataset alone held more than 2 million documents from nytimes.com. Mid-training sets contained over 91,000 copies of works from the Times, Daily News and Center for Investigative Reporting. Employees allegedly devised ways to bypass paywalls without detection. Project Mango and Project Taxi reportedly funneled news content between the two companies for training and evaluation.
Yet the data drought is real. An MIT-led study published this week by the Data Provenance Initiative examined 14,000 web domains across three major training datasets — C4, RefinedWeb and Dolma. It found 5 percent of all data and 25 percent of the highest-quality sources now restricted through robots.txt or terms of service. In the C4 set, nearly 45 percent of data faces terms-of-service blocks. “We’re seeing a rapid decline in consent,” lead author Shayne Longpre told The New York Times in coverage of the research.
Publishers, platforms and regulators push back.
Restrictions accelerated in the past year. Reddit, once a favorite source for conversational language, now licenses its data. Google pays the company roughly $60 million annually. OpenAI reportedly pays about $70 million a year for similar access. Meta’s Mark Zuckerberg has highlighted how its own Facebook and Instagram posts dwarf Common Crawl in volume and offer advantages in informal, multilingual expression.
News organizations strike direct deals. Axel Springer, the Financial Times and academic publishers like Wiley have signed multimillion-dollar licensing agreements. Shutterstock reported $138 million in data licensing revenue in 2024. The overall AI training dataset market sits near $4 billion this year and is projected to multiply several times over by the early 2030s, according to Troveo.
A new trade group wants to change the rules entirely. The Dataset Providers Alliance, formed this summer and including Rightsify, Pixta and Calliope Networks, released a position paper calling for an opt-in system. Data could be used for AI training only after explicit consent from creators and rights holders. The stance directly challenges the “publicly available” rationale that powered the first wave of generative tools.
Even Wikimedia pushed back. In a recent rebuke to OpenAI, the foundation behind Wikipedia cited the operational costs of serving free content that ends up in training runs. The message was clear. Free riding carries hidden expenses that eventually get passed along or blocked.
AI companies aren’t stopping. They pivot. Synthetic data fills some gaps but struggles with novelty and edge cases. Internal corporate records from defunct businesses become prized assets because they capture how real expert work actually happens. Robotics and physical AI demand data never posted publicly. Social platforms with real-time user behavior offer signals that static web scrapes cannot match.
But the hunger persists. Recent reporting shows AI agents from multiple labs bypass robots.txt signals despite public statements to the contrary. Australian Broadcasting Corporation told parliament that OpenAI and Anthropic had circumvented opt-out mechanisms. In Australia, the companies lobby for copyright changes that would ease training on local content. Google told one inquiry that requiring prior authorization from every rights holder would make AI development impossible.
The contradiction sits at the center. Labs publicly champion open innovation while privately acknowledging the moral and business hazards of their methods. They pay premium prices for licensed data yet continue aggressive crawling where they can. They warn internally about destroying the creative industries that produce the content they need.
And the models keep getting bigger. GPT variants proliferate. Meta, Anthropic, Google and xAI pour resources into ever-larger systems. Each advance consumes more data than the last. Demand doubles annually while fresh high-quality public content grows far slower. Something has to give.
For now the industry slurps what it can. It buys bankrupt companies’ archives. It licenses Reddit threads and news archives. It generates synthetic examples and hopes they don’t collapse into model collapse. It fights in court over fair use while negotiating private deals that effectively concede the point.
The slurp juice era won’t last forever. The juice runs out. When it does, the winners will be those who figured out how to create new high-quality material at scale. Or those who built systems that learn more efficiently from less. The rest may find themselves holding empty containers and wondering why they didn’t see the bottom sooner.
One thing is certain. The days of treating the entire internet as an unguarded buffet are ending. Consent, compensation and careful sourcing are becoming competitive advantages instead of obstacles. The labs that adapt fastest will set the terms for whatever comes after the slurp.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | Adorable AI Sidekicks Mask Growing Risks of Deception and Data Overreach | 0 | 8.82 | 02-10-2026 |
| 2 | When AI Agents Turn on Their Masters: Hackers Lose Email Harvest to Rogue Security Tools | 0 | 8.35 | 02-10-2026 |
| 3 | AI’s Cash Burn Crisis: Why Massive Losses Persist Despite Revenue Growth | 0 | 9.65 | 08-10-2026 |
| 4 | AI Agents Slip the Leash: How Frontier Labs Lost Control of Their Own Creations | 0 | 8.59 | 02-10-2026 |
| 5 | AI Comics Swipe Punchlines to Peddle Subscriptions | 0 | 12.34 | 07-10-2026 |
| 6 | The agents have jumped the fence: AI faces its Jurassic Park moment | 0 | 7.7 | 01-08-2026 |
| 7 | OpenAI hack sparks further concern over AI models going rogue | 0 | 7.01 | 28-09-2026 |
| 8 | OpenAI pauses training of latest models after agents probed US government sites in unexpected ways | 0 | 7.16 | 28-09-2026 |
| 9 | LASST sues OpenAI over rogue AI cyberattacks against Hugging Face | 0 | 16.45 | 30-09-2026 |
| 10 | AI giants probing tens of thousands of security incidents – Axios | 0 | 9.83 | 27-09-2026 |