A hedge fund buys foot-traffic counts before an earnings call. A fashion brand tracks ten thousand competitor prices a day. A language model swallows another slice of the open web before breakfast.
None of them are stealing anything. They’re all drawing from the same shrinking pool: public information that used to feel infinite.
That pool is no longer infinite, and everyone chasing it now knows it.
The Market Woke Up First
Money moved before anyone declared a crisis. The alternative data market, information sold to investors that isn’t pulled from company filings or press releases, reached about $2.8 billion in 2025, up roughly 27% from the year before. Web-scraped datasets make up the biggest slice of that spend.
Finance and retail got there first, then the habit spread. Small research shops now buy datasets that once sat behind enterprise-only price tags, and that alone has tightened supply further.
None of it happens without infrastructure most people never think about. Getting data out of the open web at scale usually runs through proxy networks, and a choice as basic as picking between residential proxies vs dedicated infrastructure ends up deciding how much a team collects before getting blocked.
Then AI Showed Up Hungry
The scale involved is hard to picture. Businesses generate an estimated 2.5 quintillion bytes of data daily, yet close to 73% of them still hit IP blocks and geographic restrictions trying to reach the data they actually want.
Language models are the biggest new buyer by far. They train on Wikipedia, news archives, forums, books — anything with enough clean text to be useful. For years the plan was simple: feed them more, every cycle, forever.
That plan has a ceiling. Researchers at Epoch AI projected that high-quality public text could run dry as early as 2026, a warning covered in detail by MIT Technology Review. Publishers noticed the same trend from the other side and started tightening their terms of service to keep AI crawlers out.
Some companies have taken the shortage in a more literal direction — asking employees to document their own workflows so the resulting notes can train the systems built to replace them. When the open web stops supplying enough raw material, human expertise itself becomes the substitute.
Demand keeps climbing while the well keeps shrinking. That squeeze is why clean, human-written data has become one of the more valuable commodities in tech.
Volume Lost, Quality Won
Nobody brags about raw record counts anymore. Collecting billions of messy rows is easy. Turning them into something a model or analyst can actually trust is not.
Harvard Business Review made a version of this argument years before the current AI wave, noting that reused, well-organized data creates value both for the company holding it and for the wider ecosystem around it. Dirty data is worse than no data at all — every decision built on top of it inherits the same errors.
Search platforms and price-comparison sites learned this the hard way years ago. Bad inputs sink the whole product, no matter how good the interface looks on top. Buyers now pay a premium for accuracy, freshness, and correct geography. A pricing feed labeled “Germany” that was actually pulled from an Austrian server isn’t a rounding error — it’s a redo, and redos cost real money.
The Legal Fog Finally Lifted
None of this scales without some shared understanding of what’s actually permitted. For years, scraping public data sat in a gray zone that kept legal teams nervous.
hiQ Labs v. LinkedIn changed that. Courts repeatedly found that pulling publicly visible profile data didn’t violate the Computer Fraud and Abuse Act. hiQ eventually lost on a separate contract claim and shut down, so the ruling didn’t settle everything — but it drew a usable line between public and gated information, one the industry still leans on today.
Where the Pressure Goes Next
Sites keep walling off access. AI companies keep buying. High-quality public data is behaving less like a free resource and more like a commodity with a price tag and a shrinking supply curve.
The winners in that market won’t be whoever scrapes fastest or hoards the most rows. They’ll be the teams that collect responsibly, keep their datasets clean, and stay inside the legal lines everyone else is still scrambling to find.
