Web data infrastructure emerges as AI's next bottleneck

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Bright Data CEO Or Lenchner argues the web was not designed for the automated discovery and retrieval that modern AI applications demand, requiring a dedicated data infrastructure layer to navigate hundreds of millions of domains and billions of new URLs created weekly.
- Lenchner says traditional training snapshots are no longer sufficient, as organizations need continuous real-time feeds to track competitor pricing, consumer sentiment, and market fluctuations — warning that 'stale answers lead to bad decisions.'
- A survey cited in the article found 56% of AI practitioners said businesses need access to real-time web data to improve trust in AI outputs, while Gartner predicts 60% of AI projects not supported by AI-ready data will be abandoned by the end of the year.
- Research referenced in the piece found 97% of AI organizations depend on real-time web data infrastructure, yet 90% feel restricted by limitations in accessing it.
- Bright Data's platform emulates human browsing behavior — mimicking IP addresses, location, and '1,000 more parameters' — to bypass antibot software and JavaScript-heavy sites, handling roughly 80 billion requests per day for millions of websites.
- The platform addresses governance by enforcing compliance with GDPR and CCPA and limiting retrieval to openly accessible public information, avoiding paywalls and private logins.
Why it matters: With 97% of AI organizations already depending on real-time web data but 90% feeling boxed in by access restrictions, enterprises face a build-vs-buy decision: dedicate full-time engineering resources to data retrieval or outsource to specialized platforms. Gartner's projection that 60% of unsupported AI projects will be abandoned by year-end gives the infrastructure gap a hard deadline.
Ask SkimNews




