Bright Data CEO Or Lenchner argues that AI's next bottleneck is not model size or compute but the live web itself, with 97% of AI organizations now dependent on real-time web data infrastructure and 90% reporting they feel boxed in by access restrictions. Gartner projects that 60% of AI projects not supported by AI-ready data — accurate, structured, and contextualized — will be abandoned by the end of the year. The frame Lenchner is pushing: a dedicated retrieval layer sitting between the open web and the models, capable of operating at the scale of hundreds of millions of existing domains and billions of new URLs created every week.
The problem starts with how the web was built. It was not designed for automated discovery and retrieval at the cadence AI applications now demand. Snapshots of training data go stale, prices and inventories and sentiment shift continuously, and a model with no live access to that movement begins producing answers that are confidently wrong. Lenchner's pitch is that solving this is an infrastructure problem, not a model problem.
Retrieval-augmented generation was supposed to close that gap by letting models pull external data at query time, but Lenchner says the execution layer underneath RAG is where most teams stall. Latency, blocking, fragmented sources, and inconsistent formatting turn a clean architecture diagram into an engineering grind. One survey cited in the piece found 56% of AI practitioners said businesses need real-time web data access specifically to improve trust in AI outputs — the link between fresh inputs and lower hallucination rates is becoming a procurement criterion, not a research curiosity.
Key facts
- 01Gartner projects 60% of AI projects not supported by AI-ready data will be abandoned by the end of the year.
- 0297% of AI organizations depend on real-time web data infrastructure, but 90% feel boxed in by access restrictions.
- 0356% of AI practitioners surveyed said businesses need real-time web data to improve trust in AI outputs.
- 04Bright Data CEO Or Lenchner says the infrastructure must mimic human browsing 80 billion times a day across millions of sites.
- 05The retrieval layer must span hundreds of millions of existing domains and billions of new URLs created each week.
Bright Data's own positioning rests on the claim that enterprises increasingly combine public web retrieval with APIs, licensed datasets, and internal proprietary data in a single AI workflow. Stitching those together in real time, without breaking on JavaScript-heavy sites or aggressive anti-bot software, is the work Lenchner says most teams underestimate when they try to build in-house. He frames the trained model and the live data feed as two halves of the same product.
The technical surface area is unusual. Lenchner describes infrastructure that emulates human browsing behavior — IP address, geographic location, and what he calls 1,000 more parameters — to access content the way a website expects a visitor to look. The volume he cites is the headline number: 80 billion such interactions per day across millions of websites, transforming raw HTML into structured feeds the model can consume.
“Think of the trained model as intelligence and relevant data as knowledge. A powerful intelligence layer sitting on top of a hollow knowledge layer is like a genius who knows nothing—useless in practice.”— Or Lenchner, CEO of Bright Data
Governance is the part that gets undersold in retrieval pitches and oversold in regulatory ones. Lenchner says compliant platforms can operate within frameworks including the EU General Data Protection Regulation and the California Consumer Privacy Act by limiting collection to openly accessible public information, avoiding paywalls and private logins, and running consent-based networks where IP-address owners are compensated. That's the version vendors describe; regulators in both jurisdictions will decide whether it holds up under enforcement.
The build-versus-buy question is where Bright Data and competitors expect to win share. Lenchner argues that once retrieval becomes critical infrastructure for a company, doing it in-house turns into a full-time engineering problem that competes directly with the AI work the team was hired to do. The implication, which is also the commercial pitch, is that orchestration, observability, and unblocking at scale are specialized enough to outsource — the same argument cloud providers made about servers fifteen years ago.
“It's basically having infrastructure that can mimic a web user with identifying information—IP address, location, and 1,000 more parameters. And at scale. Think of doing that 80 billion times a day for millions of websites.”— Or Lenchner, CEO of Bright Data
There are real reasons to be skeptical of the framing. The 97% and 90% figures come from research environments friendly to data-infrastructure vendors, and the line between scraping at scale and scraping responsibly is one that courts, not vendors, will eventually draw. Publishers, platforms, and regulators are tightening the rules around automated access in parallel, and Lenchner's own metaphor — infrastructure that looks exactly like the website expects you to look — describes the adversarial dynamic publishers will keep pushing back on.
Still, the directional point is hard to dismiss. If models are converging on similar capabilities at the frontier, then differentiation moves down the stack to whoever can deliver fresher, cleaner, more relevant inputs at lower latency. Lenchner's closing observation — that everything happening in the world is being uploaded to the public web, and that the volume of new data is accelerating — is the part that should worry every AI buyer relying on a static training cut.
The AI market spent two years pricing models as the scarce asset. The next eighteen months will test whether the scarce asset is actually the pipe that keeps those models fed, and whether enterprises end up paying more for retrieval infrastructure than they do for the inference layer sitting on top of it. If Gartner's 60% abandonment forecast lands anywhere close, the vendors selling that pipe — Bright Data among them — will be the ones quietly capturing the budget that was supposed to go to foundation models.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




