Amazon Web Services launched a redesigned OpenSearch Serverless on Thursday that decouples compute from storage so AI agents can trigger search and vector workloads in seconds and pay $0 when idle. The product is a fully managed search and vector database aimed squarely at agentic workloads — bursts of machine traffic that spike without warning and vanish just as fast. Cloudflare data puts bots at 31% of overall HTTP traffic over the last six months, with AI crawlers, search engines, and assistants making up roughly a quarter of those bot requests. The company's senior product manager Lai Yi Ohlsen expects non-human traffic to exceed human traffic in the first half of 2027.
The redesign matters because cloud infrastructure was built for humans who click, scroll, and stream at a roughly predictable pace. Agents do not. A single user prompt can spawn sub-agents that hit hundreds of databases, call APIs, and retrieve documents in parallel before going quiet again. Under the prior OpenSearch Serverless generation, customers had to keep at least one instance running at all times because compute and storage were coupled — the equivalent of paying for a parking space whether or not the car was there.
Tia White, general manager for Amazon OpenSearch Service, framed the launch as a response to agents moving from pilots into production at enterprises. The decoupling lets AWS scale compute up the moment an agent fires and back to zero when the work stops, which White compared to a metered parking spot rather than a reserved one.
Key facts
- 01AWS launched a new generation of OpenSearch Serverless that scales compute to zero when AI agents go idle, costing customers $0 during downtime.
- 02Cloudflare data shows bots accounted for 31% of overall HTTP traffic over the last six months, with AI crawlers, search engines, and assistants making up roughly a quarter of bot requests.
- 03Cloudflare projects non-human traffic will exceed human traffic in the first half of 2027.
- 04The new Serverless integrates natively at launch with AI development platforms Vercel and Kiro.
- 05Databricks, Snowflake, Microsoft Azure, and Cloudflare are all shipping similar agent-oriented infrastructure changes.
At launch, the new Serverless integrates natively with AI development platforms Vercel and Kiro, letting developers wire production-ready search and vector backends into agent stacks without managing infrastructure. That matters because retrieval — pulling the right document, the right embedding, the right tool result — is where most agent loops either work or stall. Cheaper, more elastic retrieval directly compresses the unit economics of running agents at scale.
AWS is not alone in retooling. Databricks and Snowflake are repositioning their platforms as AI memory and retrieval layers for enterprise data. Microsoft has updated Azure to handle agent traffic bursts and to share memory across agents. Cloudflare last month rolled out its own infrastructure giving agents persistent environments and instant scaling, the same problem AWS is now answering with OpenSearch Serverless.
The pressure comes from above as much as below. At Google I/O last week, Google said users will be able to delegate tasks — researching purchases, booking travel, browsing the web, interacting with apps — to AI systems. Enterprises are doing the same internally, deploying agents that constantly retrieve information, invoke tools, and generate machine-to-machine traffic that never touches a human eyeball.
“Non-human traffic will exceed human traffic sometime in the first half of 2027”— Lai Yi Ohlsen, Senior Product Manager at Cloudflare
Cloudflare's projection that machines will out-traffic humans by mid-2027 is the throughline. If the bot share keeps climbing from 31%, the systems sitting underneath — search indexes, vector stores, identity, rate-limiting, billing — all have to be rebuilt or repriced around bursts rather than steady state. That is the bet every hyperscaler is now making with product launches rather than slideware.
The open question is whether the cost savings from scale-to-zero designs actually flow through to customers or get absorbed by the platforms. Pay-per-use only beats reserved capacity when workloads are genuinely spiky; agents that run continuously could end up more expensive on a metered model than on a reserved one. Enterprises piloting agentic search will need to watch their bills closely as workloads mature from experimentation into steady production, the exact transition White cited as the trigger for the redesign.
There is also a question of lock-in. Native integrations with Vercel and Kiro are convenient for developers building today, but every agent framework wired to OpenSearch Serverless is one more workload that will not move easily to Azure or to Cloudflare's equivalent. The agent infrastructure wars are starting to look a lot like the early cloud wars — fought on developer experience, priced on idle time, and won by whoever owns the default retrieval layer.
For AWS, the strategic logic is straightforward: if agents become the dominant traffic class on the internet, the company that hosts their search and memory layer captures a structural share of the next compute cycle. Pricing idle compute at zero is a sharp move because it removes the strongest reason an enterprise might run agents on a competitor's stack. The internet rebuilt for machines will be a much larger market than the one built for humans, and the providers moving first on the plumbing are the ones setting the terms.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




