Nvidia and Amazon Web Services pushed a coordinated set of AI infrastructure upgrades into production this week, headlined by new EC2 G7 instances that deliver up to 4.6x AI inference performance and 2.1x graphics performance over the prior G6 generation. The G7 instances run on Nvidia's RTX PRO 4500 Blackwell Server Edition GPUs, and they ship alongside a quiet but consequential change in Amazon OpenSearch Serverless: GPU-accelerated vector search, powered by the Nvidia cuVS library, is now the default. AWS also earned Nvidia Exemplar Cloud status on the GB300 for training workloads, a benchmark that signals tuned performance against Nvidia's reference architecture.
The throughline is that the two companies are trying to remove the operational tax of running AI in production. G7 instances support up to 8 GPUs, 256GB of total GPU memory, 700 Gbps of EFA-enabled networking, and up to 7.6TB of local NVMe SSD storage, available in one-, two-, four-, and eight-GPU configurations with bare metal coming soon. The configuration spread is the point: AI teams, media and rendering shops, simulation and CAD users, and analytics teams can all sit on the same instance family and right-size rather than over-provision.
“Building AI systems at scale is demanding, requiring low-latency inference, fast vector search, strong GPU price-performance and infrastructure that can grow without multiplying operational complexity.”— Josiah Byers, NVIDIA
For inference workloads specifically, the 4.6x jump versus G6 is the headline number teams will use to model cost-per-token economics. Graphics-heavy workloads — high-resolution video pipelines, virtual desktop infrastructure, gaming, and spatial computing — get the 2.1x lift. Data teams running Apache Spark on Amazon EMR get GPU acceleration via the Nvidia cuDF library, putting analytics pipelines and vector database workloads on the same silicon as the model serving layer.
Key facts
- 01EC2 G7 instances deliver up to 4.6x AI inference performance and 2.1x graphics performance versus G6.
- 02G7 supports up to 8 GPUs, 256GB of GPU memory, 700 Gbps EFA networking, and 7.6TB of local NVMe SSD.
- 03NVIDIA cuVS is now the default vector compute in OpenSearch Serverless, cutting indexing cost 75% versus CPU builds.
- 04Billion-scale vector databases can now be built in under an hour on OpenSearch Serverless.
- 05AWS achieved NVIDIA Exemplar Cloud status on NVIDIA GB300 for training workloads.
G7 instances are available through AWS Deep Learning AMIs, Amazon Deep Learning Containers, Amazon EMR, Amazon EKS, Amazon ECS, and graphics AMIs, with Amazon SageMaker AI support coming soon. The distribution surface matters more than it sounds: developers do not have to rebuild their orchestration stack to pick up the new hardware, which is usually where GPU upgrades stall in enterprise environments.
The OpenSearch change is arguably the bigger story for builders. The next generation of Amazon OpenSearch Serverless makes Nvidia cuVS the default compute path for all vector collections, no infrastructure management required. AWS says vector indexing runs up to 10x faster at a quarter of the cost compared with CPU-only builds, and billion-scale vector databases can now be assembled in under an hour. For teams running retrieval-augmented generation, semantic search, recommender systems, and agentic AI, that collapses a procurement-and-tuning project into a default setting.
That is the part of the AI stack that has been quietly bottlenecking real deployments. Vector search at scale has been either expensive on CPU or operationally heavy on self-managed GPU clusters. Making cuVS the default in a serverless product means a builder gets the performance without provisioning anything — and serverless idle behavior keeps the bill from running away when workloads are bursty.
The Exemplar Cloud designation on GB300 is the training-side counterpart. Nvidia uses the Exemplar Clouds program to certify that a cloud provider's deployment of its hardware hits reference-architecture performance thresholds. The practical use case is procurement: an enterprise evaluating cloud providers for large-scale training can lean on the certification rather than running its own months-long bake-off. AWS hitting that mark on the GB300 is a signal to AI leaders comparing total cost of ownership across hyperscalers.
Nvidia detailed the lineup in a blog post by Josiah Byers, framing the work as addressing the four constraints enterprises hit at production scale: low-latency inference, fast vector search, GPU price-performance, and infrastructure that grows without multiplying operational complexity. The choice to make cuVS a default rather than an opt-in is the clearest expression of that framing.
The competitive read is that Nvidia is reinforcing its position inside the largest cloud at every layer of the AI stack rather than just at the training tier. Google Cloud and Microsoft Azure both have their own deep Nvidia integrations, and this announcement keeps AWS competitive on the inference, retrieval, and training fronts simultaneously. It also continues a pattern from earlier this year, when Nvidia launched its Agent Toolkit to anchor enterprise specialized AI builds, which AI Chat Daily covered — the company is steadily widening its software footprint above the silicon.
What is not addressed in the announcement is pricing detail for G7 instances beyond the relative performance gains, and the cuVS-default switch in OpenSearch Serverless will need real-world validation at billion-vector scale outside AWS's own benchmarks. The 10x indexing and quarter-cost claims are AWS figures; customers running mixed workloads with strict latency SLAs will want their own numbers before migrating production retrieval pipelines.
For the broader AI market, the meaningful shift is that vector search and GPU-accelerated retrieval are moving from specialized infrastructure projects into default cloud primitives. That changes the build-versus-buy math for any company shipping a RAG product or an agentic application — the moat of having a well-tuned vector index shrinks when a serverless default delivers it at a quarter of the cost. Nvidia and AWS are betting that making the hard parts of production AI boring is what unlocks the next wave of enterprise deployment, and on the evidence here, that bet is structurally sound.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




