Skip to main content
Live
Main content

DeepMind's SIMA 2 agent plays, reasons and learns in 3D game worlds

Google DeepMind extends 15 years of games research from Atari to EVE Online with a new agent that collaborates with human players.

Jaeden Schafer
Editor in Chief · · 4 min read
Google logo

Google DeepMind unveiled SIMA 2 in November 2025, a general agent that plays, reasons and learns alongside humans inside 3D virtual worlds. The release caps 15 years of the lab's games research, a stretch that started with Atari benchmarks and now stretches through titles as sprawling as EVE Online. DeepMind is framing SIMA 2 not as a game bot but as a step toward embodied AI that can hold context, take instructions and collaborate across unfamiliar environments.

The pitch is a shift from solo mastery to cooperative competence. Earlier DeepMind agents were built to beat games — AlphaGo took down Go, the Atari DQN work in 2013 cleared dozens of arcade titles, AlphaStar reached grandmaster level on StarCraft II. SIMA 2 is instead scored on how well it plays with a human partner across environments it has never seen before.

That reframing matters because 3D game worlds have become the cheapest available proxy for the physical world. They contain navigation, tool use, natural-language instructions, other agents, long-horizon goals and consequences. If a single model can hold all of that together across many games, the same architecture is a candidate for robotics, simulation and any long-running agent workflow outside gaming.

Key facts

  • 01Google DeepMind released SIMA 2 in November 2025, an agent that plays, reasons and learns alongside humans in 3D virtual worlds.
  • 02The launch marks 15 years of DeepMind AI research in games, a lineage running from Atari through EVE Online.
  • 03SIMA 2 is positioned as a collaborative agent rather than a solo player, closing the gap between game AI and general-purpose embodied agents.

SIMA 2 is the successor to the original SIMA — the Scalable Instructable Multiworld Agent — which DeepMind first detailed in 2024 as a single model trained across commercial video games. Version 2 leans harder on reasoning, framing the agent as one that can explain what it is about to do, why it is doing it, and adapt when the human partner changes the plan mid-task.

The 15-year through-line DeepMind is drawing is deliberate. Atari in 2013 established that a single neural network could learn many games from raw pixels. AlphaGo in 2016 established that self-play plus search could exceed the best humans in a closed domain. AlphaStar and OpenAI Five pushed into real-time strategy and MOBAs. EVE Online, a massively multiplayer sandbox with tens of thousands of concurrent players and no clean win condition, is closer to a live economy than a game — and closer to the messiness general agents will actually face.

There is a commercial subtext. DeepMind's parent, Alphabet, is racing OpenAI, Anthropic and Meta on agentic products, and the gap between a chat model and an agent that can take multi-step actions in a live environment is the current frontier. Games are where that gap gets narrowed with the lowest safety cost — an agent that misbehaves in EVE Online loses a spaceship, not a customer account.

The company has not published SIMA 2 benchmark tables in the initial announcement, and it is not offering the agent as a product. This is a research release, positioned alongside the lab's ongoing work on Gemini-based agents and its robotics program. Whether the underlying model is a Gemini derivative or a separate stack is not spelled out in the launch post.

The skeptic's read is that game agents have a long history of impressive demos that do not generalize. AlphaStar's StarCraft II performance did not translate directly into any shipping product. The original SIMA was constrained to games it had been trained on. Until DeepMind shows SIMA 2 transferring skills into a genuinely novel environment — ideally one outside gaming — the claim of a general instructable agent is a research thesis rather than a shipped capability.

Related · from this week
Google bakes computer use into Gemini 3.5 Flash as a native tool
Jaeden Schafer · 4 min read →

For the broader AI market, SIMA 2 is a marker of where the agent race is actually being run. The interesting work is no longer about single-turn benchmarks; it is about whether a model can hold a plan, take feedback and operate across environments without being retrained for each one. DeepMind has 15 years of infrastructure aimed at exactly that problem, and the labs competing on agentic products will now have to show comparable range in a domain where the ground truth is unambiguous — either the agent completed the task with its human partner, or it did not.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Google logo
Models

Google bakes computer use into Gemini 3.5 Flash as a native tool

DeepMind folds its standalone agent model into Flash, letting developers build agents that drive browsers, mobile apps and desktops via one API call.

Jaeden Schafer4 min read
Anthropic logo
Models

Anthropic logs 8x code merge jump as researchers benchmark AI gaming society's rules

Anthropic sees early signs of recursive self-improvement, a new benchmark tests AI loophole-hunting, and RL drones beat a human champion at 22 m/s.

Jaeden Schafer5 min read
Prime Intellect raises $130M at $1B to build enterprise AI agents
Business

Prime Intellect raises $130M at $1B to build enterprise AI agents

Radical Ventures led the Series A as the startup hit a $100M revenue run rate helping companies train their own models instead of renting frontier labs.

Jaeden Schafer4 min read