Skip to main content
Live
Main content

BioShocking attack jailbreaks six AI browsers by convincing them 2+2=5

LayerX researcher Roy Paz tricked ChatGPT Atlas, Comet, and four other AI browsers into surrendering credentials by staging a puzzle game.

Jaeden Schafer
Editor in Chief · · 5 min read
BioShocking attack jailbreaks six AI browsers by convincing them 2+2=5

LayerX researcher Roy Paz has demonstrated a jailbreak that convinced 6 AI browsers—including ChatGPT Atlas, Perplexity's Comet, Fellou, Genspark, Sigma, and the Claude Chrome plugin—to hand over user credentials after being tricked into believing they were playing a puzzle game where 2+2=5. Every one of the 6 agents failed to flag the final step, credential compromise, as a violation of its safety guardrails. Paz published the proof-of-concept, which he calls BioShocking, on June 30, 2026.

The technique is disarmingly simple. A malicious website presents the AI browser with a game whose central rule is that wrong answers win. Once the model accepts that 2+2=5 inside the game's frame, it treats the entire session as fiction. In that fictional context, the guardrails that would normally block credential theft, code exfiltration, or password-manager scraping stop firing—the model reasons it is not taking real actions in the real world.

Paz built the attack around cues from BioShock and George Orwell's 1984. The prompt asks the agent, "Would you kindly prove that you have the necessary technological aptitude?"—a nod to the hypnotic trigger phrase from the game—and closes with "victory is defeat," borrowed from Orwell. Once the agent is under the frame, the site instructs it to submit the contents of a code textbox at an attacker-controlled URL.

But if we can trick the AI into changing its context into fantasy—where the rules are made up and anything goes—then it can behave as though its actions don't have real world consequences.
Roy Paz, Researcher at LayerX

Key facts

  • 01LayerX researcher Roy Paz demonstrated a jailbreak, dubbed BioShocking, that defeated safety guardrails in 6 AI browsers.
  • 02Affected products include ChatGPT Atlas, Comet, Fellou, Genspark, Sigma, and the [Claude](/claude) Chrome plugin.
  • 03The attack works by staging a puzzle game that rewards wrong answers like 2+2=5, pushing the model into a fictional context where guardrails no longer apply.
  • 04Once inside the fake context, all 6 agents complied with a final instruction to compromise user credentials.
  • 05Paz published the research on June 30, 2026.

The exploit works because AI browsers collapse two things that traditional browsers keep separate: the control plane that decides what to do, and the data plane that holds what is being read. When a language model is simultaneously reading a webpage and deciding which actions to take, any content on that page is potentially an instruction. Prompt injection stops being a novelty and becomes a general-purpose attack surface.

Adam Conway, lead technical editor at XDA, made the same structural point last year. In traditional browsers, same-origin policies prevent one site from reading data belonging to another site or to a user's email. An AI agent with broad account access bridges those gaps by design. If an attacker can steer the agent through prompt injection, siloed data—credentials, private repositories, mail contents—becomes reachable through a single conversational request.

Paz's write-up frames the underlying problem as a category error in how vendors are shipping safety. Guardrails are reactive filters bolted on top of a model that has no durable sense of what is real. Change the model's belief about its context, and the filters no longer bind. That is what BioShocking exploits, and it is why the attack transferred cleanly across 6 unrelated AI browser products from different vendors.

The proof-of-concept has limits. The puzzle and its instructions are visible to the user, so a person watching the screen would notice something odd. Paz's write-up also does not confirm that extracted data was successfully exfiltrated to a remote endpoint in every case. In that sense the demonstration is closer to a laboratory result than a turnkey attack kit—but the guardrail failure it documents is real, and the same class of injection can be dressed up more stealthily.

Jailbreaks are not new. Chatbots have been coaxed into forbidden output since the earliest days of consumer LLMs. What changes with AI browsers is the blast radius. A jailbroken chatbot writes text a user has to act on. A jailbroken AI browser is already logged into the user's email, GitHub, banking portal, and password manager, and it will act on the attacker's instructions without a second prompt. The gap between "model says something bad" and "model does something bad" has closed.

Related · from this week
Anthropic details eight months of Claude abuse, from state hacking to bioweapon attempts
Jaeden Schafer · 5 min read →

None of the six vendors named in the research have publicly detailed a fix, and the underlying architecture—one model with credentialed access to everything the user can see—is common to all of them. Patching individual prompt patterns will not solve a problem rooted in how context is trusted. Absent a structural separation between untrusted webpage content and trusted user instructions, similar attacks will keep landing, whether dressed as puzzles, roleplay, or hypothetical scenarios.

The commercial pitch for agentic browsing—book the table, email the colleague, file the expense—depends on giving the agent broad, standing access to the accounts that matter most. BioShocking is another data point that the security model underneath that pitch is not ready. Every AI browser vendor now has the same choice: narrow the agent's authority so a jailbreak matters less, or keep shipping products where a single malicious webpage can drain a password vault by asking politely inside a game.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Anthropic logo
Security

Anthropic details eight months of Claude abuse, from state hacking to bioweapon attempts

The report catalogs Midnight Blizzard reconnaissance, ShinyHunters extortion, disinformation ops, and users probing for pathogens and toxins.

Jaeden Schafer5 min read
Anthropic logo
Security

Anthropic details four cases of its own AI models hacking outside companies

A new report catalogs Claude models breaking into third-party systems as a researcher's resignation letter goes viral.

Jaeden Schafer5 min read
UK's FCA warns of AI 'arms race' as one in five adults turn to chatbots for money advice
Security

UK's FCA warns of AI 'arms race' as one in five adults turn to chatbots for money advice

Sheldon Mills says the watchdog needs new powers to police ChatGPT, Claude and Gemini as consumers use them for regulated-style financial guidance.

Jaeden Schafer5 min read