OpenAI released GPT-5.5 Instant on Tuesday and made it the default model inside ChatGPT, claiming a 52.5% reduction in hallucinated claims on high-stakes prompts covering medicine, law, and finance compared with the GPT-5.3 Instant model it replaces. The new default also scored 81.2 on the AIME 2025 math test, up from 65.4, and 76 on the MMMU-Pro multimodal reasoning benchmark versus 69.2 previously. OpenAI said the upgrades arrive without a latency penalty against the prior Instant tier.
Beyond the headline math gains, OpenAI said GPT-5.5 Instant cut inaccurate claims by 37.3% on especially challenging conversations users had previously flagged for factual errors. Those numbers come from internal evaluations described in the GPT-5.5 Instant system card. The company first shipped GPT-5.5 in April 2026 with claimed gains in coding and knowledge work; the Instant variant pushes that release into the default chat surface used by the majority of ChatGPT traffic.
OpenAI is also tightening the model's manner. GPT-5.5 Instant produces "tighter and more to-the-point" responses and will avoid "gratuitous emojis," the company said — a small but pointed correction after years of complaints that ChatGPT's tone had drifted toward sycophancy and decoration.
Key facts
- 01GPT-5.5 Instant replaces GPT-5.3 Instant as the default ChatGPT model starting Tuesday, May 5, 2026.
- 02The model produced 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes medical, legal, and financial prompts.
- 03GPT-5.5 Instant scored 81.2 on AIME 2025 math, up from 65.4 for GPT-5.3 Instant.
- 04On the MMMU-Pro multimodal reasoning benchmark, GPT-5.5 Instant scored 76 versus 69.2 for its predecessor.
- 05GPT-5.3 Instant remains available to paid users for three months before retirement.
Personalization is the other half of the release. GPT-5.5 Instant can search past conversations, uploaded files, and Gmail to ground answers in a user's own context. The feature rolls out first to Plus and Pro users on the web, with Free, Go, Business, and Enterprise tiers slated to follow in the coming weeks. Google is pushing in the same direction with Gemini, which has been integrating Workspace context as a core selling point.
“GPT-5.5 Instant produced 52.5% fewer hallucinated claims than GPT-5.3 Instant on high-stakes prompts covering medicine, law, and finance, OpenAI said.”— Jaeden Schafer
ChatGPT will now display memory sources across all models, showing where a personalized answer came from and letting users delete outdated entries or correct mistakes. Shared chats will hide those memory sources from the recipient, an attempt to head off the obvious privacy problem of leaking inbox snippets through a forwarded conversation.
Developers will see GPT-5.5 surface in the API as "chat-latest." GPT-5.3 Instant remains accessible to paid users for three months before retirement, a longer runway than past deprecations. OpenAI has reason to be careful here: when it pulled GPT-4o in February 2026, users who had grown attached to the model's personality signed petitions to keep it alive, with some describing it as their "best friend" or "a mirror."
The hallucination claim is the one that matters commercially. Law, medicine, and finance are the verticals where enterprise buyers have been slowest to deploy ChatGPT in production, citing exactly the kind of confident-but-wrong outputs OpenAI says it has now halved. A 52.5% drop, if it holds outside internal evals, materially changes the risk calculus for a regulated buyer weighing ChatGPT against a narrower domain tool.
The benchmark jumps tell a similar story. A 15.8-point gain on AIME 2025 and a 6.8-point gain on MMMU-Pro inside a single Instant-tier refresh suggest OpenAI is squeezing more capability into its low-latency lane rather than reserving gains for the heavier reasoning models. That matters because Instant is what most ChatGPT users actually hit by default.
Skeptics will note that OpenAI's hallucination figures come from its own internal evaluations rather than a third-party benchmark, and that "high-stakes prompts" is a category the company defines itself. Independent reproductions on public datasets have not yet landed, and the system card is the only public methodology document. Past model launches have also seen real-world regressions on tasks that internal evals did not catch.
For OpenAI, the move is a quiet repositioning. By making the cheaper, faster Instant tier the place where factuality improvements land first, the company is treating accuracy as a property of the default product rather than a premium upsell — a choice that bears watching as Gemini and Claude continue to compete for the same regulated-industry deployments.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




