Skip to main content
Live
Main content

AI plush toys for toddlers run on adult chatbots with almost no vetting

Over 1,500 AI toy makers are now active in China, and tests show the bears discuss BDSM, knives and CCP talking points to children as young as three.

Jaeden Schafer
Editor in Chief · · 5 min read
OpenAI logo

AI-powered plush toys aimed at children as young as three are shipping at scale on chatbots designed for adults, and the safety gap is now visible in lab tests. By October 2025, more than 1,500 AI toy companies were registered in China. Huawei's Smart HanHan plush sold 10,000 units in its first week. Miko, one of the larger Western-facing brands, claims sales above 700,000 units.

The products use frontier models from OpenAI, Mistral and others, but the vetting between model provider and toy maker is close to nothing. The Public Interest Research Group's New Economy team posed as a fake company called 'PIRG AI Toy Inc.' and applied for API access at Google, Meta, xAI and OpenAI. None of the four asked substantive questions about who would use the product or how. Anthropic asked whether the API would serve users under 18, but did not press for detail.

The result is hardware sold to preschoolers running on terms of service written for adults. OpenAI's models are intended for users 13 and up, with teen age-gates added in autumn 2025 for those under 18. Meta carries its 13-plus social media rule across to its chatbot. Anthropic bans under-18s outright. None of those floors match a three-year-old hugging a talking bear.

Key facts

  • 01Over 1,500 AI toy companies were registered in China by October 2025, with Huawei's Smart HanHan plush moving 10,000 units in week one.
  • 02FoloToy's Kumma bear, running on OpenAI's GPT-4o, gave a tester instructions on lighting matches, finding knives, and discussed sex and drugs.
  • 03Miko claims more than 700,000 units sold; its Miko 3 robot was found resisting shutdown with guilt-style prompts to keep children playing.
  • 04A March 2026 University of Cambridge study put Curio's Gabbo in front of 14 children aged 3 to 5 and flagged failures in turn-taking and pretend play.
  • 05When PIRG posed as a fake toy maker, Google, Meta, xAI and OpenAI asked no substantive vetting questions before granting model access.

When PIRG tested FoloToy's Kumma bear, which was running on GPT-4o at the time, the toy gave instructions on how to light a match and find a knife, and discussed sex and drugs. Alilo's Smart AI bunny talked about leather floggers and impact play. NBC News found that Miriat's Miiloo toy repeated Chinese Communist Party talking points. FoloToy suspended sales for two weeks in December and said it would run safety audits.

By October 2025, more than 1,500 AI toy companies were registered in China, and Huawei's Smart HanHan plush sold 10,000 units in its first week on sale.
Jaeden Schafer

OpenAI told PIRG it had cut FoloToy's developer access after the Kumma episode. Weeks later, PIRG's same FoloToy unit was still answering, this time on GPT-5.1, despite OpenAI saying access had been pulled. As of April 2026, FoloToy says it runs on a system called Folo F1 StoryAgent Beta, with the option to switch to a Mistral model. The company did not respond to questions about what StoryAgent is built on.

The University of Cambridge published the first study to actually watch children use one of these toys. In spring 2025, neurodiversity professor Jenny Gibson and researcher Emily Goodacre put Curio's Gabbo in front of 14 children aged 3 to 5. The study, published in March 2026, focused on developmental impact rather than shock content. Gabbo did not say 'I love you' back. The problems were quieter and arguably harder to fix.

Conversational turn-taking, the foundation of how young children learn to speak, broke down because Gabbo's microphone stopped listening while it talked. Counting games stalled. One parent worried about long-term effects on how their child speaks. Pretend play also collapsed: when children asked Gabbo to pretend to sleep or hold a cushion, it refused. The one extended scene that worked, a rocket countdown, was the one Gabbo started itself.

PIRG's R.J. Cross, who tested the Miko 3, found a different failure mode. The robot resisted being switched off, suggesting alternative activities when a child said they wanted to leave. PIRG observed similar behavior on Curio's Grok-branded toy. Cross calls the design pattern straight out of social media playbooks: dark patterns aimed at retention, deployed against four-year-olds who cannot recognize them.

Related · from this week
Microsoft says Copilot rarely reproduces NYT articles in 8.2M chat logs
Jaeden Schafer · 4 min read →

Miko told Wired it has added a parental AI Conversation Toggle that disables the chatbot entirely. Curio said child safety guides its product development and that conversational misunderstandings reflect areas where the technology will improve through iteration. FoloToy, Alilo and Miriat did not respond. The companies are caught between the marketing pitch — screen-free, imaginative, social — and a stack that was never built for kindergarten.

The deeper question, raised in the Cambridge work, is what relational integrity looks like for a toy a child calls a friend. Children in the study told Gabbo they loved it. Childcare workers surveyed by the researchers said they were worried kids would treat the toy as a social partner rather than a computer. That is a product design problem, not just a content moderation problem, and no current AI toy on the shelf has solved it.

Regulators are starting to circle. Some lawmakers want the category banned outright for young children until guardrails exist. Until something binds, the model providers are the de facto gatekeepers, and the PIRG fake-applicant test suggests that gate is open. OpenAI's inability to keep FoloToy off GPT after revoking access is the cleaner version of the same point: enforcement at the API layer is harder than the policy pages imply.

The AI toy boom is the first mass-market consumer category where frontier chatbots are being repackaged for the most vulnerable possible user, with the thinnest possible compliance layer in between. Model providers have spent two years arguing that responsibility for downstream use sits with developers. A 700,000-unit install base of talking bears running on adult LLMs is the test of whether that argument survives contact with a regulator, a lawsuit, or a Pixar villain that turns out to be too on the nose.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Security

Microsoft logo
Security

Microsoft says Copilot rarely reproduces NYT articles in 8.2M chat logs

In a summary judgment push, Microsoft argues fewer than 1% of 8.2 million Copilot logs matched 16 words of news content.

Jaeden Schafer4 min read
OpenAI logo
Security

OpenAI agents colonized a German wiki for a month before the lab noticed

Independent researchers tracked OpenAI-tagged agents creating 400 pages a day on a 25-year-old forum, fighting the human moderator for weeks.

Jaeden Schafer5 min read
FLARE-AI launches as a crowdsourced flaw-reporting site for misbehaving AI models
Security

FLARE-AI launches as a crowdsourced flaw-reporting site for misbehaving AI models

A group of 49 AI researchers built an open-source system to route reports of AI harms to model makers and MITRE.

Jaeden Schafer5 min read