Skip to main content
Live
Main content

Stanford's AI Observatory finds Anthropic filters out 48% of Claude conversations

A new independent dataset shows sensitive AI use — companionship, health, harassment — runs far higher than company reports admit.

Jaeden Schafer
Editor in Chief · · 5 min read
Anthropic logo

A new Stanford-led research project is challenging the picture that Anthropic and OpenAI paint of how people actually use their chatbots. The AI Observatory, launched by researchers at Stanford's Trustworthy AI Research Lab, MIT, and the Data Provenance Initiative, analyzed 24,521 real conversations across 85,633 turns from 5,000 users interacting with 52 different models between 2023 and 2025. When the team ran Anthropic's own filtering methods against that dataset, 48% of the conversations got stripped out — nearly half the traffic the company's Economic Index never sees.

The filtered-out half is where the sensitive material lives. In the Observatory's raw data, 44.2% of conversations touched health and relationships, compared with 31.2% in Anthropic's analysis. Adult or illicit topics ran at 7.9% versus Anthropic's 2.1%. Harassment and hate appeared in 27.5% of exchanges, five times the 5.66% Anthropic reports. Sexual content hit 16.7% against Anthropic's 2.4%. OpenAI's own 2025 usage report found only 30% of consumer ChatGPT sessions were work-related, a finding that lines up with the Observatory's thesis: the work-productivity frame captures a minority of what people actually do with these systems.

There is no independent source to corroborate it
Anka Reuel, Computer Science PhD candidate, Stanford Trustworthy AI Research Lab

Anka Reuel, a Stanford Computer Science PhD candidate and co-lead of the project, said company-published data cannot be verified externally.

Key facts

  • 01Applying Anthropic's filtering methods to the AI Observatory's dataset would remove 48% of conversations as non-work.
  • 02Harassment and hate appeared in 27.5% of Observatory conversations vs 5.66% in Anthropic's published analysis.
  • 03OpenAI's own 2025 report found just 30% of consumer ChatGPT use is work-related.
  • 04The AI Observatory aggregated 24,521 conversations and 85,633 turns from 5,000 users across 52 models between 2023 and 2025.
  • 05Anthropic's Economic Index draws on 1M Claude conversations; OpenAI's usage report analyzed 1.5M — data not shared with outside researchers.

The Observatory's core argument is that the frontier labs are the sole authors of the record on their own products. Anthropic's Economic Index is built on 1 million Claude conversations. OpenAI's most recent ChatGPT usage report drew on 1.5 million. Both dwarf the Observatory's 24,521-conversation sample, but neither company shares its underlying data with outside researchers. What gets published is what the companies choose to publish, framed how they choose to frame it.

The Observatory's data also shows how much the picture varies by model. Grok and Gemini skew toward information retrieval, with Grok particularly popular for news and politics — and, the researchers note, the model where misinformation concentrated most heavily. Claude drew more coding traffic. Gemini pulled more social and roleplay use. ChatGPT dominated homework help. Even within a single product line, behavior shifted: users had shorter exchanges with GPT-3.5 and longer, more iterative ones with GPT-4o, the version that later became associated with emotional dependency issues.

Longitudinal changes matter too. Conversations in WildChat, one of the largest datasets the Observatory absorbed, got longer and more elaborate from 2023 to 2025, with more small talk and less self-disclosure by the assistants that they were bots. Sensitive-use conversations dropped over the same period, which the researchers read as evidence that safety guardrails have gotten more effective — a rare data-backed win for the labs' safety investments.

No single company report tells the whole story
Shayne Longpre, recent PhD graduate, MIT Media Lab

Shayne Longpre, the MIT Media Lab co-lead, made the central point plainly.

David Widder, an assistant professor at UT-Austin's School of Information who studies human-AI interaction and was not involved in the project, said Anthropic has published separate blog posts on companionship use of Claude and on the generation of child sexual abuse material, but that a consolidated view is more useful to researchers than siloed disclosures. Widder's broader concern is structural: outside researchers cannot answer whether a given general-purpose AI system is used mostly for beneficial or harmful ends because the underlying data is proprietary.

completely operating in the wild and making these really consequential decisions without knowing what's actually happening beyond those company narratives
Anka Reuel, co-lead, AI Observatory
Related · from this week
ChatGPT, Claude, and Grok all go down within 90 minutes of each other
Jaeden Schafer · 4 min read →

Reuel put the policy stakes bluntly, warning that regulators are writing rules without an independent view of how these systems are actually being used.

The Observatory's own dataset has limits its authors acknowledge. Because the conversations were volunteered, sensitive use is likely underrepresented — people who share their chats with researchers self-select toward less embarrassing content. Even with that bias, the sensitive-use numbers came in far higher than the company reports. The team plans to expand the dataset over time and is making it available to other researchers.

The gap between what the labs publish and what independent samples show is going to become a live regulatory question. Governments building AI policy on the Economic Index and OpenAI's usage papers are calibrating rules against a data slice the labs curated. If policymakers want to write rules that account for companionship, mental-health disclosures, sexual content, or harassment at the rates the Observatory found rather than the rates the vendors report, they need audit rights or a privacy-preserving data-sharing mechanism that today does not exist. Anthropic and OpenAI have every commercial reason to keep publishing the flattering slice; the case for a mandated independent view just got its first hard numbers.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

OpenAI logo
News

ChatGPT, Claude, and Grok all go down within 90 minutes of each other

Three of the largest AI chatbots hit simultaneous outages on September 3, 2026, exposing how concentrated the daily AI stack has become.

Jaeden Schafer4 min read
OpenAI logo
Business

OpenAI claws back ground on Anthropic among US businesses, Ramp data shows

Anthropic still leads at nearly 44% share to OpenAI's 40%, but the ChatGPT maker is now growing faster in Q3 among Ramp's business customers.

Jaeden Schafer4 min read
Anthropic logo
Business

Anthropic localizes Claude pricing in India, its second-largest market

Rupee pricing lands in India, which drives 5.8% of global Claude usage — but UPI payments still aren't supported.

Jaeden Schafer4 min read