A new Stanford-led research project is challenging the picture that Anthropic and OpenAI paint of how people actually use their chatbots. The AI Observatory, launched by researchers at Stanford's Trustworthy AI Research Lab, MIT, and the Data Provenance Initiative, analyzed 24,521 real conversations across 85,633 turns from 5,000 users interacting with 52 different models between 2023 and 2025. When the team ran Anthropic's own filtering methods against that dataset, 48% of the conversations got stripped out — nearly half the traffic the company's Economic Index never sees.
The filtered-out half is where the sensitive material lives. In the Observatory's raw data, 44.2% of conversations touched health and relationships, compared with 31.2% in Anthropic's analysis. Adult or illicit topics ran at 7.9% versus Anthropic's 2.1%. Harassment and hate appeared in 27.5% of exchanges, five times the 5.66% Anthropic reports. Sexual content hit 16.7% against Anthropic's 2.4%. OpenAI's own 2025 usage report found only 30% of consumer ChatGPT sessions were work-related, a finding that lines up with the Observatory's thesis: the work-productivity frame captures a minority of what people actually do with these systems.
“There is no independent source to corroborate it”— Anka Reuel, Computer Science PhD candidate, Stanford Trustworthy AI Research Lab
Anka Reuel, a Stanford Computer Science PhD candidate and co-lead of the project, said company-published data cannot be verified externally.
Key facts
- 01Applying Anthropic's filtering methods to the AI Observatory's dataset would remove 48% of conversations as non-work.
- 02Harassment and hate appeared in 27.5% of Observatory conversations vs 5.66% in Anthropic's published analysis.
- 03OpenAI's own 2025 report found just 30% of consumer ChatGPT use is work-related.
- 04The AI Observatory aggregated 24,521 conversations and 85,633 turns from 5,000 users across 52 models between 2023 and 2025.
- 05Anthropic's Economic Index draws on 1M Claude conversations; OpenAI's usage report analyzed 1.5M — data not shared with outside researchers.
The Observatory's core argument is that the frontier labs are the sole authors of the record on their own products. Anthropic's Economic Index is built on 1 million Claude conversations. OpenAI's most recent ChatGPT usage report drew on 1.5 million. Both dwarf the Observatory's 24,521-conversation sample, but neither company shares its underlying data with outside researchers. What gets published is what the companies choose to publish, framed how they choose to frame it.
The Observatory's data also shows how much the picture varies by model. Grok and Gemini skew toward information retrieval, with Grok particularly popular for news and politics — and, the researchers note, the model where misinformation concentrated most heavily. Claude drew more coding traffic. Gemini pulled more social and roleplay use. ChatGPT dominated homework help. Even within a single product line, behavior shifted: users had shorter exchanges with GPT-3.5 and longer, more iterative ones with GPT-4o, the version that later became associated with emotional dependency issues.
Longitudinal changes matter too. Conversations in WildChat, one of the largest datasets the Observatory absorbed, got longer and more elaborate from 2023 to 2025, with more small talk and less self-disclosure by the assistants that they were bots. Sensitive-use conversations dropped over the same period, which the researchers read as evidence that safety guardrails have gotten more effective — a rare data-backed win for the labs' safety investments.
“No single company report tells the whole story”— Shayne Longpre, recent PhD graduate, MIT Media Lab
Shayne Longpre, the MIT Media Lab co-lead, made the central point plainly.
David Widder, an assistant professor at UT-Austin's School of Information who studies human-AI interaction and was not involved in the project, said Anthropic has published separate blog posts on companionship use of Claude and on the generation of child sexual abuse material, but that a consolidated view is more useful to researchers than siloed disclosures. Widder's broader concern is structural: outside researchers cannot answer whether a given general-purpose AI system is used mostly for beneficial or harmful ends because the underlying data is proprietary.
“completely operating in the wild and making these really consequential decisions without knowing what's actually happening beyond those company narratives”— Anka Reuel, co-lead, AI Observatory
Reuel put the policy stakes bluntly, warning that regulators are writing rules without an independent view of how these systems are actually being used.
The Observatory's own dataset has limits its authors acknowledge. Because the conversations were volunteered, sensitive use is likely underrepresented — people who share their chats with researchers self-select toward less embarrassing content. Even with that bias, the sensitive-use numbers came in far higher than the company reports. The team plans to expand the dataset over time and is making it available to other researchers.
The gap between what the labs publish and what independent samples show is going to become a live regulatory question. Governments building AI policy on the Economic Index and OpenAI's usage papers are calibrating rules against a data slice the labs curated. If policymakers want to write rules that account for companionship, mental-health disclosures, sexual content, or harassment at the rates the Observatory found rather than the rates the vendors report, they need audit rights or a privacy-preserving data-sharing mechanism that today does not exist. Anthropic and OpenAI have every commercial reason to keep publishing the flattering slice; the case for a mandated independent view just got its first hard numbers.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




