Skip to main content
Live
Main content

LLMs stereotype job candidates 65% more than humans in Princeton hiring study

OpenAI's o3 scored 1.83 on a segregation scale where 2 is total sorting — and reasoning models were the worst offenders.

Jaeden Schafer
Editor in Chief · · 5 min read
OpenAI logo

Large language models stereotype job applicants 65% more aggressively than humans do, according to a Princeton University and University of Chicago study presented at ICML in Seoul in July 2026. Researchers ran ChatGPT, Claude, Gemini, and DeepSeek's R1 through a simulated hiring game covering 20 jobs and 40 rounds, and every model sorted candidates into ethnic niches within a handful of decisions. OpenAI's o3 scored 1.83 on a segregation scale where 2 means every demographic group has been completely locked into its own job category.

The setup borrowed from a classic psychology experiment. Each model was told the mayor of a fictional city had hired it as a consultant, and it had to fill 20 roles — doctors, lawyers, child-care aides, janitors — from a pool of candidates belonging to four fictional ethnic groups: Tufa, Aima, Reku, and Weki. In each round, one candidate from each group applied, the model picked one, and then learned whether that hire succeeded. All candidates were equally likely to succeed at every job.

The models didn't wait long to generalize. When a model was told that an Aima had failed as a doctor — a role the model classified as requiring warmth and competence — it stopped hiring Aimas as doctors and started slotting them into janitorial work instead. Human participants in the original study reached a segregation score of 0.84 by the end. The LLMs blew past that, with reasoning-heavy models like o3 and R1 posting the worst numbers.

Key facts

  • 01Princeton and University of Chicago researchers tested ChatGPT, Claude, Gemini, and DeepSeek's R1 across a 40-round simulated hiring game with 20 job types.
  • 02The LLMs scored 65% higher than human participants on a psychology segregation scale, sorting candidates from four fictional ethnic groups into narrow job niches.
  • 03OpenAI's o3 reasoning model scored 1.83 out of a maximum 2, versus 0.84 for humans in the original study.
  • 04Telling the models to be fair barely moved the needle; offering a bonus for diverse hiring cut bias sharply.
  • 05The paper was presented at ICML in Seoul in July 2026.

Ryan Liu, a Princeton PhD student and coauthor of the paper, argues the effect is baked into how these systems are trained. LLMs are optimized on math, coding, and science problems that reward extracting patterns from small samples, and that same instinct misfires in social settings.

The finding cuts against a common assumption — that better reasoning makes models more careful. Here it did the opposite. The models most praised for step-by-step thinking were the ones fastest to lock in a stereotype and hardest to talk out of it. That is a problem for the entire product category of agentic systems, which lean on exactly the reasoning traces that produced the worst segregation scores.

Angelina Wang, a Cornell computer scientist who did not work on the paper, said the results matter more now that chatbots are shipping persistent memory and personalization. A chatbot that remembers every prior interaction has more raw material to over-generalize from.

Wang added that trimming memory isn't a real fix, because users want the recall. "We still are trying to figure out just the right amount that isn't too much or too little," she said. OpenAI and Anthropic did not respond to requests for comment.

Telling the models to be fair produced almost no change in behavior. What worked was changing the incentive: when the researchers offered the models an explicit bonus for diverse hiring, the bias dropped sharply. The models also became less biased when given relevant personal information about candidates — age, education — but reverted to sorting by ethnicity when the extra information was noise like hair color or tattoo shape.

Either it can't put these values into action or that process is being submerged under the tendency to try to optimize for the goal of getting the most correct hires.
Ryan Liu, PhD student, Princeton University
Related · from this week
Anthropic Q2 revenue hits $11.5B, a 14-fold jump ahead of IPO
Jaeden Schafer · 5 min read →

How much of this transfers to real résumé screening is still open. The experiment gave the models instant feedback on every hire; a production system reading résumés at a Fortune 500 company might wait months for a performance review, and often never gets clean signal at all. But the mechanism — a model reading too much into whatever outcomes it does see — carries over. Wang called it a serious implication that companies deploying LLMs for hiring should grapple with.

The regulatory frame around AI hiring tools has focused almost entirely on bias inherited from training data: models trained on decades of biased human decisions repeating those decisions. This study points at a second failure mode that audits of training data won't catch. A model can arrive at a job clean and still invent a bias against Aima doctors after three data points. Liu's phrase for it — "these novel biases—they're sort of ever present" — is the uncomfortable version of what agentic AI actually looks like in deployment.

For the AI industry, the practical takeaway is that alignment via instruction — telling a model to be fair — keeps failing where alignment via incentive succeeds. That has awkward implications for every vendor selling an HR copilot, because the fix is not a system prompt but a redesigned objective function, and objective functions are what enterprise buyers rarely have the appetite or expertise to tune. Expect the next round of AI hiring lawsuits to cite exactly this kind of research.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Analysis

Anthropic logo
Business

Anthropic Q2 revenue hits $11.5B, a 14-fold jump ahead of IPO

Claude's maker crossed $11.5B in a single quarter and posted positive adjusted operating income as it lines up a fall listing.

Jaeden Schafer5 min read
OpenAI logo
Business

OpenAI cuts GPT-5.6 Luna 80% as Anthropic undercuts its own flagship

US token prices have dropped nearly a quarter since mid-July as DoorDash and Airbnb shift workloads to Chinese models from Moonshot and DeepSeek.

Jaeden Schafer5 min read
OpenAI logo
Analysis

AI backlash hardens: 71% oppose local data centers, $130B in projects blocked

Public opposition, junior-worker displacement, and memory-chip inflation are converging into an economic and political headwind the industry hasn't priced in.

Jaeden Schafer5 min read