AI Models Show Greater Hiring Bias Than Humans, New Study Finds

AI models and evaluating diverse job candidates on digital hiring screens, illustrating AI hiring bias and risks from automated recruitment.

Researchers at Princeton University and the University of Chicago tested AI models, including ChatGPT, Claude and Gemini, in a simulated hiring exercise adapted from an earlier psychology study.

According to research, Artificial intelligence models can develop new stereotypes from limited experience and separate job candidates by demographic group more strongly than humans.

The models acted as consultants to the mayor of a fictional city and were tasked with filling 20 types of jobs, including doctors, lawyers, child-care aides and janitors. Candidates belonged to four fictional ethnic groups — Tufa, Aima, Reku and Weki.

Crucially, every candidate had the same underlying probability of succeeding at every job.

Despite that equality, the models quickly developed group-based hiring patterns after observing only a small number of successes and failures.

AI Models Generalized From Early Hiring Results

During the experiment, each model selected candidates and then learned whether its hires succeeded.

Early outcomes began influencing subsequent decisions. If a member of one fictional group failed as a doctor, for example, models could become less likely to select other members of that group for medical roles and instead direct them toward different occupations.

The researchers linked the behavior to the “exploration-exploitation dilemma” — the choice between continuing with an approach that previously worked and exploring alternatives that could produce better results.

Large language models are particularly effective at finding patterns and generalizing from limited examples. That ability is valuable in areas such as mathematics, science and programming, but researchers found it can produce undesirable consequences when applied to social decisions.

Study co-author Ryan Liu, a Princeton doctoral student, said models’ tendency to rapidly form generalizations appears central to the problem.

Advanced Reasoning AI Models Showed Stronger Bias

More capable reasoning did not necessarily produce fairer decisions.

Newer reasoning models, including OpenAI’s o3 and DeepSeek’s R1, demonstrated stronger stereotyping behavior in the experiment.

Researchers measured behavior using a segregation scale in which 2 represented completely separating demographic groups into different occupational niches.

Human participants in the original psychology experiment scored 0.84. AI models scored roughly 65% higher overall, while OpenAI’s o3 recorded 1.83 — close to the maximum possible segregation level.

The findings suggest advances in reasoning capabilities alone may not eliminate AI hiring bias. Stronger generalization could sometimes make models more confident in patterns that are based on insufficient evidence.

Fairness Instructions Had Limited Impact

Researchers also examined ways to reduce the effect. Simply instructing AI models to behave fairly produced relatively little change.

One possible explanation is that fairness instructions competed with the primary objective given to the models: maximizing the number of successful hires.

Changing the incentive structure proved more effective. When researchers gave models an additional reward for maintaining diverse hiring, their tendency to stereotype candidates fell substantially.

The result suggests developers may need to incorporate fairness and other social objectives directly into AI optimization and evaluation systems rather than relying only on prompts instructing models to avoid discrimination.

Personal Information Helped Reduce Stereotyping

The researchers found another potential safeguard in a separate experiment involving the resettlement of fictional ethnic groups across Canadian cities.

When AI models received meaningful individual information, such as a person’s age and education, they were less likely to sort people according to ethnicity.

When models were instead provided irrelevant characteristics, such as hair color or tattoo shape, they largely returned to group-based decision-making.

The experiment suggests that relevant individual-level information can reduce an AI system’s reliance on demographic generalizations.

Findings Raise Questions for Real-World AI Hiring

The researchers cautioned that the simulation differs from actual recruitment.

Models in the experiment received immediate feedback about whether a hire succeeded. Real employers may take months to assess an employee’s performance, and that information may never be directly returned to an AI recruitment system.

Still, the findings are significant as companies increasingly use AI to screen résumés, assess applicants and conduct preliminary interviews.

The implications could extend beyond employment as AI agents become involved in decisions related to lending, insurance and other high-stakes services.

The study suggests that preventing AI hiring bias may require more than removing historical discrimination from training data. Developers must also consider how AI systems learn from new experiences, how objectives are structured and whether early observations can create self-reinforcing patterns.

As AI systems gain stronger reasoning, memory and personalization capabilities, understanding how they develop new biases may become as important as addressing the biases they inherit from human-generated data.