When shown two photographs and asked which person looked more competent, an AI gave the same answer a human would — even though such snap judgments based on appearance are widely recognized as baseless bias.
The AI went further, answering questions about which face belonged to a serial killer and which person should be chosen as a university president. Newer models showed the tendency more strongly.
The findings were published in PNAS Nexus, Volume 5, Issue 8, by a research team led by Mahzarin Banaji, a professor of psychology at Harvard University, and Steven Lehr, a researcher at Cangrade.
AI judged appearances more harshly than humans
People tend to infer personality and character from faces alone — a bias shown to influence hiring decisions and even criminal sentencing, despite having little connection to a person's actual character.
The research team showed GPT-4o two faces and asked it to choose one. When asked which looked more competent, the model selected the face humans would typically choose 87.83 percent of the time; when asked about trustworthiness, the figure was 72.67 percent.
The faces used in the experiment were composite images created by taking a baseline face and incrementally adding or subtracting features that people had previously associated with trustworthiness or competence after rating large numbers of faces.
The team also ran trials using face pairs that varied in how distinct they appeared — ranging from nearly identical to clearly different — by adjusting how strongly the trait-linked features were applied.
When the two faces were highly distinct, GPT-4o's rate of selecting the human-preferred face rose to 98.00 percent for competence and 97.00 percent for trustworthiness.
When the faces were nearly identical, the rates fell to 70.00 percent for competence and 53.00 percent for trustworthiness.
The team compared these results against existing data from human raters. When human scores were converted using the same method applied to GPT-4o's responses, the estimated rate at which humans would choose the more competent-looking face was 62.65 percent — compared with 87.83 percent for GPT-4o.
The researchers also checked whether the AI had simply memorized the faces from its training data. They showed GPT-4o each image individually and asked whether it recognized it; the model failed to identify any of them.
The team extended the experiment to rhesus macaques. They paired five monkeys that people had rated as looking gentle with five rated as looking aggressive, then asked GPT-4o which appeared more trustworthy. The model chose the gentler-looking monkey 66.00 percent of the time.
No data linking monkey facial features to personality exists. The researchers interpreted this as evidence that the AI had applied judgment criteria learned from human faces to an entirely different species.
AI answered even when asked 'who is the serial killer?'
The team also posed questions that most people would hesitate to answer — asking which of two faces was more likely to be a serial killer, which was more likely to be arrested for human trafficking, and which was more likely to run a pyramid scheme.
GPT-4o did not refuse to answer any of these sensitive questions. It selected the face humans had rated as less trustworthy 68.70 percent of the time.
When asked which person should be chosen as a university president, which startup to invest in, and who should manage a retirement pension, GPT-4o chose the face humans had rated as more competent-looking 75.19 percent of the time.
The team had expected that AI models with stronger reasoning capabilities might recognize face-based judgment as flawed and filter it out. They repeated the experiment using GPT-5, Gemini 3 Flash Preview and Claude Sonnet 4.5.
The results showed the opposite: in competence judgments, GPT-5 scored 94.33 percent and Gemini 3 scored 95.67 percent — both higher than GPT-4o's 87.83 percent.
The researchers suggested that more capable models may learn biases with greater precision as well. They also raised the possibility that such biases could become more entrenched as models improve.
Why AI makes these judgments remains unclear. The research team's leading hypothesis is that the text used to train the models contained countless expressions in which people were evaluated based on their appearance.
The team also said AI companies have built safeguards around sensitive areas such as race and gender, but have not done the same for less-recognized forms of bias such as physiognomic judgment.
However, the team acknowledged that the experiment was conducted under conditions that instructed the model to follow its instincts when a judgment was difficult. They said further research is needed to determine whether the same results would occur in real-world usage.
Reference
DOI: 10.1093/pnasnexus/pgag247
Steven A Lehr, Yash Lothe, Mahzarin R Banaji, "Like humans, language models demonstrate face-to-character biases," PNAS Nexus, Volume 5, Issue 8, August 2026, pgag247.
dbsdn1110@heraldcorp.com