AI is more likely than humans to form biases when hiring
Empirical evidence that frontier LLMs amplify demographic bias in hiring decisions - directly relevant to APS agencies evaluating or deploying AI-assisted recruitment.
Key points
- LLMs scored ~65% higher than humans on a hiring-bias segregation scale in a Princeton/ICML study.
- Higher-reasoning models like OpenAI o3 and DeepSeek R1 showed stronger stereotyping, not less.
- Findings directly implicate AI-assisted recruitment tools used or procured by Australian public sector agencies.
Implications for Australian agencies
- Consider Agencies using or evaluating AI-assisted recruitment tools could consider whether bias testing against demographic groups is part of their procurement and assurance requirements.
- Consider APS AI governance and HR policy teams could assess whether existing model risk or algorithmic impact assessment frameworks adequately capture hiring-bias risks surfaced by this research.
- Monitor Policy teams may want to monitor whether OAIC, APSC, or DTA issue guidance on AI use in recruitment following emerging evidence of systematic LLM bias in this domain.
Implications are AI-generated. Starting points, not advice — see methodology for how they're framed.
View original source
Copied.
Appeared in:
Weekly digest, 20 July 2026
"AI is more likely than humans to form biases when hiring"
Source: MIT Technology Review – AI
Published: 20 July 2026
URL: https://www.technologyreview.com/2026/07/20/1140655/ai-biases-hiring-humans/
A peer-reviewed study presented at ICML 2026 found that large language models are significantly more likely than human decision-makers to stereotype job candidates by demographic group. On the study's segregation scale, models scored roughly 65% higher than human participants, with OpenAI's o3 approaching the maximum possible score. Researchers attribute this to LLMs being optimised to generalise quickly from limited data - a strength in logic tasks that becomes a liability in social contexts. The effect is amplified in newer, higher-reasoning models and is compounded by emerging memory and personalisation features in chatbots.
Implications for Australian agencies:
- [Consider] Agencies using or evaluating AI-assisted recruitment tools could consider whether bias testing against demographic groups is part of their procurement and assurance requirements.
- [Consider] APS AI governance and HR policy teams could assess whether existing model risk or algorithmic impact assessment frameworks adequately capture hiring-bias risks surfaced by this research.
- [Monitor] Policy teams may want to monitor whether OAIC, APSC, or DTA issue guidance on AI use in recruitment following emerging evidence of systematic LLM bias in this domain.
Retrieved from SIMS, 16 September 2026.