LLMs have been set their toughest test yet. What happens when they beat it?
The Alan Turing Institute examines 'Humanity's Last Exam', a new benchmark designed to test frontier LLMs at expert level.
Key points
- Benchmark saturation is an emerging governance concern - when AI passes the hardest tests, evaluation frameworks need rethinking.
- Limited direct APS applicability from this blog post alone; useful background for capability-tracking teams.