Item Catalogue

AI governance, regulation, strategy, and practice developments from monitored sources.

Last updated 27 Aug 2026, 06:05 AM AEST
Clear
Saved (0)
Filters 1 active
Jurisdiction
Category
Source (1)

Date range

primary source commentary 26 items · Page 2 of 2

Week of 3 February 2025

Alan Turing Institute – Blog(Global) 6 Feb 2025 42

LLMs have been set their toughest test yet. What happens when they beat it?

The Alan Turing Institute examines 'Humanity's Last Exam', a new benchmark designed to test frontier LLMs at expert level.

Key points
  • Benchmark saturation is an emerging governance concern - when AI passes the hardest tests, evaluation frameworks need rethinking.
  • Limited direct APS applicability from this blog post alone; useful background for capability-tracking teams.