UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
A joint UK-US pre-release cyber evaluation of a PRC frontier model sets a benchmark for AI safety testing that Australian agencies and AISI should track.
Key points
- UK AISI and US CAISI jointly assessed Kimi K3's cyber capabilities, finding it below leading US models but ahead of prior open-weight models.
- Kimi K3 autonomously completed a simulated corporate network attack in 1 of 10 attempts, signalling growing open-weight cyber risk.
- Australia's AISI is absent from this joint evaluation - a notable gap as peer safety institutes deepen bilateral testing collaboration.
Implications for Australian agencies
- Monitor Australia's AISI and cyber security policy teams may want to monitor the UK-US joint evaluation program as open-weight model cyber capabilities continue to advance.
- Consider Agencies with cyber risk responsibilities could consider whether open-weight model capability benchmarks like ExploitBench and TLO could inform threat assessments for AI-enabled cyber attacks on government systems.
Implications are AI-generated. Starting points, not advice — see methodology for how they're framed.
View original source
Copied.
"UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities"
Source: NIST – AI News (topic 2753736)
Published: 23 July 2026
URL: https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities
The UK AI Security Institute and US CAISI have published a preliminary joint evaluation of Moonshot AI's Kimi K3, a Chinese model released on 16 July 2026 with an open-weight release planned shortly after. The evaluation focused on cyber capabilities using ExploitBench and the 'Last Ones' cyber range. Kimi K3 outperforms previous open-weight models on exploit development but falls well short of leading US closed-weight models, notably failing to achieve arbitrary code execution. It completed a 32-step simulated corporate network attack in 1 of 10 attempts, indicating a baseline capacity for autonomous attack against weakly defended systems. The evaluation highlights a continuing trend of open-weight models approaching capabilities previously exclusive to frontier closed-weight systems.
Implications for Australian agencies:
- [Monitor] Australia's AISI and cyber security policy teams may want to monitor the UK-US joint evaluation program as open-weight model cyber capabilities continue to advance.
- [Consider] Agencies with cyber risk responsibilities could consider whether open-weight model capability benchmarks like ExploitBench and TLO could inform threat assessments for AI-enabled cyber attacks on government systems.
Retrieved from SIMS, 16 September 2026.