• 4 mins read
  • Published

China's Kimi K3 Lags Behind US AI Models in Cybersecurity Tests

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

China's Kimi K3 Lags Behind US AI Models in Cybersecurity Tests Science.Report
China's Kimi K3 Lags Behind US AI Models in Cybersecurity Tests

A UK-US evaluation of Moonshot AI's Kimi K3 large language model found it significantly underperformed leading US models on offensive cybersecurity benchmarks, raising questions about the current state of China's AI capabilities in this domain

A joint evaluation by the UK Artificial Intelligence Security Institute (UK AISI) and the US Center for AI Standards and Innovation (CAISI) has found that Moonshot AI's Kimi K3 large language model performs well below the leading US models in offensive cybersecurity tasks. The assessment, which comes amid heightened international scrutiny of AI capabilities, focused on the ability of large language models to autonomously develop software exploits and conduct simulated cyberattacks.

The evaluation used ExploitBench, a public benchmark from Carnegie Mellon University, to measure how effectively AI systems can generate end-to-end exploits for known software vulnerabilities. Kimi K3 achieved an overall cyber capability score of 32.2%, while the most cyber-capable US models averaged 76.2% on the same tasks. For context, GLM-5.2, another Chinese open-weight model, scored 24%, placing Kimi K3 ahead of some domestic competitors but still far behind the US frontier.

Exploit Development and Attack Simulation

Researchers tested whether Kimi K3 could achieve arbitrary code execution (ACE), a critical exploit outcome that allows attackers to take control of a target system. Kimi K3 failed to achieve ACE on any of the 41 ExploitBench tasks, whereas the top US models succeeded on an average of 20 out of 41 samples. This gap highlights a substantial difference in the models' ability to autonomously generate high-severity exploits.

The evaluation also included "The Last Ones" (TLO), a 32-step simulated corporate network attack designed to test end-to-end cyber operations. Kimi K3 reached step 17 on average, compared to 28.5 steps for the most cyber-capable US models. In only one out of ten attempts did Kimi K3 complete the full simulated attack. The report noted that, despite its lower performance, Kimi K3 was able to autonomously attack small, weakly defended enterprise systems when provided with initial network access and explicit instructions. Notably, the model's built-in safeguards did not prevent it from attempting offensive cyber operations during testing.

Benchmark Limitations and Policy Implications

The findings are based on a single benchmark and a limited set of tasks, with Kimi K3's overall score estimated from ExploitBench alone. In contrast, some US models were evaluated across a broader range of cybersecurity tasks, which may affect direct comparisons. The results are preliminary and do not capture the full range of possible model behaviors or real-world deployment risks.

These results arrive at a time of growing concern in Washington and other capitals about China's progress in artificial intelligence. However, the evaluation suggests that, at least in the domain of offensive cybersecurity, leading US models retain a significant technical advantage. This assessment echoes recent analysis of China's AI ambitions in other high-stakes sectors, such as nuclear energy, where a recent report on China's AI roadmap for nuclear safety highlighted both advances and ongoing limitations.

Context and Remaining Questions

Kimi K3 has attracted attention for its performance on general AI benchmarks, but this evaluation provides a more nuanced view of its capabilities in specialized, high-risk domains. The model's ability to attempt offensive cyber operations without effective internal safeguards raises questions about the adequacy of current safety measures in large language models, especially as they are adapted for sensitive applications. The report also notes that fears of Chinese AI dominance in cybersecurity may be overstated, at least for now, and cautions against overregulation based on incomplete evidence.

Understanding the limitations of current benchmarks is essential for interpreting these results. ExploitBench is designed to test a model's ability to generate software exploits, but it does not fully represent the complexity of real-world cyber operations, which often involve unpredictable environments, human oversight, and evolving defensive measures. Benchmark scores provide a useful snapshot of technical capability, but they should not be mistaken for comprehensive assessments of operational risk or safety. As AI systems are increasingly evaluated for dual-use and security-sensitive applications, transparent, multi-dimensional testing will be critical for responsible deployment and governance.

Related articles