UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
- AI
- Security
- Open Source
- Developer Tools
The linked post points to a UK AISI and CAISI preliminary assessment of Kimi K3’s cyber capability, republished by NIST. The core finding is simple: Kimi K3 is not at the level of the strongest U.S. cyber-capable models on these tests, but it is good enough to autonomously make meaningful progress through multi-step attack paths and, at least occasionally, compromise a small vulnerable enterprise environment. That puts it well past toy status even if it still misses the frontier. Several people argued the report itself is narrower than the headline implies. The evaluation used a selective set of cyber tests because of Kimi K3’s hosting setup, and its overall estimate leaned heavily on ExploitBench rather than a broad battery. That made readers question how much confidence to put in a single aggregate bar chart, especially when the report groups “top U.S. models” in a way that hides exactly which systems and safeguard settings are being compared. The main substantive dispute was about whether Kimi K3’s score is being understated by the harness. One camp said UK AISI likely under-elicits quirky open or open-weight models, especially token-hungry ones, because the 100 million token budget includes cache hits and may cut off progress before those models plateau. The pushback was that budget limits are part of the point in agentic cyber evals, and every model gets better with more tokens, so relative rankings still carry signal. Even with that disagreement, the broad read was that the strategic gap looks measured in months, not years. The stronger takeaway was practical rather than benchmark-centric. Closed U.S. models may still have much broader and more reliable offensive cyber ability, possibly because some were explicitly tuned for it, but access and guardrails are now central to real-world usefulness. A model that is somewhat weaker yet consistently willing to help can be more operationally valuable than a stronger API model that refuses unpredictably. That is why the conversation kept circling back to open-weight and Chinese models as tools defenders can actually use, and to the expectation that once these capabilities exist in available models they will persist rather than vanish with a provider policy change.
Treat this as a timing signal, not a comfort signal. Even if Kimi K3 is meaningfully behind the best closed models today, multiple commenters think open and less-restricted models are close enough that security teams should start using them now for internal testing and plan for stronger open-weight cyber models within months.
-
nist.gov
- Discuss on HN