Will It Mythos?
- AI
- Security
- Open Source
- Developer Tools
- Infrastructure
The post builds a benchmark from nine real security bugs that Anthropic said its internal Mythos system found in open source projects. The author used those bugs as a test set for public models, asking them to audit the relevant file with access to the rest of the repo but without being told what the bug was. The point was not to prove Mythos was fake. It was to check whether the gap between Mythos and public models was dramatic, or mostly hype. The answer landed in the middle. Public models did find some of the bugs, but the best results were still only around four out of nine in a single pass, and some leaderboard placements were distorted by models timing out or burning through a cost cap before finishing all cases. Cheap models from DeepSeek and MiMo looked much better than many expected on bang for buck, Gemini underperformed badly in this setup, and later replication tests mentioned in comments suggest Gemma 4 31B may be the strongest self-hostable option the author has tried when given multiple attempts.
If you rely on LLMs for security review, treat model choice and cost as moving targets rather than assuming the biggest US frontier model wins. For now, the practical play is to benchmark your own workflow, keep an eye on self-hostable models like Gemma, and not confuse bug finding with full autonomous exploitation or secure code generation.
-
swelljoe.com
- Discuss on HN