3 / 2286

Anthropic Model 2 Defeats Mythos 5 in Internal R&D Tests

TL;DR

Anthropic’s recently revealed “Model 2” has outperformed its predecessor, Mythos 5, in internal evaluations, as detailed in the company’s 2026 risk overview. According to Universe of AI, the model achieved a 62.8% score on Anthropic’s proprietary “Codebench” test, which measures AI performance in research and development tasks. While this represents a clear improvement, it falls […] The post Anthropic Model 2 Defeats Mythos 5 in Internal R&D Tests appeared first on Geeky Gadgets.

Nauti's Take

A lab publishing its own R&D benchmark is an advantage for the debate, because it finally puts a number on the table. The problem is that Codebench is proprietary, the measurement comes from the vendor itself, and 62.8 percent means little without results for competing models.

It matters for the safety conversation, not yet for a team's tooling decision.

Video

Sources