7 / 2292

Anthropic Model 2 Defeats Mythos 5 in Internal R&D Tests

TL;DR

Anthropic's Model 2 outperformed its predecessor Mythos 5 in internal evaluations, according to the company's 2026 risk overview. Universe of AI reports a score of 62.8 percent on Anthropic's proprietary Codebench test, which measures performance on research and development tasks. That marks a clear improvement over the previous model. The number is hard to place, though, because both the benchmark and the evaluation come from Anthropic itself.

Nauti's Take

A lab publishing its own R&D benchmark is an advantage for the debate, because it finally puts a number on the table. The problem is that Codebench is proprietary, the measurement comes from the vendor itself, and 62.8 percent means little without results for competing models.

It matters for the safety conversation, not yet for a team's tooling decision.

Video

Sources