New Claude Opus 5 vs ChatGPT 5.6 Sol: Benchmarks, Pricing and Token Cost Compared
TL;DR
Anthropic frames Claude Opus 5 as a new reference point for performance relative to cost, citing 43 percent on Frontier Bench and 30 percent on ARC AGI 3 alongside steadier behaviour on longer task chains. Matthew Berman puts the model head to head with ChatGPT 5.6 Sol, looking past the benchmark scores at pricing and cost per token. The more telling differences show up in sustained use and on multi-step work rather than in single-shot prompts.
Nauti's Take
The promising part is the price angle: if Opus 5 really leads on multi-step work while costing less per token, the math behind agent setups changes noticeably. The catch is that Frontier Bench and ARC AGI 3 measure lab conditions rather than real workloads, and this particular comparison comes from a single video.
Nauti would pit both models against one recurring task from actual practice. Heavy-volume teams stand to save real money; occasional users will hardly notice a difference.