1 / 2405

Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?

TL;DR

Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Groups such as METR and Redwood Research are pushing for access to training checkpoints, evaluation transcripts and employee interviews, plus the right to publish without company edits. Neither company has said which evaluators it will work with, when, or what they will see.

Nauti's Take

Outside evaluators with access to training checkpoints would be real progress for AI safety, since risks could surface before a model ships. The catch is that without fixed timelines, publication rights and protection from restrictive NDAs, the setup stays a voluntary gesture that can end at any time.

Teams building on Claude or ChatGPT should watch for the first published evaluation reports before we treat this as a trust signal.

Tweets

Sources