The idea is to post a safety referee inside each AI lab. Anthropic and OpenAI are both proposing embedded AI evaluators, a term for independent reviewers placed within a company to check whether its models could cause catastrophic harm to society. Both companies have used the word "neutral" to describe these evaluators, and that choice is where the concern starts.

Embedded means the evaluator operates from inside the organization being reviewed. That is a different arrangement from outside regulators, where independence is structural: the reviewer does not depend on the company for access or continued standing. The proposal asks observers to accept that an inside reviewer can still be genuinely neutral. That is the gap neither company has closed.

What "neutral" is being asked to carry

Both Anthropic and OpenAI have framed this as a response to the risk that advanced AI models could cause large-scale, hard-to-reverse harm to society. The structure they have chosen is internal. The label they have attached is neutral. A reviewer embedded inside the company being evaluated is in a structurally different position from one who answers to an outside authority, regardless of how that role is described on paper.

Calling an evaluator neutral does not resolve the structural question. The proposal offers a mechanism and applies a label, and the case for the idea rests on accepting that those two things hold together. Neither company has explained how, in practice, they do.

Related reading