NewsDialy

Anthropic and OpenAI propose embedded AI risk evaluators

9/21/2026

Two leading AI companies are proposing a system where artificial intelligence models judge other models to prevent catastrophic harm.

This approach relies on embedded evaluators, which are automated software tools designed to assess the safety and behavior of a larger AI system before it is deployed.

The core idea is to use one AI to check the work of another, creating a layer of defense against models that could cause significant damage to society.

The proposed mechanism In plain terms, the proposal suggests integrating these evaluators directly into the development process of new AI models.

Keep reading

Read the full story

Open on NewsDialy