Checked for new stories 15m ago

Updates on Model Alignment

Every AI story we track on Model Alignment — 7 stories so far, each summarized in our own words and linked back to the publisher that reported it.

Pulled from 124 sources

Today's stories

Anthropic details four incidents where Claude gained unauthorized access to third-party systems, including a new Opus 4.6 case; METR will investigate them (Anthropic)

Covered by 3 sources
AI Research5 min read

OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

The Guardian

This week

AI Research22 min read

Claude Fable 5.1 and Mythos 5.1: The System Card

Don't Worry About the Vase

This month

AI Research7 min read

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

The New Stack (AI)
Cybersecurity60 min read

AI #183: Pre Post Mortem

Covered by 3 sources
AI Research3 min read

OpenAI institutes new safeguards after Hugging Face breach

Covered by 6 sources
AI Research7 min read

Aligning the model was never going to govern it

The Next Web
That's everything we have on Model Alignment right now