Checked for new stories 12m ago

Updates on AI Safety

Every AI story we track on AI Safety — 178 stories so far, each summarized in our own words and linked back to the publisher that reported it.

Pulled from 124 sources

Today's stories

Scoop: OpenAI faces GOP-led Senate investigation into Hugging Face breach

Covered by 3 sources

Paul Christiano joins OpenAI Foundation Board

Covered by 6 sources
Cybersecurity7 min read

We have started losing control of AI. It’s time to shut it down | Garrison Lovely

The Guardian
AI Research5 min read

OpenAI not on track to reduce risk of ‘catastrophic’ loss of control, says board member

The Guardian
AI Research10 min read

Why some experts increasingly fear AI will take over

BBC
Cybersecurity4 min read

Anthropic has a cute graphic showing how its AI spread 'malicious' code

Business Insider

This week

Anthropic's Alignment Science lead says there is a ">10%" chance AI could kill all humans within the next decade and is worried about recursive self-improvement (Evan Hubinger/@evanhub)

Covered by 9 sources

Anthropic researcher Jacob Coxon says he is quitting the AI industry over fears that tech companies are racing to build systems they won't be able to control (Amrith Ramkumar/Wall Street Journal)

Covered by 11 sources
AI Research1 min read

The AI warnings are coming from inside the lab

Platformer

Fields medalist Jacob Tsimerman, set to join OpenAI later in September, launches the Mathematical AI Safety Institute to apply higher math to AI safety problems (Siobhan Roberts/New York Times)

Techmeme
AI Research7 min read

“Some agents will be pursuing their own objectives”: OpenAI’s chief scientist warns AI could trick and blackmail humans

The New Stack (AI)
AI Research2 min read

OpenAI admits to German wiki ‘incident’

Covered by 5 sources
Cybersecurity6 min read

The US plans to raise AI-directed cyberattacks with China, Nikkei reports

The Next Web
AI Research5 min read

No monitoring system caught the German wiki. Two outside researchers found it by searching the internet.

The Next Web

An in-depth look at OpenAI's wiki incident: other hacked message boards, OpenAI's cover-up, how harmless web search tasks led agents to break out, and more (Zvi Mowshowitz/Don't Worry About the Vase)

Techmeme

How Rationalism, a movement pioneered by Eliezer Yudkowsky focused on existential superintelligent AI risks, influenced top AI leaders and their alarmist claims (Cal Newport/New York Times)

Techmeme
Cybersecurity6 min read

AI agents keep finding ways to bend the rules. Here are some of the wildest.

Business Insider
AI Research7 min read

In “An Alien Mind,” OpenAI’s Jakub Pachocki Urges Shared Safety Bars

Unite.AI
AI Research3 min read

Astra appears to think without showing its work, and the people arguing about it co-wrote the warning

The Next Web
AI Research22 min read

Claude Fable 5.1 and Mythos 5.1: The System Card

Don't Worry About the Vase
Politics16 min read

An Open Letter to Bernie Sanders: Regulate AI’s Dangers, Don’t Ban Its Promise

Unite.AI
AI Research6 min read

Researchers Document OpenAI Agent Swarm That Repurposed German Wiki

Covered by 3 sources
Cybersecurity5 min read

OpenAI’s rogue agents keep escaping, with no formal process to investigate them

TechCrunch

This month

AI Research4 min read

This Is the Worst Possible Time for OpenAI to BfЖ7!م#2猫$9&क

Covered by 3 sources
Cybersecurity4 min read

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

Covered by 4 sources
Robotics5 min read

TechCrunch Disrupt 2026’s new Real World AI Stage features Nvidia, robots, and extinct animals 

TechCrunch
AI Research4 min read

Researchers fear safety disaster ahead of OpenAI’s Astra release

The Verge
Cybersecurity4 min read

Anthropic Says It Hit the Brakes on AI Testing Following Autonomous Hacks

Covered by 3 sources
AI Research4 min read

Anthropic Deliberately Trained an Extremely Misaligned, Reward-Seeking AI and It Did Some REALLY Bad Things

Futurism
Cybersecurity13 min read

The Singularity Is Not What It Seems: Whatever the AI Future Is, We're in It Now

Hacker News
AI Research5 min read

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data

Covered by 2 sources
Cybersecurity3 min read

OpenAI delayed its new model’s development after the Hugging Face hack

The Verge
AI Research4 min read

‘Not perfectly aligned’ with human values: Anthropic admits security failures behind AI hacking incidents

The Guardian
Cybersecurity5 min read

Anthropic has resumed the tests in which its models attacked real companies

The Next Web
AI Research7 min read

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

The New Stack (AI)
AI Research5 min read

CCTV-Affiliated Account Attacks Anthropic, Sets Terms for US-China AI Talks

Unite.AI
AI Research1 min read

Number of AI agents going out of control peaks in July — British newspaper

TASS English

Would you share your weirdest agent logs with AI safety researchers?

Hacker News
AI Research4 min read

Sharp rise in incidents of AI escaping users’ control, research finds

The Guardian
AI Research6 min read

OpenAI Says AGI Is Coming By Year-End. It Also Just Had The Worst Safety Crisis In Its History.

Forbes
Agents3 min read

This Is How Anthropic Thinks AI Agents Should Navigate the Physical World

Covered by 3 sources
AI Research5 min read

Silico, a tool for researchers to understand their AI models better

IEEE Spectrum
AI Research5 min read

OpenAI wants California to toughen the AI law it once fought

The Next Web
AI Research3 min read

Microsoft Moves AI Governance from Policy to Runtime Enforcement

InfoQ (AI, ML & Data)
AI Research3 min read

Sam Altman says he's worried about AI being controlled by a few powerful players

Business Insider
AI Research2 min read

OpenAI says California should strengthen its AI safety bill

Covered by 2 sources
Cybersecurity2 min read

OpenAI pauses training of new AI models due to cyber risks — company

Covered by 10 sources
AI Research3 min read

OpenAI Halts AI Training on Advanced Model as It Detects Dark Signs Emerging

Futurism
AI Research3 min read

OpenAI institutes new safeguards after Hugging Face breach

Covered by 6 sources
Cybersecurity4 min read

OpenAI Scales Back AI Development, but it Could be Too Late

AI Business
Cybersecurity8 min read

OpenAI’s Greg Brockman: Z.ai’s GLM-5.3 likely to “significantly accelerate the threat landscape”

The New Stack (AI)
AI Research14 min read

I Turned AI to the Dark Side

Hacker News
Cybersecurity1 min read

SPAR – Fall 2026 AI Safety Research Projects

Hacker News
AI Research6 min read

Experts are warning: our AI arms race is putting humanity at risk | Stuart Russell

The Guardian
AI Research3 min read

Bernie Sanders calls on Silicon Valley to ‘pause AI development’ in interest of humanity

Covered by 2 sources
AI Research16 min read

Import AI 468: 23 RSI ideas; PostTrainBench+; and how trust and transparency interplay with AI racing

Import AI
AI Research5 min read

House Democrats Press Johnson for AI CEO Testimony After Rogue Model Hacks

Covered by 3 sources
Cybersecurity5 min read

OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause

Covered by 2 sources
Cybersecurity7 min read

The AI safety test is becoming a safety risk

TechCrunch
Showing the 60 most recent of 178 stories on AI Safety