AI Research2 min reading time
Now we have a timeline of the OpenAI accidental attack against Hugging Face
Covered by 2 sources
Read full postOpenAI began training an experimental, unreleased model on May 7, using reinforcement learning with verifiable rewards (RLVR) for cybersecurity tasks. This training approach may explain the lack of safety measures and monitoring that led to an accidental attack on Hugging Face.

Covered by 2 sources
- Further Developments About Internal AI Models Hacking Things· Don't Worry About the Vase
- OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause· 2 sources
- OpenAI Astra pause 🚨, Claude Code cross-session 🤖, how Cursor Router works 🔀· 2 sources
- The AI safety test is becoming a safety risk· TechCrunch
- Why Aren’t Any AI Companies Watching Their Frontier Models to Make Sure They Don’t Go on Hacking Sprees?· Futurism
- Why are so many AI models going 'rogue'? The experts weigh in · TechRadar

