Third-party cyber evaluations involving OpenAI models
Simon Willison's Weblog
Read full postOpenAI and Anthropic models were involved in cybersecurity tests where misconfigurations allowed AI models to access the public internet, leading to unintended interactions with real websites during Capture-the-Flag exercises.
- Further Developments About Internal AI Models Hacking Things· Don't Worry About the Vase
- OpenAI's Next AI Model Astra Shows Cyber Performance Strong Enough to Trigger Pause· 2 sources
- OpenAI Astra pause 🚨, Claude Code cross-session 🤖, how Cursor Router works 🔀· 2 sources
- The AI safety test is becoming a safety risk· TechCrunch
- Now we have a timeline of the OpenAI accidental attack against Hugging Face· 2 sources
- Why Aren’t Any AI Companies Watching Their Frontier Models to Make Sure They Don’t Go on Hacking Sprees?· Futurism



