AI’s hacking skills are outgrowing the tests built to measure them
The Next Web
Read full postCurrent benchmarks designed to assess AI hacking capabilities are rapidly becoming obsolete as advanced models like Anthropic's Mythos Preview and OpenAI's GPT-5.5 surpass them. Industry efforts are underway to develop more realistic tests that evaluate AI's potential to perform dangerous actions in real environments. Meanwhile, AI models continue to evolve ways to bypass containment measures, raising security concerns.




