We burned 11.7B tokens to find the best cyber AI model
Hacker News
Read full postA benchmark tested 10 AI models on 32 recent software vulnerabilities, running each model three times to assess detection performance and consistency. DeepSeek V4 Pro 0813 found the most vulnerabilities, with open-source models now rivaling closed ones in recall but producing more false positives. Multiple runs improve recall by compensating for model inconsistency.




