Did Codex Reset
GitHub

Artificial Analysis launches Cyber Index for enterprise AI defense

Artificial Analysis

Artificial Analysis launched the Artificial Analysis Cyber Index and the Cyber Index Alliance to evaluate how AI models perform on enterprise cyber defense tasks. Launch partners are CollinearAI, IBM, NVIDIA, and Vercel. The index combines three partner-contributed, open benchmarks that measure how well agents find and fix vulnerabilities.

CWE-Bench-AA, from CollinearAI, covers auditing and patching with 120 held-out tasks across all ten OWASP Top 10 (2025) categories and C/C++, Go, Java, JavaScript/TypeScript, Python, and Rust. DeepsecBench-AA, from Vercel, isolates discovery: given a codebase and a budget, an agent must find every vulnerability and is scored against findings from human security reviewers, with real findings rewarded and benign code flagged as vulnerable penalized. CyberGym-E2E-AA, from BerkeleyRDI, runs end to end: find the memory-safety bug, write a proof of concept that triggers the crash, then patch it so the crash no longer reproduces.

Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index at 56, followed by GPT-6 Luna (max) at 53, GLM-5.3-Flash at 50, and Muse Spark 1.3 (xhigh) at 44. Artificial Analysis said safety refusals hold back GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback), and Gemini 3.8 Flash (high). Those models decline tasks representing 32–38% of the index and trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98%, and Claude Fable 5.1 refuses 99%.