Did Codex Reset
GitHub

Artificial Analysis averages Cyber Index across three benchmarks and scores safety refusals as zero

Artificial Analysis

Artificial Analysis says the Cyber Index averages scores across CWE-Bench-AA, DeepsecBench-AA, and CyberGym-E2E-AA. A task that a model declines on safety grounds scores zero. Those safety blocks are reported separately from failures, so refusals can be distinguished from cases where the model attempts a task and does not succeed.

The scoring rules are in the Cyber Index methodology. Full results are on the Artificial Analysis Cyber Index evaluation page.