Paper uses one Jev yes/no score to flag alignment failures
Researchers built RLCDAlignBench from 44 existing benchmarks covering ten failure types, including sycophancy, jailbreaks, deception, prompt injection, and reward hacking. The method asks Jev one generic yes/no question about a model’s response and uses that probability as the score. Jev, TypeSafe AI’s calibrated decision model, can answer many typed questions about one input in a single call, each with its own probability.
On StrongREJECT, the paper says Jev matches the GPT-4o-mini scorer’s agreement with human labels and ranks responses better, with an AUROC of 0.971 versus 0.929. Where Jev confidently disagreed with benchmark labels, it found label errors in three benchmarks. Question wording mattered little, but thresholds did not transfer across benchmarks; fitting one on 10 labelled items raised median F1 from 0.706 to 0.793.
The results are reported in the arXiv paper, with a paper chat also available.