Paper uses one Jev yes/no score to flag alignment failures
A paper reports that TypeSafe AI’s Jev, with no extra training, separates alignment failures from good responses at a median AUROC of 0.886. Across 19 benchmarks, one Jev pass cost $0.30, compared with $18.96 for the LLM judges those benchmarks use.