Did Codex Reset
GitHub

DTap red team reports 70.1% direct-misuse ASR on Jev 1.13

TypeSafe AI

Zhaorun Chen says a red-team evaluation of Jev 1.13 on DTap, the DecodingTrust-Agent Platform, found a safety gap: a 70.1% attack success rate under direct misuse and 43.5% under indirect prompt injection. In those indirect-injection evaluations, he says Jev followed attacker-injected instructions, including instructions to exfiltrate user data, delete files, or take other harmful actions.

Chen says a safer integration is to use Jev as a self-gating layer for its own tool calls. He says that approach significantly reduced the attack success rate while preserving most of Jev’s utility. TypeSafe AI questions the idea that Jev is “not working right,” saying Jev is built for composability and that the situation needs more Jev.