# Arena.ai evaluates Jev Router across more than 4,700 agentic sessions on Agent Arena

Arena.ai evaluated typesafeai's Jev Router across more than 4,700 agentic sessions on Agent Arena, finding that while it does not improve on the current Pareto frontier, it demonstrates strong steerability in response to user feedback.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/d2efd8a2a7364c27949b00b4.html

Content language: en
Localization state: sameLanguage

Source: Arena.ai · Published 10/7/2026, 15:33:08

Arena.ai evaluated Jev Router by typesafeai on Agent Arena across more than 4,700 real-world agentic sessions. The tests showed that Jev Router does not improve on the existing Pareto frontier: matching the performance of DeepSeek V4.1 Flash (Max) cost 38% more, with 1.7 times higher median model request latency.

The evaluation noted that the router mostly directs queries to Pareto-efficient models, selecting DeepSeek V4.1 Flash most frequently, alongside regular selections of GPT-6.1 Sol and GPT-6 Luna. Jev Router's primary strength was steerability, scoring +10% and nearly matching Claude Opus 5.5 (High) at +10.48 by routing to more capable models following user feedback.

Tags: Arena.ai, Jev Router, typesafeai, Benchmarks

[View original post](https://x.com/arena/status/2107961555363213482)
