# NVIDIA paper details VERA framework for co-evolving agent harnesses and models

An NVIDIA research paper presents VERA, a system that converts benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring while co-evolving agent harnesses alongside model weights. The team open-sourced the environment corpus, reporting that a 27B agent scored 71.6 on AutoCoWorkBench to top Claude Opus 4.8.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/10433a5ea5fa405a9d584306.html

Content language: en
Localization state: sameLanguage

Source: elvis · Published 10/7/2026, 08:30:48

VERA converts benchmark trajectories into more than 9,000 restartable sandboxes with rubric scoring, retaining only environments that execute and can be evaluated from observable evidence. During training, the framework co-evolves the agent harness and model weights simultaneously. A harness modification is kept only if it passes self-tests and yields at least a 5-point gain on the development set, while model checkpoints are rejected if performance drops by more than 20%.

According to the [research paper](https://arxiv.org/abs/2610.05923), a 9B co-evolved agent outperformed the strongest single-axis baseline by 10.3 points on AutoCoWorkBench and 13.0 points on AutoMedBench. At 27B, the agent reached a score of 71.6 on AutoCoWorkBench, surpassing Claude Opus 4.8. NVIDIA has open-sourced the environment corpus.

Tags: NVIDIA, AI Agents, Benchmarks, Open Source

[View original post](https://x.com/omarsar0/status/2107849444897243141)
