Best Taste-Bench model chooses the better direction 59.7% of the time
A paper from Microsoft and colleagues measures how often frontier agents choose the better direction at a decision point in a long task. On Taste-Bench, each question is a point in a trajectory where several directions are open and one leads to a better outcome. The forks are mined automatically from parallel attempts and detours in engineering and research runs. The best model answers 59.7% correctly.
Forks whose deciding evidence appears later in the trajectory are much harder, and a larger reasoning budget does not raise accuracy. The authors distill a teacher’s judgment of the outcome into a student model, which improves end-to-end success on held-out SWE-bench Pro tasks.
The paper is available on arXiv, with a companion Chat with Paper page.