Did Codex Reset
GitHub

AI News

TGRSS feeds413 briefs

Best Taste-Bench model chooses the better direction 59.7% of the time

A paper from Microsoft and colleagues reports that the best model answers 59.7% of Taste-Bench questions correctly when choosing the better open direction in long engineering and research tasks. Forks decided by later evidence are much harder, a larger reasoning budget does not raise accuracy, and distilling a teacher’s outcome judgment improves held-out SWE-bench Pro success.