PrimeScientist gets 10.3% more reward with 50.6% fewer attempts than AutoResearch
PrimeScientist treats deciding where a research agent spends its budget as part of the agent's job. It keeps an executable plan tree of competing research directions and their outcomes. An adaptive MCTS policy reads experimental feedback and the remaining budget, then chooses whether to explore a new direction or continue a promising one.
DAIR.AI reports that, against AutoResearch on 12 AI research tasks and under the same budget, PrimeScientist achieved 10.3% more reward with 50.6% fewer research attempts. DAIR.AI also says the gains hold on systems, code optimization, and ML engineering tasks, without separate figures for those settings. The account points to the PrimeScientist paper.