# Harness-Aware Distillation trains small model agents to 63.4% success on ALFWorld

A research paper presents Harness-Aware Distillation, a training framework that teaches small language model agents to leverage harness information by contrasting teacher decisions. The resulting student model reached 63.4% on unseen ALFWorld tasks, outperforming its 8B teacher.

Language: en
Time zone: America/Los_Angeles

HTML: https://didcodexreset.com/news/69ad27f580d24dbfb8438d77.html

Content language: en
Localization state: sameLanguage

Source: elvis · Published 10/6/2026, 09:57:17

Researchers introduced Harness-Aware Distillation, a framework designed for small language model agents operating with an execution harness, according to an [arXiv paper](https://arxiv.org/abs/2610.02858). The method queries the same teacher model with and without harness information, training the student model to prefer actions chosen when harness data is present. A filter discards candidate pairs where the preferred action contradicts harness records, operating without task rewards or success labels.

The researchers observed that simply adding harness data to on-policy distillation raised the student's harness usage on ALFWorld from 65.7% to 73.1%, but left task success flat at 43.1% to 43.5%. With Harness-Aware Distillation, the student model achieved a 63.4% success rate on unseen ALFWorld tasks, outperforming the best baseline at 47.0% and surpassing its 8B teacher. The trained student also escaped 59.7% of stalls, compared to baselines remaining near the untrained student baseline of 46.8%.

Tags: Model Distillation, AI Agents, ALFWorld, LLM Research

[View original post](https://x.com/omarsar0/status/2107507684270600394)
