X PAPER / Full edition 中文

AI & Technology

Same weights, 62% in one harness and 33% in another

Hugging Face · @huggingface

Edited by X Paper

Hugging Face’s multi-harness experiment put identical LFM2.5-2.6B weights at 62.1% success in Mini-SWE-Agent and 33.2% in Claude Code. The test comprised 250 held-out data-analysis tasks, rather than a general coding leaderboard.

A harness governs context, tools, retries and stopping. Keeping weights fixed does not keep the whole agent fixed. A capture proxy records exact generated tokens and probabilities; OpenEnv, Harbor and TRL connect environments, tasks and training.

Training across OpenCode, Claude Code, Codex and Mini-SWE-Agent lifted average success from 42.2% to 54.2%. OpenCode-only training averaged 52.3% and reached 58% within OpenCode; mixed training did better in Claude Code and Codex. The authors put the 1.9-point overall gap within noise.

The reported 31.1% reduction in tool calls applies to tasks both baseline and trained models solved, not all tasks or bills. Each run had one seed, and mixed training saw more distinct tasks and processed more tokens. There was no LFM run without the efficiency bonus to isolate its effect.

The practical lesson is to evaluate the model in its actual harness. Open code and models permit further testing; these results do not establish universal superiority at equal compute.

Sources and further reading

Original post and supporting sources read