AIarXiv
Heuristic editor, no API keyVerdict: RoutineMid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution.
Key numbers
- 50.00% for the base agent
- 68.03% with 8 sampled actions
VerdictCompetent work. Briefs at most.
Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ██░░░ 2 | 14% | Limited: single setting, weak baselines, or an observational association presented as causal. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (relative gain); scale (scalable); stakes (general AI).
How the score was computed
- Merit
- 4.0 / 10
- Adjusted merit
- 4.0 / 10
- Attention
- 89%
- Freshness
- 26%
- Citations0 (reference 15, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
- Influential citations0 (reference 3, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
- Hugging Face upvotes111 (reference 25, via hf-daily, Oct 3, 2026, 13:49 UTC)