AIarXiv

Heuristic editor, no API keyVerdict: Routine

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution.

By Kang, Hachiuma, Zhang +8

Score█████░░░░░5.1

Key numbers

  • 50.00% for the base agent
  • 68.03% with 8 sampled actions

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (relative gain); scale (scalable); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.1

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.0 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
89%
Citations, upvotes, points, mentions.
Freshness
26%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Hugging Face upvotes111 (reference 25, via hf-daily, Oct 3, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 02:16 UTC. Paper type: method.
  • Categories: cs.CL, cs.AI, cs.MA
  • BRIEF, No.2 in the Artificial Intelligence edition of October 5, 2026.