AIarXiv

Heuristic editor, no API keyVerdict: Notable

Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment.

By Li, Shi, Li +7

Score█████░░░░░5.4

Key numbers

  • 34.3% more than the strongest
  • 57.0% to 74.2% on Terminal-Bench
  • 1.5% to 9.1% on Terminal-Bench

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.

Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong Huang, Fuxiao Liu, Haitao Mi, Jordan Boyd-Graber, LeoweiLiang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (relative gain, outperforms); verification (multiple benchmarks); scale (scalable).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.4

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.5 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
82%
Citations, upvotes, points, mentions.
Freshness
39%
Half-life decay since publication.
  • Hugging Face upvotes59 (reference 25, via hf-daily, Oct 5, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 5, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.AI
  • TOP, No.5 in the Artificial Intelligence edition of October 5, 2026.
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite | Humanity's List · Humanity's List