AIHugging Face

Heuristic editor, no API keyVerdict: Routine

ActiveSaddler: Automated Curriculum Learning for Agent Harness Optimization

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback.

By Park, Kim, Zhang +8arXiv

Score█████░░░░░5.0

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

Automated harness optimization can substantially improve LLM agents by iteratively updating their prompts, tool interfaces, and control logic from execution feedback. However, existing methods primarily optimize how the harness is updated while largely fixing which training scenarios generate the feedback that drives those updates. As the harness evolves, the scenarios most useful for further optimization can change, suggesting that the training curriculum itself should adapt alongside the harness. We formulate this missing dimension of harness optimization as an automated curriculum learning problem and introduce ActiveSaddler. ActiveSaddler models the evolving curriculum as a non-stationary bandit with dynamically instantiated optimization targets. It abstracts recurring failures into reusable failure-pattern arms, estimates the potential learning progress from further targeting each pattern, and adaptively balances revisiting known weaknesses with exploring unseen scenarios for new ones. Optimization outcomes continually update both the set of discovered failure patterns and their priorities, allowing the curriculum to co-evolve with the harness. Experiments on GAIA2 and Terminal-Bench 2.0 show that ActiveSaddler consistently discovers stronger harnesses, improving test Pass@1 by 4.4 and 7.5 percentage points over the same harness optimizer using a scenario order fixed before optimization, respectively. Ablations further show that these gains depend on dynamically constructing optimization targets, estimating their evolving utility, and balancing continued optimization with new failure discovery. Together, these results establish automated curriculum learning as a new crucial optimization dimension for harness optimization.

Sungho Park, Wonjoong Kim, Jue Zhang, Wook-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Victor Rühle

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage██░░░ 224%Reusable within one subfield (a technique, dataset, or protocol a few groups will adopt).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: novelty (discovery); verification (ablations, held-out test); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.3 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.1 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
80%
Citations, upvotes, points, mentions.
Freshness
29%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 2, 2026, 13:49 UTC)
  • Hugging Face upvotes56 (reference 25, via hf-daily, Oct 4, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 13:49 UTC. Paper type: empirical.
  • BRIEF, No.5 in the Artificial Intelligence edition of October 5, 2026.