AIHugging Face

Heuristic editor, no API keyVerdict: Notable

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time.

By Qian, Li, Luo +11arXiv

Score██████░░░░5.7

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and long-horizon execution, which keeps the iteration effective over many rounds. We hypothesize that both capabilities are domain-agnostic, and can therefore be learned in scenarios that are well suited for supervision. Accordingly, we synthesize long-horizon improvement trajectories from machine learning and algorithmic programming tasks, two domains that offer verifiable feedback and reward sustained iteration. Trained on this data, our agent, built on Qwen3.8-27B, achieves strong results on MLE-bench Lite (81.8) and Frontier-CS (70.7), transfers to deep research with 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA, and keeps improving as its budget of rounds grows. These results show that long-horizon reflective data is an effective route toward self-improving agents.

Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (general-purpose, many tasks); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.7

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
88%
Citations, upvotes, points, mentions.
Freshness
49%
Half-life decay since publication.
  • Hugging Face upvotes88 (reference 25, via hf-daily, Oct 1, 2026, 13:49 UTC)
  • GitHub stars10 (reference 250, via hf-daily, Oct 1, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 13:49 UTC. Paper type: method.
  • BRIEF, No.4 in the Artificial Intelligence edition of October 1, 2026.