AIarXiv

Heuristic editor, no API keyVerdict: Routine

InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns…

By Lin, Xu, Bian +2

Score█████░░░░░4.9

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.

Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks, zero/few-shot); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░4.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.4 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.1 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
74%
Citations, upvotes, points, mentions.
Freshness
35%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 4, 2026, 02:05 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 4, 2026, 02:05 UTC)
  • Hugging Face upvotes43 (reference 25, via hf-daily, Oct 4, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.RO, cs.CV, cs.GR
  • BRIEF, No.8 in the Artificial Intelligence edition of October 5, 2026.