AIarXiv

Heuristic editor, no API keyVerdict: Notable

Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions.

By Gao, Zhou, Zheng +10

Score██████░░░░5.7

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.

Hongcheng Gao, Jingjing Zhou, Zelin Zheng, Shijia Ge, Jay Zhu, Yazhe Wang, Jianshu Zeng, Xuan Shangguan, Di Wu, Lingyu He, Zhiqi Jia, Sihang Wu, Xiao He

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (outperforms); novelty (discovery); verification (held-out test); scale (efficient); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.7

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.4 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
90%
Citations, upvotes, points, mentions.
Freshness
51%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Hugging Face upvotes100 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars45 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record