AIHugging Face

Heuristic editor, no API keyVerdict: Routine

AgentGarten: Code Worlds for Evolving Agents

Interactive virtual worlds allow agents to learn through exploration and interaction.

By Chi, Miao, Shi +11arXiv

Score██████░░░░5.6

Caveats

  • Preprint; not yet peer reviewed.

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.

Jiawei Chi, Shangchen Miao, Zhiyuan Shi, Kailu Wu, Hanyang Wang, Weiliang Chen, Qiyu Dai, Jinshan Ren, Jun Gao, Mingsheng Long, Yueqi Duan, Jiangran Lyu, Jialong Wu, Fangfu Liu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes███░░ 310%Meaningful benefit to many people within a few years.

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); scale (scalable, efficient); stakes (global scale, general AI). Red flags: derivative (we apply).

How the score was computed

rank-2026-10-07

Score██████░░░░5.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.1 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.1 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
92%
Citations, upvotes, points, mentions.
Freshness
65%
Half-life decay since publication.
  • Hugging Face upvotes128 (reference 25, via hf-daily, Oct 9, 2026, 13:49 UTC)
  • GitHub stars97 (reference 250, via hf-daily, Oct 9, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 9, 2026, 03:46 UTC. Paper type: method.
  • Categories: cs.CV
  • TOP, No.1 in the Artificial Intelligence edition of October 9, 2026.