AIHugging Face

Heuristic editor, no API keyVerdict: Notable

CompoWorld: Compositional Environment Scaling for General Agents

Automatically generated environments provide a scalable source of interaction data for training general agents.

By Yang, Xu, Da +9arXiv

Score█████░░░░░5.1

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (CompoWorld), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.

Xiao-Wen Yang, Weiyi Xu, Wen Da, Hang Xu, Canwei Li, Hong-Jie You, Pusen Dong, Yucheng Zeng, Zhaokai Luo, Yu-Feng Li, Yao Hu, Mu Chuan

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (relative gain, outperforms); verification (multiple benchmarks); scale (scalable); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.1

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
6.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.9 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
64%
Citations, upvotes, points, mentions.
Freshness
32%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Hugging Face upvotes31 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars0 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • BRIEF, No.4 in the Artificial Intelligence edition of September 29, 2026.