AIarXiv

Heuristic editor, no API keyVerdict: Notable

RealtimeWAM: One-Step Asynchronous World Action Models

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction.

By Lv, Du, Feng +7

Score█████░░░░░5.5

Key numbers

  • 1 % drop

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

World Action Models (WAMs) incorporate visual representations from video generation backbones to guide action prediction. Recent efficient WAMs adopt Mixture-of-Transformers (MoT) architectures and compute video representations once for reuse by the action expert. However, intra-expert iteration (\ie, multi-step action denoising) and inter-expert waiting (\ie, sequential execution of the video and action experts) still limit inference efficiency. To this end, we present RealtimeWAM, an extremely efficient WAM variant with one-step action generation and asynchronous inference, addressing these two bottlenecks. To reduce intra-expert iteration, we propose Teacher-Anchored Consistency Distillation (TACD) to address a local-global error gap: low local consistency error alone does not guarantee accurate final actions. TACD supplements local consistency with explicit supervision from the frozen teacher's multi-step rollout endpoint, enabling accurate one-step action generation. Additionally, we propose Cross-Expert Wavefront Pipelining (CEWP) to eliminate unnecessary expert-level waiting. It overlaps the two experts through block-wise sharing of the video KV cache, synchronizing only immediately before the corresponding action attention consumes it. Extensive experiments across diverse benchmarks (\eg, LIBERO, LIBERO-Plus and RoboTwin) and model variants (\eg, Fast-WAM and Faster-WAM) demonstrate the superiority of RealtimeWAM. Notably, RealtimeWAM maintains near-lossless performance (\ie, <1% drop) across these benchmarks while delivering significant end-to-end speedup (\eg, ~25× on H100). Our code and checkpoints are available via this \href{https://github.com/ModelTC/LightX2V/tree/main/examples/realtimewam}{link}.

Chengtao Lv, Jinyang Du, Shuyi Feng, Yang Yong, Shiqiao Gu, Shunzi Yang, Ruihao Gong, Shen Ren, Tianwei Zhang, Wenya Wang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); verification (code released); scale (efficient); stakes (global scale).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.5

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
69%
Citations, upvotes, points, mentions.
Freshness
78%
Half-life decay since publication.
  • Hugging Face upvotes7 (reference 25, via hf-daily, Oct 6, 2026, 13:49 UTC)
  • GitHub stars2.9k (reference 250, via hf-daily, Oct 6, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 6, 2026, 13:49 UTC. Paper type: method.
  • Categories: cs.CV, cs.LG, cs.RO
  • TOP, No.3 in the Artificial Intelligence edition of October 6, 2026.
RealtimeWAM: One-Step Asynchronous World Action Models | Humanity's List · Humanity's List