AIarXiv

Heuristic editor, no API keyVerdict: Routine

ROWBench: Do Video Models Render What the Program Specifies?

Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines.

By Huang, Lin, Tsai +6

Score█████░░░░░4.9

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. However, their visual adherence to explicit rules and interactions remains insufficiently evaluated. Existing benchmarks assess visual quality, controllability, and instruction or physical adherence, but rarely test fidelity to fine-grained, program-specified world events. We introduce PROWBench, comprising 170 programmatically constructed episodes and 600 proxy videos covering diverse scenes and interactions. PROWBench logs entity states and timestamped events, including those outside the camera's field of view, as replayable world records, from which it renders synchronized views and proxy representations. This enables generated videos to be checked against the observable consequences of program execution. An extensible framework constructs scenes, controls behaviors, and can render each camera view in different representations, such as coarse 3D, and bounding boxes. The benchmark covers first- and third-person perspectives, with synchronized multi-view observations available for a subset of episodes. Grounded in these records, PROWBench evaluates entity control, long-horizon memory, and, with two VLM-based metrics, Logic-Render Alignment and Interaction Success Rate, adherence to the prescribed timeline and the visual realization of timestamped engine-recorded events.

Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (programmable).

How the score was computed

rank-2026-09-29

Score█████░░░░░4.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
3.7 / 10
Weighted rubric, evidence-gated.
Adjusted merit
3.9 / 10
Shrunk toward the desk prior by editor confidence (34%).
Attention
82%
Citations, upvotes, points, mentions.
Freshness
35%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 4, 2026, 02:05 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 4, 2026, 02:05 UTC)
  • Hugging Face upvotes58 (reference 25, via hf-daily, Oct 4, 2026, 13:49 UTC)
  • GitHub stars30 (reference 250, via hf-daily, Oct 4, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.CV
  • BRIEF, No.7 in the Artificial Intelligence edition of October 5, 2026.