AIarXiv

Heuristic editor, no API keyVerdict: Routine

WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents

As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects…

By Jiang, Yang, Ji +5

Score██████░░░░5.8

Key numbers

  • 6.6% to 42.3%

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

As interactive 3D worlds are increasingly used to study intelligent behavior, it becomes important to develop efficient pipelines for identifying anomalies in these simulated environments, such as floating objects, traversable walls, or objects inconsistent with the surrounding scene. Multimodal AI systems, including vision-language models (VLMs) and vision-language-action models (VLAs), have shown potential for automating this task. However, 3D world auditing is complex, requiring the close coupling of two distinct capabilities: action, to navigate the 3D world and search for anomalies systematically and efficiently; and visual reasoning, to understand the environment and identify anomalies from multimodal observations. It remains largely unexplored whether multimodal agents can effectively couple these two capabilities, using visual reasoning to identify potential anomalies while taking actions to validate them. In this paper, we introduce WorldAuditBench, a benchmark for 3D world auditing comprising 213 anomaly tasks across 13 environments built with Unreal Engine 5 and Three.js, spanning five anomaly families. We evaluate five frontier models under a fixed exploration budget using two auditing paradigms: VLA-based exploration followed by VLM-based anomaly identification, and an end-to-end VLM agent in which visual reasoning directly guides action selection. Across the evaluated models and two paradigms, success rates range from 6.6% to 42.3%, substantially below human performance (83.4%). Through the task of world auditing, WorldAuditBench provides a testbed for studying how multimodal agents couple action and visual reasoning in interactive 3D environments, while highlighting current limitations in their ability to gather and interpret evidence during exploration.

Ziyan Jiang, Jingbo Yang, Jiabao Ji, Yujian Liu, Qiucheng Wu, Tommi Jaakkola, Yang Zhang, Shiyu Chang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); verification (multiple benchmarks); scale (efficient); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.8

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.1 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.5 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
84%
Citations, upvotes, points, mentions.
Freshness
79%
Half-life decay since publication.
  • Hugging Face upvotes66 (reference 25, via hf-daily, Oct 1, 2026, 13:49 UTC)
  • GitHub stars2 (reference 250, via hf-daily, Oct 1, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 02:16 UTC. Paper type: method.
  • Categories: cs.AI
  • TOP, No.5 in the Artificial Intelligence edition of October 1, 2026.