AIarXiv

Heuristic editor, no API keyVerdict: Routine

Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness

Recent game design agents have made substantial progress in generating playable games.

By Chen, Wu, Jia +1

Score█████░░░░░5.2

Key numbers

  • 34.1% improvement over the same-model
  • 93.4% among compared methods

Caveats

  • Preprint; not yet peer reviewed.

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.

Jiajun Chen, Haoyu Wu, Mingda Jia, Xihui Liu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence██░░░ 214%Limited: single setting, weak baselines, or an observational association presented as causal.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (state of the art); scale (efficient); stakes (general AI).

How the score was computed

rank-2026-10-07

Score█████░░░░░5.2

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.0 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
79%
Citations, upvotes, points, mentions.
Freshness
65%
Half-life decay since publication.
  • Hugging Face upvotes51 (reference 25, via hf-daily, Oct 8, 2026, 05:48 UTC)
  • GitHub stars6 (reference 250, via hf-daily, Oct 8, 2026, 03:47 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 8, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.MA, cs.SE
  • BRIEF, No.1 in the Artificial Intelligence edition of October 8, 2026.