AIarXiv
Heuristic editor, no API keyVerdict: RoutineRecursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness
Recent game design agents have made substantial progress in generating playable games.
Key numbers
- 34.1% improvement over the same-model
- 93.4% among compared methods
Caveats
- Preprint; not yet peer reviewed.
VerdictCompetent work. Briefs at most.
Abstract
Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's feedback into detailed plans. The Builder turns these plans into candidate games. The coding-native Player creates and executes reusable policies through programmatic interfaces to efficiently collect diverse gameplay trajectories, mitigating evaluation bias caused by slow GUI-based collection. The Reviewer uses carefully designed trajectory-based metrics to induce player preferences, integrating with visual evidence and explicit textual preferences to evaluate games against game-specific criteria. Finally, the Reviewer accepts the better version and provides improvement reviews for the next round, closing the recursive loop. Our method achieves state-of-the-art overall performance of 77.89 on GameCraft-Bench. On GameASG-Bench, it achieves a strict task success rate of 53.2%, a 34.1% improvement over the same-model baseline, and the highest mean runtime-check pass rate at 93.4% among compared methods. A user study shows longer playtime and higher ratings. Code is coming soon.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ██░░░ 2 | 14% | Limited: single setting, weak baselines, or an observational association presented as causal. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (state of the art); scale (efficient); stakes (general AI).
How the score was computed
- Merit
- 4.0 / 10
- Adjusted merit
- 4.0 / 10
- Attention
- 79%
- Freshness
- 65%
- Hugging Face upvotes51 (reference 25, via hf-daily, Oct 8, 2026, 05:48 UTC)
- GitHub stars6 (reference 250, via hf-daily, Oct 8, 2026, 03:47 UTC)