AIHugging Face

Heuristic editor, no API keyVerdict: Notable

MaLiang-Harness: A Programmable Path to Image and Video Generation

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation.

By Zhao, Zhang, Wang +4arXiv

Score██████░░░░5.6

Key numbers

  • 100% generation success on both
  • 96.0% of image tasks and
  • 76.9% of video tasks meeting

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.

Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu, Yu-Gang Jiang, Shuicheng Yan

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (new method); breadth (programmable); verification (error bars, code released).

How the score was computed

rank-2026-09-29

Score██████░░░░5.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.2 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.5 / 10
Shrunk toward the desk prior by editor confidence (36%).
Attention
90%
Citations, upvotes, points, mentions.
Freshness
42%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 1, 2026, 02:16 UTC)
  • Hugging Face upvotes279 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars15 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record