AIarXiv

Heuristic editor, no API keyVerdict: Notable

EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?

Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric…

By Gao, Qu, Lei +19

Score██████░░░░5.8

Key numbers

  • 3.6% of multi-software attempts succeed

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Autonomous agents have made rapid progress in general-purpose computer use, but reliable automation of professional industrial engineering remains out of reach, as engineering workflows demand reasoning over geometric and physical constraints and dependencies preserved across software and design stages. We present EngiWorld, the first benchmark structured around the complete design loop: 1,301 expert-curated tasks spanning 6 engineering domains (CAD, CAE, CAM, BIM, EDA, and 3D visualization) and 26 professional software platforms, with both GUI and CLI interfaces and 6 task types ranging from software-selection to open-ended tasks. We further introduce an artifact-centric evaluation methodology built on a unified domain-verifier suite, which programmatically checks the geometric validity, physical feasibility, and rule compliance of final and intermediate artifacts, and scores quantitative design tasks continuously by specification attainment rather than binary success. Evaluation of seven frontier models reveals a substantial capability gap: the strongest model achieves an EngiScore of only 44.3, and just 3.6% of multi-software attempts succeed. EngiWorld provides the first rigorous foundation for measuring progress toward agents that operate professional engineering software end to end.

Hongcheng Gao, Hailong Qu, Yu Lei, Henghui Sun, Haoyang Li, Yipeng Wei, Naihao Xue, Xiaohan Yu, Zhuo Tao, Yihe Zang, Yajiao Wang, Jingyi Tang, Yi Li, Jingjing Zhou, Jie Luo, Bohan Zeng, Chengyu Shen, Hao Jiang, Chong Chen, Bowen Qu, Olive Huang, Zeqiang Wang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (general-purpose, many tasks); novelty (alternative to status quo); verification (multiple benchmarks); stakes (general AI). Red flags: weak evidence (preliminary).

How the score was computed

rank-2026-09-29

Score██████░░░░5.8

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.9 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.8 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
80%
Citations, upvotes, points, mentions.
Freshness
66%
Half-life decay since publication.
  • Hugging Face upvotes52 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars33 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 30, 2026, 11:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.CL
  • TOP, No.3 in the Artificial Intelligence edition of October 1, 2026.
  • BRIEF, No.4 in the Artificial Intelligence edition of September 30, 2026.
EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments? | Humanity's List · Humanity's List