AIarXiv

Heuristic editor, no API keyVerdict: Notable

NeutronGym: Physics-Graded Neutron Instrument Design for LLM Agents

Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with.

By Ding, Do

Score████░░░░░░3.6

Key numbers

  • 11% to 77% of held-out
  • 69% at a second seed
  • 77% against 81%

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Designing a scientific instrument tests whether language-model agents can do physics rather than recall it, provided the grading cannot be argued with. We introduce NeutronGym, to our knowledge the first executable environment for neutron instrument design: agents build instruments through validating tools, McStas ray-traces what they build, and a level-resolved ladder grades syntax, runtime, structure and science with no LLM judge. Procedural families supply unlimited instances of a fixed layout whose design parameters the agent must set, with held-out parameter regimes; a curated slice, McStasBench, adds 16 tasks from published instruments behind memorization probes and a sandbox. Seven models reproduce at most 7 of the 16, none retrieves a reference, and none meets an improvement target. The environment also trains. Reinforcement learning on its reward takes Qwen3-8B from 11% to 77% of held-out instances of a family whose targets come from a hidden design (69% at a second seed), past an untrained Qwen3-32B, and the recipe holds, at one seed each, on three further gated families. The analysis says what that gain is. Without the ladder's partial credit it collapses by 60 points. From reward alone the trained model reaches what a classical optimizer reaches, at the agent's simulation budget, only when handed the closed-form physics (77% against 81%, a gap that does not separate at this size), while frontier models still solve 98-99%. Getting a trustworthy result meant failing four task designs that no-model baselines could solve, and we release the probes that found them.

Lijie Ding, Changwoo Do

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (versus baseline); novelty (alternative to status quo); verification (multiple benchmarks, held-out test, code released); stakes (general AI).

How the score was computed

rank-2026-09-29

Score████░░░░░░3.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.9 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.8 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
45%
Half-life decay since publication.

No attention signals recorded yet.

The record

  • Reviewed by heuristic-v2 on Oct 5, 2026, 02:05 UTC. Paper type: theory.
  • Categories: cs.AI, physics.ins-det
  • TOP, No.6 in the Physics edition of October 5, 2026.