AIarXiv

Heuristic editor, no API keyVerdict: Notable

SciExam for ENSO: Can AI Agents Build Climate Models?

Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell…

By Zhang, Liu, Xiu +4

Score████░░░░░░3.8

Caveats

  • Preprint; not yet peer reviewed.

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Language-model agents are increasingly asked to carry out open-ended scientific research, yet their results are usually graded against a known answer, a rubric, or a language-model reviewer, none of which can tell whether a new scientific model is valid. The AI Science Exam for El Nino-Southern Oscillation (SciExam for ENSO) is a benchmark in which agents build low-order stochastic models of ENSO, the dominant mode of interannual climate variability, from real observations. Within a six-hour budget, agents process the observations, write their own diagnostics, which are then frozen, and develop a model using only these diagnostics as feedback. Hidden graders then test whether the model reproduces ENSO's statistics, recovers unobserved variables, and forecasts held-out years, and score a published model in the same way. Across twelve agent systems, six produce models that score higher than the published model, mainly through better reconstruction and forecasting. The simplified forms of the stronger models are each compatible with one of the two competing explanations of ENSO's warm-cold asymmetry, an open debate that the task never mentions. Controlled runs of the top system under varied information suggest that its scores do not come from recalling the dated observational record and that the information it receives shapes how it builds its model. SciExam for ENSO can thus evaluate agent research where no answer is known, and the results suggest that agents can already build competitive models whose structures bear on questions that scientists still debate.

Yinling Zhang, Langchen Liu, Dongbin Xiu, Xueyan Zou, Xu Kuang, Mengdi Wang, Shilong Liu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (new method); breadth (many tasks); verification (multiple benchmarks, held-out test, code released); stakes (general AI).

How the score was computed

rank-2026-10-07

Score████░░░░░░3.8

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.2 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.5 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
87%
Half-life decay since publication.

No attention signals recorded yet.

The record

  • Reviewed by heuristic-v5 on Oct 8, 2026, 05:48 UTC. Paper type: method.
  • Categories: cs.AI, cs.LG, physics.ao-ph
  • TOP, No.5 in the Climate & Energy edition of October 8, 2026.