AIarXiv

Heuristic editor, no API keyVerdict: Routine

OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific…

By Dai, Qi, Liu +28

Score█████░░░░░5.0

VerdictCompetent work. Briefs at most.

Read the originalPDFCode

Abstract

Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.

Dingyuan Dai, Heli Qi, Lei Liu, Yinxi Li, Baiding Chen, Zijun Dou, Qingcheng Zeng, Qi Kang, Oliver Sun, Eric Wang, Bo Zhou, Haixin Wang, Yufan Du, Shi Bo, Ruihan Lin, Mengqi Yuan, Dunjie Lu, Steven Dillmann, Yiming Shi, Tina Su, Amy Xin, Minghao Liu, Xi Wang, Xu Huang, Ge Zhang, Pengyu Nie, Zhen Yang, Jie Tang, Juanzi Li, Weihao Xuan, et al.

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty█░░░░ 116%A minor twist on a known approach.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: breadth (many tasks); gains (state of the art); verification (error bars); scale (efficient); stakes (general AI). Red flags: derivative (comparative study).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
4.8 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.3 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
77%
Citations, upvotes, points, mentions.
Freshness
25%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Oct 3, 2026, 02:05 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Oct 3, 2026, 02:05 UTC)
  • Hugging Face upvotes47 (reference 25, via hf-daily, Oct 3, 2026, 13:49 UTC)
  • GitHub stars5 (reference 250, via hf-daily, Oct 3, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 02:16 UTC. Paper type: empirical.
  • Categories: cs.AI
  • BRIEF, No.4 in the Artificial Intelligence edition of October 5, 2026.