AIarXiv

Heuristic editor, no API keyVerdict: Routine

HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents

We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers.

By Aji, Ramadhan, Zuhri +14

Score█████░░░░░5.3

VerdictCompetent work. Briefs at most.

Read the originalPDF

Abstract

We introduce HyperBrowseComp, a multilingual and multimodal browsing benchmark comprising 423 manually authored and human-validated questions across 13 languages, written by native or highly proficient speakers. Questions are designed to be extremely challenging. Each question targets a concise, publicly verifiable answer whose discovery requires locating obscure evidence, following multi-step clue chains, or inspecting heterogeneous sources such as videos, scanned documents, images, or maps. Easier questions are filtered out by evaluating them with models without internet access to reduce the likelihood that they can be answered with parametric knowledge alone. We evaluate several models using provider-native search and a shared external retrieval harness under a common agent protocol. To contextualize model performance and effort, we also conduct a human evaluation on a sample of the questions. HyperBrowseComp provides a challenging testbed for persistent information seeking across languages and evidence modalities, with difficulty arising from discovering and connecting evidence on the open web.

Alham Fikri Aji, Faiz Rizki Ramadhan, Zayd M. K. Zuhri, Seung Hun Eddie Han, Ryandito Diandaru, Qinrong Cui, Jan Christian Blaise Cruz, Badrinath Chandana, Peerawat Chomphooyod, Ahmed Attia, Jonibek Mansurov, Emilio Villa-Cueva, Canh Duong Nguyen, Imran Turganov, Minghao Wu, Peerat Limkonchotiwat, Irina Nikishina

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude██░░░ 218%Solid incremental gain on a meaningful problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); novelty (discovery); verification (multiple benchmarks, human evaluation, experimental validation); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░5.3

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.1 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.5 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
78%
Citations, upvotes, points, mentions.
Freshness
39%
Half-life decay since publication.
  • Hugging Face upvotes51 (reference 25, via hf-daily, Oct 6, 2026, 01:15 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 5, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.LG
  • TOP, No.6 in the Artificial Intelligence edition of October 6, 2026.