AIarXiv

Heuristic editor, no API keyVerdict: Notable

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks.

By Lee, Baek, Jeong +5

Score██████░░░░5.9

Key numbers

  • 74.1% to 78.0% with GPT-5
  • 61.3% to 82.3% with Gemini-3

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves solutions and search queries with fixed model parameters. At each iteration, a retrieval gate lets the LLM assess its knowledge gap and choose to retrieve new documents, reuse stored ones, or proceed without them. An inner loop refines queries and ranks documents by the solution scores they are predicted to yield; an outer loop generates candidates in parallel from these documents and records the evaluated outcomes for later searches. Across 21 optimization tasks with one candidate per iteration, EvoDuet raises OpenEvolve's normalized discovery gain from 74.1% to 78.0% with GPT-5.6-Luna and from 61.3% to 82.3% with Gemini-3.8-Flash, whereas Qwen3.5-9B does not benefit. Our best runs surpass the previously reported best scores on eight tasks, including Swap Reduction on Q20 and Rosetta, and match them on three more. EvoDuet also improves with other scaffolds (e.g., Top-K, EvoX) on Sums/Diffs and Denoising, demonstrating its applicability across evolutionary search scaffolds.

Young-Jun Lee, Jinheon Baek, Soyeong Jeong, Minki Kang, Seungyeon Jwa, Jonghyun Choi, Seungho Han, Dongyeop Kang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (relative gain, outperforms); novelty (discovery); verification (multiple benchmarks); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.4 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
84%
Citations, upvotes, points, mentions.
Freshness
80%
Half-life decay since publication.
  • Hugging Face upvotes66 (reference 25, via hf-daily, Oct 1, 2026, 13:49 UTC)
  • GitHub stars0 (reference 250, via hf-daily, Oct 1, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 1, 2026, 02:16 UTC. Paper type: method.
  • Categories: cs.CL
  • LEAD, No.1 in the Artificial Intelligence edition of October 1, 2026.
EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery | Humanity's List · Humanity's List