AIarXiv
Heuristic editor, no API keyVerdict: RoutineSearchJev: A Fast and Calibrated System-1 Model for Search Agents
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions.
Key numbers
- 5.3 times faster decisions
- 4.7 times speedup in active search
- 45% to up to 54%
Caveats
- Preprint; not yet peer reviewed.
VerdictCompetent work. Briefs at most.
Abstract
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (x-fold, relative gain); verification (multiple benchmarks); scale (efficient); stakes (general AI).
How the score was computed
- Merit
- 5.1 / 10
- Adjusted merit
- 4.5 / 10
- Attention
- 51%
- Freshness
- 36%
- Citations0 (reference 15, via semantic-scholar, Oct 8, 2026, 02:05 UTC)
- Influential citations0 (reference 3, via semantic-scholar, Oct 8, 2026, 02:05 UTC)
- Hugging Face upvotes20 (reference 25, via hf-daily, Oct 8, 2026, 02:06 UTC)
- GitHub stars8 (reference 250, via hf-daily, Oct 8, 2026, 02:06 UTC)