AIHugging Face

Heuristic editor, no API keyVerdict: Notable

Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve.

By Lee, Kang, Park +3arXiv

Score██████░░░░5.6

Key numbers

  • 2.7 times lower serving cost
  • 45.96% pass
  • 16.85% vs

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.

Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude████░ 418%A qualitative jump: a capability or regime that did not exist before.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose, we report); breadth (many tasks); gains (x-fold, relative gain, versus baseline); novelty (alternative to status quo); verification (multiple benchmarks); scale (scalable, low cost); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
6.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
5.3 / 10
Shrunk toward the desk prior by editor confidence (50%).
Attention
68%
Citations, upvotes, points, mentions.
Freshness
42%
Half-life decay since publication.
  • Citations0 (reference 15, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Influential citations0 (reference 3, via semantic-scholar, Sep 29, 2026, 23:53 UTC)
  • Hugging Face upvotes35 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars5 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record