AIarXiv

Heuristic editor, no API keyVerdict: Notable

Sharpening Tax in Post-Training

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of…

By Oh, Zeng, Qi +7

Score██████░░░░5.7

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to agentic tasks, where multi-turn tool use and interaction may require capabilities newly acquired during post-training. Our surprising finding is that pre-trained LLMs, equipped with a light inference harness, can serve as capable agents. Despite far lower accuracy (pass@1), they often surpass their post-trained counterparts in solution coverage (pass@K) given a sufficient test-time budget. We further analyze the underlying mechanism and show that post-training pushes tasks toward two extremes, always solved or never solved, and thereby improves sampling efficiency and consistency at the cost of solution coverage. To measure this cost, we propose Sharpening Tax, a diagnostic metric that quantifies the loss in test-time scalability after post-training. Across 14 base/post-trained model pairs from four families and three agentic benchmarks (42 cases in total), the tax is prevalent in most settings, can be estimated from a few rollouts, and correlates well with other metrics. Finally, we present posterior-tempered group sampling (PTGS), a simple plug-and-play Bayesian sampler that adapts the sampling temperature per prompt to its estimated difficulty. Applied during RL training in two agentic environments, PTGS pays a smaller tax than the fixed-temperature baseline, solving more tasks under repeated sampling while also improving single-shot accuracy.

Changdae Oh, Qi Zeng, Qi Qi, Andrey Zhmoginov, Deren Lei, Yun He, Hoang Phan, Hangoo Kang, Azalia Mirhoseini, Sharon Li

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 316%A genuinely new approach to an open problem.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (outperforms); novelty (unexpected); verification (multiple benchmarks); scale (scalable, efficient, improves with scale); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.7

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.8 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.9 / 10
Shrunk toward the desk prior by editor confidence (48%).
Attention
70%
Citations, upvotes, points, mentions.
Freshness
74%
Half-life decay since publication.
  • Hugging Face upvotes38 (reference 25, via hf-daily, Oct 2, 2026, 13:49 UTC)
  • GitHub stars1 (reference 250, via hf-daily, Oct 2, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.LG
  • BRIEF, No.2 in the Artificial Intelligence edition of October 2, 2026.