AIarXiv

Heuristic editor, no API keyVerdict: Notable

Towards Looped Models Done Right, Part II: Rethinking at Fixed Points

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL).

By Huang, Shi, Chen +4

Score█████░░░░░4.9

Key numbers

  • 1.79x faster
  • 2x faster than backpropagating through
  • 3x smaller KV cache matches

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Every recurrence of a looped language model adds cost in training, decoding, prefill, and reinforcement learning (RL). The closer recurrent states get to fixed points, the less the path to them matters. This enables truncated backpropagation in training; terminal key-value (KV) sharing for decoding with almost no loss in accuracy; a distilled student that prefills up to 1.79x faster; and RL updates that compute gradients from saved rollout states, 2x faster than backpropagating through the replayed trajectory. We therefore improve the two components of training that shape these fixed points: the depth prior and input injection. Fixed-depth training breaks KV sharing, and Huginn's broad depth prior supports sharing but dilutes supervision at the target depth more than sharing requires; we learn the prior from prediction feedback, with an entropy term that keeps it broad. Existing injection schemes let the state's component along the input amplify or cancel the injection; we remove this component with orthogonal injection. From 100M to 1.6B parameters, the learned prior and orthogonal injection lower perplexity at every scale relative to Huginn's prior and existing injection schemes, respectively. At 1.6B, the learned prior with a 3x smaller KV cache matches the downstream average of fixed-depth training with the full cache.

Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: gains (x-fold); verification (code released); scale (scalable, efficient); stakes (general AI).

How the score was computed

rank-2026-09-29

Score█████░░░░░4.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.5 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (38%).
Attention
48%
Citations, upvotes, points, mentions.
Freshness
70%
Half-life decay since publication.
  • Hugging Face upvotes17 (reference 25, via hf-daily, Oct 7, 2026, 01:27 UTC)
  • GitHub stars30 (reference 250, via hf-daily, Oct 7, 2026, 01:27 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 6, 2026, 13:49 UTC. Paper type: empirical.
  • Categories: cs.LG
  • BRIEF, No.6 in the Artificial Intelligence edition of October 7, 2026.