AIarXiv

Heuristic editor, no API keyVerdict: Notable

Recurrent Looped Transformer

State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length.

By Zhang, Feng, Qin

Score██████░░░░5.5

Key numbers

  • 100% accuracy in every seed
  • 97% final-state accuracy versus under
  • 1% for the Transformer

Caveats

  • Preprint; not yet peer reviewed.

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length. We introduce the Recurrent Looped Transformer (RLT), which splits its layers between a parallel causal encoder and a recurrent decoder. At each token, the decoder merges the encoder output with the previous token's final decoder state, so the computation path grows with sequence length at a fixed per-token cost. On six algorithmic tasks, we compare five splits of eight layers with an eight-layer Transformer over three seeds. Trained on at most 40 bits, two RLT splits generalize parity to 256 bits with 100% accuracy in every seed, while the Transformer stays at chance. On swap-based S₅ permutation tracking at eight times the training length, RLT reaches 97% final-state accuracy versus under 1% for the Transformer, and accuracy increases with decoder depth. On modular arithmetic beyond the training lengths, RLT reaches up to 93% versus 33% for the Transformer. Ablations show that these gains depend on the feedback: removing it drops parity and swap-based S₅ to chance at every split. Updating the feedback once per four-token chunk lets known tokens in a chunk run in parallel and keeps 64-bit parity at 99%, while permutation tracking depends on per-token feedback: chunking lowers length-64 swap-based S₅ from 100% to 20%.

Yifan Zhang, Jichen Feng, Shihan Qin

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (relative gain, versus baseline); verification (multiple benchmarks, ablations, code released). Red flags: derivative (comparative study).

How the score was computed

rank-2026-10-07

Score██████░░░░5.5

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.7 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
76%
Citations, upvotes, points, mentions.
Freshness
56%
Half-life decay since publication.
  • Hugging Face upvotes16 (reference 25, via hf-daily, Oct 8, 2026, 03:47 UTC)
  • GitHub stars908 (reference 250, via hf-daily, Oct 8, 2026, 03:47 UTC)

The record

  • Reviewed by heuristic-v5 on Oct 8, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.CL, cs.AI, cs.LG
  • TOP, No.1 in the Artificial Intelligence edition of October 8, 2026.