AIarXiv

Heuristic editor, no API keyVerdict: Notable

Context Language Models

We introduce Context Language Models (CLMs), language models that natively manage their own context.

By Shao, Shen, Yin +10

Score██████░░░░5.6

Key numbers

  • 11.4% higher accuracy with 21.5%
  • 5% higher scores with 59%
  • 65% greater improvement with the

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety of tasks: 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus, 5% higher scores with 59% fewer FLOPs on 12-hour EdgeBench, and 65% greater improvement with the same compute on a 24-hour multi-repository agent-swarm task. Moreover, by shifting context management from external harness control to intrinsic model behavior, CLMs naturally enable both in-context and parametric learning of context-management strategies. We show that CLMs can be steered with natural-language instructions evolved through a standard skill-optimization loop, improving held-out accuracy by up to 35.9 points on a context-management task while reducing compute. We also introduce an online reinforcement learning method for CLMs, improving Qwen3.5-9B performance on BrowseComp-Plus by 47.6% while using 12% fewer FLOPs. Finally, we co-design Suffix Cache Reuse for CLM serving, further reducing server-side compute by 35% relative to standard SGLang at matched performance.

Rulin Shao, Shannon Zejiang Shen, Junjie Oscar Yin, Yuetai Li, Minheng Wang, Hamish Ivison, Radha Poovendran, Nathan Lambert, Teng Xiao, Mike Lewis, Wen-tau Yih, Luke Zettlemoyer, Pang Wei Koh

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude████░ 418%A qualitative jump: a capability or regime that did not exist before.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose, we report); breadth (zero/few-shot); gains (relative gain, state of the art, outperforms); verification (held-out test); stakes (general AI).

How the score was computed

rank-2026-09-29

Score██████░░░░5.6

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
6.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.9 / 10
Shrunk toward the desk prior by editor confidence (46%).
Attention
76%
Citations, upvotes, points, mentions.
Freshness
50%
Half-life decay since publication.
  • Hugging Face upvotes27 (reference 25, via hf-daily, Oct 2, 2026, 02:05 UTC)
  • GitHub stars277 (reference 250, via hf-daily, Oct 2, 2026, 02:05 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 30, 2026, 11:05 UTC. Paper type: method.
  • Categories: cs.AI, cs.CL, cs.LG
  • BRIEF, No.8 in the Artificial Intelligence edition of October 2, 2026.