AIarXiv

Heuristic editor, no API keyVerdict: Notable

Language Models that Play Chess and Explain Their Moves

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play.

By Bhaskar, Cheng, Chen

Score█████░░░░░4.8

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.

Adithya Bhaskar, Jeffrey Cheng, Danqi Chen

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory███░░ 318%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (orders of magnitude, outperforms); novelty (unexpected); verification (code released); scale (orders of magnitude, larger models); stakes (general AI). Red flags: derivative (we apply).

How the score was computed

rank-2026-09-29

Score█████░░░░░4.8

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.5 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.7 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
54%
Citations, upvotes, points, mentions.
Freshness
40%
Half-life decay since publication.
  • Hugging Face upvotes23 (reference 25, via hf-daily, Oct 6, 2026, 01:15 UTC)
  • GitHub stars1 (reference 250, via hf-daily, Oct 6, 2026, 01:15 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 5, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.CL
  • BRIEF, No.9 in the Artificial Intelligence edition of October 6, 2026.