PhysicsarXiv

Heuristic editor, no API keyVerdict: Notable

Pre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale

We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a…

By Lu

Score████░░░░░░4.3

Key numbers

  • 96% perplexity

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.

Jun-qiang Lu

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 318%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 320%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 322%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 322%A genuinely new approach to an open problem.
Trajectory██░░░ 210%Some room to improve with obvious engineering.
Stakes██░░░ 28%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we report); gains (relative gain); novelty (alternative to status quo); design (registered); verification (error bars, code released); scale (scalable). Red flags: weak evidence (preliminary).

How the score was computed

rank-2026-09-29

Score████░░░░░░4.3

Score = 10 × (75% × adjusted merit / 10 + 15% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.7 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
80%
Half-life decay since publication.
  • Citations0 (reference 20, via semantic-scholar, Sep 29, 2026, 23:37 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • Categories: cond-mat.dis-nn, cond-mat.mtrl-sci, cond-mat.other, cs.LG
  • BRIEF, No.1 in the Physics edition of September 29, 2026.