PhysicsarXiv
Heuristic editor, no API keyVerdict: NotablePre-registered tests of solid-state-physics-inspired LLM compression: a cluster-level negative result at small-language-model scale
We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a…
Key numbers
- 96% perplexity
VerdictWorth a reader's time today.
Abstract
We report a three-month autonomous research-agent program testing five solid-state-physics-inspired compression mappings on pretrained language models, with predictions committed to git before any pilot data and a 3-sigma gate deciding PASS or SHELVE. The common anchor -- area-law / Kohn-nearsighted decay of the one-particle density matrix -- has a distance face (P001 Wannier, P002 tight-binding) and a rank face (P003 DMRG-truncated MLPs, P005 Wilson-RG, P011 tensor-train embeddings). P005 was pre-empted at Phase 1; three of four Phase-3 pilots were falsified. On the attention face, GPT-2-medium attention-versus-distance is best fit by a stretched exponential in 12 of 16 median-layer heads once probe padding is excluded, and a tight-binding cutoff costs +96% perplexity (P002); on Pythia-160M the Wannier sparsity 0.054 +/- 0.004 is indistinguishable from PCA, random-Haar and identity baselines (P001). On the rank face, per-token tensor-train bond dimension does not track surprisal (r = 0.016 vs a pre-registered 0.65) and the format inflates rather than compresses (P011). P003 is mixed: its scaling claim shelved (r = -0.434), its MPO premise died at stage-0, and its cross-paper check, r = 0.523 as first written, collapses to 0.047 under the same correction, leaving both cross-paper checks null. The results invert the pre-registered prediction that most attention heads behave like Kohn-nearsighted insulators, pointing instead to critical, glassy or heavy-tailed regimes; the inversion is specific to the <= 350M scale tested, while the rank-face no-gain result held to 7-8B. We contribute the pre-registration + 3-sigma + cluster-framing + append-only-catalogue discipline -- including why our own enforcement gate was designed but not deployed -- four pre-registered negative results with full data release, and the inversion. The catalogue holds eighteen concluded studies, seventeen negative.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 18% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 20% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 22% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ███░░ 3 | 22% | A genuinely new approach to an open problem. |
| Trajectory | ██░░░ 2 | 10% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 8% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we report); gains (relative gain); novelty (alternative to status quo); design (registered); verification (error bars, code released); scale (scalable). Red flags: weak evidence (preliminary).
How the score was computed
- Merit
- 5.6 / 10
- Adjusted merit
- 4.7 / 10
- Attention
- 0%
- Freshness
- 80%
- Citations0 (reference 20, via semantic-scholar, Sep 29, 2026, 23:37 UTC)