How we decide what matters
Google published the Transformer and moved on. We will not. Every story on this site was scored against one rubric, calibrated on the work that changed the world.
Principles
- Merit first. What the work makes possible outranks how loudly it was announced.
- Evidence gates everything. A spectacular claim with weak proof is discounted before it can lead.
- Attention is a signal, not a verdict. The overlooked slot exists for high-merit work nobody has noticed yet.
- Every desk ranks by its own lights. Medicine weighs evidence and stakes; AI weighs leverage and trajectory.
- Every score shows its work. Each story page lists the rubric levels, the weights, the signals, and the formula.
- The editor is named. Claude reviews what matters most; a deterministic heuristic triages the rest, and we say which.
The rubric
Leverage
How many other people can build on this, and how much does it multiply their work?
- 0Nothing reusable. A one-off result.
- 1Narrow application of an existing method to a new dataset or setting.
- 2Reusable within one subfield (a technique, dataset, or protocol a few groups will adopt).
- 3A method or resource many groups across the field will adopt within a year.
- 4A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
- 5A new primitive or platform that reorganizes fields (the Transformer, CRISPR-Cas9, the mRNA-LNP platform).
Magnitude
How large is the advance over the best prior work, taking the claims at face value?
- 0No improvement, or a null result on a question nobody doubted.
- 1Marginal: within noise or a few percent on a saturated benchmark.
- 2Solid incremental gain on a meaningful problem.
- 3Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
- 4A qualitative jump: a capability or regime that did not exist before.
- 5A discontinuity (AlexNet cutting ImageNet top-5 error from 26% to 15%, AlphaFold 2 at median GDT 92, 95% vaccine efficacy).
Evidence
How strong is the proof that the claims are true? Judge this independently of magnitude.
- 0Assertion without data, or data that cannot support the claim.
- 1Anecdotal: tiny n, one cherry-picked benchmark, no controls or baselines, or an extraordinary claim with ordinary proof.
- 2Limited: single setting, weak baselines, or an observational association presented as causal.
- 3Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
- 4Strong: large scale, preregistered, independently replicated, or a well-powered randomized trial.
- 5Definitive: phase 3 randomized evidence on hard endpoints, multi-lab replication, or community verification at scale.
Novelty
How far does this depart from what the field expected?
- 0Replication or derivative work.
- 1A minor twist on a known approach.
- 2A new combination of known ideas.
- 3A genuinely new approach to an open problem.
- 4Challenges a prevailing assumption with evidence.
- 5Overturns consensus (H. pylori causes ulcers, nucleoside modification makes mRNA viable, attention without recurrence).
Trajectory
Does this get better with scale, time, or investment? Is there headroom?
- 0A dead end, or a result at its ceiling.
- 1Limited headroom.
- 2Some room to improve with obvious engineering.
- 3A clear path to scale.
- 4Improves predictably with data, compute, or investment.
- 5Opens a compounding frontier (scaling laws, sequencing cost curves, programmable editing).
Stakes
If this holds up, how many people's lives change, and how directly?
- 0No plausible effect outside the paper.
- 1Of interest to a niche academic community.
- 2Benefits a professional community (practitioners, clinicians, engineers).
- 3Meaningful benefit to many people within a few years.
- 4Large-scale human impact: millions of people, or a major industry.
- 5Civilization-scale: pandemics, energy, food, or general-purpose intelligence.
The evidence gate
Leverage, magnitude, novelty, trajectory, and stakes are scored as if the claims are true. Evidence is scored on its own. Then weak evidence discounts everything else before merit is computed:
- Evidence 0×0.3
- Evidence 1×0.55
- Evidence 2×0.8
- Evidence 3×1
- Evidence 4×1
- Evidence 5×1
That is why LK-99 never leads this paper: a room-temperature superconductor would be a level-5 story, but a single group’s partial data is evidence 1, and the gate cuts its other dimensions by nearly half.
Verdicts
- Verdict: Landmark8.4+Belongs in the canon if it holds up.
- Verdict: Major6.8+A leading story on any desk.
- Verdict: Notable5.2+Worth a reader's time today.
- Verdict: Routine3.0+Competent work. Briefs at most.
- Verdict: Skip0.0+Not for this paper.
The desks
| Desk | Leverage | Magnitude | Evidence | Novelty | Trajectory | Stakes | Blend (merit / attention / fresh) |
|---|---|---|---|---|---|---|---|
| AI | 24% | 18% | 14% | 16% | 18% | 10% | 65% / 25% / 10% |
| Biology | 24% | 16% | 20% | 20% | 10% | 10% | 75% / 15% / 10% |
| Medicine | 10% | 20% | 32% | 8% | 5% | 25% | 80% / 10% / 10% |
| Climate | 16% | 20% | 20% | 10% | 14% | 20% | 75% / 15% / 10% |
| Physics | 18% | 20% | 22% | 22% | 10% | 8% | 75% / 15% / 10% |
| News | 15% | 20% | 20% | 10% | 10% | 25% | 50% / 25% / 25% |
The canon
The editor is calibrated against these. Landmarks show what the top of the scale means. Cautionary tales show what fooled people. Scores are ours, on the same rubric every story gets.
AI · 2017 · Landmark
merit 8.9Verdict: LandmarkAttention Is All You Need
Replaced recurrence and convolution in sequence models with stacked self-attention. The Transformer set translation records (28.4 BLEU English-German, 41.8 English-French) after 3.5 days of training on eight GPUs, a fraction of the cost of earlier systems.
LessonThe primitive that reorganized AI arrived as a translation paper with two benchmarks and a parsing side result. The signals were there on day one: a simpler architecture, a large efficiency gain, parallelism that invites scale, and an early hint of generality. Weight a new, general, parallelizable primitive above a bigger number on a narrow task.
Leverage 5 · Magnitude 4 · Evidence 3 · Novelty 5 · Trajectory 5 · Stakes 4
AI · 2012 · Landmark
merit 8.7Verdict: LandmarkImageNet Classification with Deep Convolutional Neural Networks
A 60-million-parameter convolutional network trained on two GPUs with rectified linear units and dropout won ILSVRC-2012 with 15.3% top-5 error, against 26.2% for the runner-up.
LessonA discontinuity on a hard, unsaturated benchmark, scored blind by the organizers, is the clearest landmark signal there is. The ingredients were not new; the jump came from scale, GPUs, and data. When an old idea suddenly works at a new scale, the trajectory is the story.
Leverage 4 · Magnitude 5 · Evidence 4 · Novelty 4 · Trajectory 5 · Stakes 4
AI · 2015 · Landmark
merit 7.7Verdict: MajorDeep Residual Learning for Image Recognition
Identity shortcut connections let networks learn residual functions, making 152-layer models trainable. An ensemble reached 3.57% top-5 error and won ILSVRC 2015 classification, plus the ImageNet detection and localization and COCO detection and segmentation tracks.
LessonRates MAJOR, not LANDMARK, as published: the idea was a one-line architectural change, and simple changes read as incremental. Its leverage (the default backbone for most vision models since) and blind-test wins across five competition tracks were the tells. Do not confuse simplicity with smallness.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 4 · Stakes 4
AI · 2018 · Landmark
merit 7.0Verdict: MajorImproving Language Understanding by Generative Pre-Training
Pre-trained a Transformer language model on unlabeled books, then fine-tuned it for each task with a small output head. One task-agnostic recipe improved the state of the art on 9 of 12 language-understanding benchmarks, including +8.9 points on story completion and +5.7 on RACE.
LessonPretrain a general model, then put a classifier on top: the recipe BERT and every later GPT followed arrived as an unreviewed technical report with modest per-task gains. The signal was breadth (one model, many tasks, minimal architecture changes) and a trajectory that plainly improved with more data and parameters.
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 4
AI · 2018 · Landmark
merit 7.4Verdict: MajorBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Masked-token pretraining gave a Transformer encoder context from both directions; fine-tuned with one extra output layer, it set new records on eleven language tasks, lifting GLUE by 7.7 points and SQuAD v2.0 F1 by 5.1.
LessonReleased weights plus gains across a whole leaderboard at once is how leverage announces itself. BERT became the default starting point for language work within months because anyone could fine-tune it. Count open, reusable artifacts as leverage, not as a footnote.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 4
AI · 2020 · Landmark
merit 7.2Verdict: MajorAn Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Applied a standard Transformer directly to sequences of 16x16 image patches. With large-scale pretraining it matched or beat the best convolutional networks on ImageNet, CIFAR-100, and VTAB while using substantially less compute to train.
LessonChallenging a field's core inductive bias (convolutions for vision) with a plain, general architecture is a novelty signal even when the headline number is only competitive. The caveat that it needed very large pretraining data was also the trajectory signal: it kept improving with data where convolutional networks flattened.
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 3
AI · 2020 · Landmark
merit 8.6Verdict: LandmarkLanguage Models are Few-Shot Learners
Scaled an autoregressive language model to 175 billion parameters, ten times any previous dense model, and showed it could perform new tasks from a few examples in the prompt, with no gradient updates, sometimes rivaling fine-tuned systems.
LessonA qualitative capability (in-context learning) emerging from scale alone is a landmark signal, and the paper's own contamination analysis is the kind of self-scrutiny that earns evidence points. Closed weights limited direct reuse, which is why leverage stops at 4.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 4 · Trajectory 5 · Stakes 5
AI · 2020 · Landmark
merit 7.8Verdict: MajorScaling Laws for Neural Language Models
Found that language-model loss falls as a smooth power law in parameters, data, and compute across more than seven orders of magnitude, and derived how to split a fixed compute budget between model size and data.
LessonPredictability is itself a capability: a result that let labs forecast returns before spending became the basis for enormous investment. The recommended allocation was later revised (Chinchilla, 2022), a reminder that single-lab empirical laws earn evidence 3, not 5, even when the trend holds.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 5 · Stakes 4
AI · 2020 · Landmark
merit 6.8Verdict: MajorDenoising Diffusion Probabilistic Models
Trained diffusion models with a weighted variational bound linked to denoising score matching, reaching a state-of-the-art FID of 3.17 on unconditional CIFAR-10 and GAN-level samples on 256x256 LSUN, with code released.
LessonAs published it only matched GANs, so it rates MAJOR by a hair. The tells were a principled, stable training objective (no adversarial game) and a clean path to scale; within two years diffusion displaced GANs for images and spread to audio, video, and protein design. Weight stable, general objectives for their trajectory.
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 3
AI · 2021 · Landmark
merit 7.5Verdict: MajorLearning Transferable Visual Models From Natural Language Supervision
Contrastively pretrained image and text encoders on 400 million image-caption pairs. The model classifies images zero-shot from text prompts, matching the original ResNet-50 on ImageNet without any of its 1.28 million labeled training examples.
LessonZero-shot transfer across more than 30 datasets, with released weights, made CLIP a component in countless systems, from search to image generation. Evidence breadth and open weights were the signals; the idea (contrastive image-text learning) was not new, the scale and the evaluation were.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 4 · Stakes 3
AI · 2022 · Landmark
merit 7.6Verdict: MajorTraining language models to follow instructions with human feedback
Fine-tuned GPT-3 on labeler demonstrations and then with reinforcement learning from human preference rankings. Labelers preferred the 1.3-billion-parameter InstructGPT to the 175-billion-parameter GPT-3, a model with 100 times more parameters.
LessonA small model beating a 100x larger one on what users actually want is a qualitative signal, and the recipe (supervised fine-tuning, then RLHF) was simple enough for every lab to copy. Evidence rests on the authors' own prompts and labelers, so it stays at 3; the stakes were visible because the method already served real traffic.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 5
AI · 2014 · Landmark
merit 5.8Verdict: NotableAdam: A Method for Stochastic Optimization
Combined momentum with per-parameter adaptive step sizes from running estimates of the gradient's first and second moments, giving an optimizer that needs little tuning and copes with noisy, sparse gradients.
LessonOnly NOTABLE as published: small-scale experiments, a combination of known ideas, and a convergence proof later shown to be flawed. It became the default optimizer for a decade anyway. The lesson is leverage: a drop-in tool every practitioner can adopt tomorrow deserves a look even when its paper is modest.
Leverage 4 · Magnitude 2 · Evidence 3 · Novelty 2 · Trajectory 3 · Stakes 3
AI · 2021 · Landmark
merit 6.2Verdict: NotableLoRA: Low-Rank Adaptation of Large Language Models
Froze pretrained weights and trained small low-rank update matrices instead. On GPT-3 175B it cut trainable parameters 10,000-fold and GPU memory 3-fold while matching full fine-tuning quality, with no added inference latency.
LessonAnother NOTABLE-as-published tool that became infrastructure. The numbers that mattered were efficiency ratios (10,000x fewer trainable parameters, 3x less memory) on the largest available model, plus released code: cheap adaptation of big models was a need every lab had.
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 2 · Trajectory 3 · Stakes 3
AI · 2022 · Landmark
merit 6.5Verdict: NotableFlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Computed exact attention in tiles that stay in fast on-chip GPU memory, cutting traffic to high-bandwidth memory. Training ran 3x faster on GPT-2 and 15% faster on BERT-large, and longer contexts produced the first better-than-chance results on Path-X.
LessonSystems work that removes a bottleneck compounds: exact, IO-aware attention now sits inside nearly every training and inference stack. The rubric rates it NOTABLE as published because the gains were 2-3x; notice when a speedup is exact (no quality trade-off) and general (every Transformer).
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 3 · Trajectory 3 · Stakes 3
AI · 2022 · Landmark
merit 6.9Verdict: MajorChain-of-Thought Prompting Elicits Reasoning in Large Language Models
Showed that prompting large language models with a few worked examples of intermediate reasoning steps sharply improves arithmetic, commonsense, and symbolic reasoning; with eight exemplars a 540-billion-parameter model set a state of the art on GSM8K math word problems.
LessonA zero-cost method whose benefit appears only at scale is a trajectory signal: it pointed straight at reasoning-focused training. Evidence (three model families, several task types) was solid, and the gain was largest on the hardest benchmark.
Leverage 3 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 4
AI · 2023 · Landmark
merit 5.8Verdict: NotableSelf-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
Learned image representations by predicting, in embedding space, the representations of masked target blocks from a single context block, with no hand-crafted augmentations. A ViT-Huge/14 trained in under 72 hours on 16 A100 GPUs transferred well to classification, counting, and depth tasks.
LessonA mid-level anchor: a genuinely different self-supervised objective, solid evidence, and real efficiency gains, but reuse within one subfield and competitive rather than dominant numbers. It rates NOTABLE on the anchors; MAJOR would need adoption beyond vision self-supervision or a clear margin over its contemporaries.
Leverage 3 · Magnitude 3 · Evidence 3 · Novelty 3 · Trajectory 3 · Stakes 2
AI · 2025 · Landmark
merit 7.9Verdict: MajorDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Showed that large-scale reinforcement learning with rule-based rewards, without supervised reasoning traces, makes long chains of reasoning emerge (AIME 2024 pass@1 rose from 15.6% to 71.0% for R1-Zero). The released R1 model matched OpenAI o1 on reasoning benchmarks, and six distilled dense models were open-sourced.
LessonOpen weights plus a simple, reproducible recipe turned a frontier capability into a public baseline within weeks. Evidence stays at 3 because the benchmarks were self-reported and the training data unreleased; leverage and novelty (pure RL instead of human reasoning traces) carried it to MAJOR.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 5
Biology · 2012 · Landmark
merit 8.8Verdict: LandmarkA programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity
Showed that the bacterial protein Cas9, guided by a pair of RNAs, cuts DNA at a site chosen by the guide sequence, and that the two RNAs can be joined into a single guide that still works, making the cutter reprogrammable by changing a short RNA.
LessonThe abstract's last clause (the potential for RNA-programmable genome editing) was the whole story. Test-tube evidence only, yet leverage, novelty, and trajectory were unmistakable: retargeting by editing a 20-letter RNA instead of engineering a protein. Programmability is the signal.
Leverage 5 · Magnitude 5 · Evidence 3 · Novelty 4 · Trajectory 5 · Stakes 5
Biology · 2013 · Landmark
merit 8.5Verdict: LandmarkMultiplex genome engineering using CRISPR/Cas systems
Adapted two CRISPR/Cas systems to edit genes in human and mouse cells, and showed that several guides placed in one array could edit multiple sites at once.
LessonThe step from a test tube to mammalian genomes, published alongside an independent demonstration from another lab (Mali et al., same issue), is what made CRISPR a platform. Independent replication at the moment of publication is rare and worth an evidence point.
Leverage 5 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 5 · Stakes 5
Biology · 2006 · Landmark
merit 8.5Verdict: LandmarkInduction of pluripotent stem cells from mouse embryonic and adult fibroblast cultures by defined factors
Turned mouse skin cells back into embryonic-like stem cells by adding four defined factors, narrowed down from 24 candidates. The induced cells behaved like embryonic stem cells and contributed to developing mouse embryos.
LessonFour genes undoing differentiation contradicted a long-held view that development runs one way outside eggs and nuclear transfer. A recipe any lab could repeat (several did within a year) made it a platform for disease modeling and regenerative medicine.
Leverage 5 · Magnitude 4 · Evidence 3 · Novelty 5 · Trajectory 4 · Stakes 4
Biology · 2005 · Landmark
merit 7.4Verdict: MajorMillisecond-timescale, genetically targeted optical control of neural activity
Put a light-gated channel from algae (channelrhodopsin-2) into cultured mammalian neurons and used brief flashes of light to make them fire, with millisecond precision.
LessonA narrow-looking demonstration in cultured neurons became the standard way to test cause and effect in brain circuits. Genetic targeting plus millisecond control was a capability neuroscience had never had; score what a tool enables, not only the system it was first tested in.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 3
Biology · 2016 · Landmark
merit 7.2Verdict: MajorProgrammable editing of a target base in genomic DNA without double-stranded DNA cleavage
Introduced base editing, which changes a single DNA letter at a programmed site without cutting both strands. In several human and mouse cell lines it corrected 15-75% of target sites with typically 1% or fewer unwanted insertions or deletions.
LessonMost disease-linked variants are single-letter changes, so a precise way to change one letter was the obvious next platform. Clear efficiency and purity numbers across several cell lines were the evidence; the trajectory followed quickly, with new editor types within a year.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 4 · Stakes 4
Biology · 2019 · Landmark
merit 7.6Verdict: MajorSearch-and-replace genome editing without double-strand breaks or donor DNA
Introduced prime editing, in which the guide RNA itself carries the new sequence to be written. The authors reported more than 175 edits in human cells covering every kind of single-letter change plus small insertions and deletions, and estimated the approach could in principle address up to 89% of known disease-linked variants.
LessonA new mechanism rather than a better cutter. Breadth of demonstrated edits and head-to-head comparisons with earlier methods were the evidence; the 89% figure is a theoretical ceiling, not an achievement, and copy should say so.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 4
Biology · 2005 · Landmark
merit 8.7Verdict: LandmarkSuppression of RNA recognition by Toll-like receptors: the impact of nucleoside modification and the evolutionary origin of RNA
Showed that the innate immune sensors that react to RNA largely ignore RNA built with naturally occurring modified nucleosides such as pseudouridine, so immune cells exposed to modified RNA release far fewer inflammatory signals.
LessonPublished as basic immunology in a specialist journal and largely overlooked for years, it removed the main obstacle to mRNA medicines. The overlooked-gem archetype: a simple mechanistic finding with an enormous implication, from an unfashionable lab. The Overlooked slot exists for papers like this.
Leverage 5 · Magnitude 4 · Evidence 3 · Novelty 5 · Trajectory 4 · Stakes 5
Biology · 2021 · Landmark
merit 8.8Verdict: LandmarkHighly accurate protein structure prediction with AlphaFold
A redesigned neural network predicted protein structures with near-experimental accuracy in the blind CASP14 assessment (median backbone GDT around 92), far ahead of every other method, and its code was released with the paper.
LessonA blind, independently scored assessment is the gold standard for evidence in computational science. When a method wins by a discontinuous margin under those conditions and the code is public, it is a landmark on day one.
Leverage 5 · Magnitude 5 · Evidence 4 · Novelty 4 · Trajectory 4 · Stakes 4
Biology · 2023 · Landmark
merit 7.6Verdict: MajorDe novo design of protein structure and function with RFdiffusion
Turned a protein structure predictor into a generative diffusion model for protein backbones and used it to design binders, symmetric assemblies, and scaffolds; hundreds of designs were tested in the lab, and an electron-microscopy structure of one binder matched its design almost exactly.
LessonTwo known ideas (diffusion models and a structure predictor) combined into a general design tool, backed by hundreds of laboratory tests. Experimental validation at scale is what separates protein-design landmarks from computer-only demonstrations.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 4 · Stakes 4
Medicine · 2021 · Landmark
merit 9.1Verdict: LandmarkEfficacy and Safety of the mRNA-1273 SARS-CoV-2 Vaccine
A phase 3, observer-blinded, placebo-controlled trial randomized 30,420 adults at 99 U.S. sites. Two doses of the Moderna mRNA vaccine were 94.1% efficacious against symptomatic Covid-19 (95% CI, 89.3 to 96.8), and all 30 severe cases occurred in the placebo group.
LessonThe top of the evidence scale: a large, registered, randomized, blinded trial on a clinical endpoint with a huge effect. It also validated a platform built on basic immunology from 2005. Landmarks in medicine usually look like this; single-arm surprises rarely do.
Leverage 4 · Magnitude 5 · Evidence 5 · Novelty 2 · Trajectory 3 · Stakes 5
Medicine · 2020 · Landmark
merit 9.1Verdict: LandmarkSafety and Efficacy of the BNT162b2 mRNA Covid-19 Vaccine
In a multinational, placebo-controlled, observer-blinded trial of 43,548 participants aged 16 and older, two doses of the Pfizer-BioNTech modified-mRNA vaccine were 95% effective against Covid-19 (8 cases versus 162 with placebo).
LessonThe first phase 3 readout for an mRNA vaccine and a direct descendant of the 2005 nucleoside-modification finding. A 95% effect in a trial this large is a discontinuity; the right caveats at the time were duration of protection and rare adverse events, not efficacy.
Leverage 4 · Magnitude 5 · Evidence 5 · Novelty 2 · Trajectory 3 · Stakes 5
Medicine · 2011 · Landmark
merit 3.8Verdict: RoutineChimeric antigen receptor-modified T cells in chronic lymphoid leukemia
A patient with treatment-resistant chronic lymphocytic leukemia received a low dose of their own T cells engineered to target the CD19 marker on leukemia cells. The cells multiplied more than 1,000-fold in the body, persisted for six months, and produced a complete remission still ongoing at ten months.
LessonOne patient. The rubric rates it ROUTINE as published, and that is correct triage for a single-patient report: a dramatic n=1 is a reason to watch, not to lead. The evidence came with the trials that followed and the first approval in 2017. Flag first-in-human responses with objective readouts as stories to follow.
Leverage 4 · Magnitude 5 · Evidence 1 · Novelty 3 · Trajectory 4 · Stakes 4
Medicine · 2021 · Landmark
merit 3.6Verdict: RoutineCRISPR-Cas9 Gene Editing for Sickle Cell Disease and β-Thalassemia
Reported the first two patients treated with their own gene-edited blood stem cells, designed to switch fetal hemoglobin back on. More than a year later, both no longer needed transfusions, and the patient with sickle cell disease had no further pain crises.
LessonTwo patients: ROUTINE as published, however striking. As with CAR-T, objective and durable first-in-human responses are a story to follow until the full trial reports (approval came in 2023). The headline should say 'first patients', never 'cure'.
Leverage 3 · Magnitude 5 · Evidence 1 · Novelty 3 · Trajectory 4 · Stakes 4
Medicine · 2021 · Landmark
merit 7.2Verdict: MajorOnce-Weekly Semaglutide in Adults with Overweight or Obesity
In a 68-week double-blind trial, 1,961 adults with obesity or overweight and no diabetes were randomized 2:1 to weekly semaglutide 2.4 mg or placebo, both with lifestyle support. Mean weight change was -14.9% versus -2.4%, a difference of 12.4 percentage points (95% CI, 11.5 to 13.4).
LessonA well-powered randomized trial with a large effect on a condition affecting hundreds of millions. It stops at MAJOR because body weight is not a hard outcome and the drug class was known; the cardiovascular-outcomes evidence came later. Keep surrogate and hard endpoints apart even when the effect is big.
Leverage 2 · Magnitude 4 · Evidence 4 · Novelty 2 · Trajectory 3 · Stakes 4
Medicine · 2024 · Landmark
merit 9.1Verdict: LandmarkTwice-Yearly Lenacapavir or Daily F/TAF for HIV Prevention in Cisgender Women
A phase 3 double-blind trial randomized 5,338 adolescent girls and young women in South Africa and Uganda. None of the 2,134 who received a lenacapavir injection every six months acquired HIV, against a background incidence of 2.41 per 100 person-years; daily oral F/TAF did not beat background.
LessonZero infections on a hard endpoint in a large randomized trial is as clean a discontinuity as medicine gets. It also solved what daily pills could not (adherence to both oral arms was low), which is the stakes argument: two injections a year reach people pills do not.
Leverage 3 · Magnitude 5 · Evidence 5 · Novelty 3 · Trajectory 3 · Stakes 5
Medicine · 1984 · Landmark
merit 5.5Verdict: NotableUnidentified curved bacilli in the stomach of patients with gastritis and peptic ulceration
Found spiral bacteria in stomach biopsies from 58 of 100 consecutive patients undergoing gastroscopy and grew a previously undescribed species from 11. The bacteria were present in almost all patients with active chronic gastritis, duodenal ulcer, or gastric ulcer.
LessonObservational evidence in 100 patients against a firm consensus that ulcers came from stress and acid. The rubric rates it NOTABLE as published because evidence 2 caps it. This is the archetype the Overlooked slot exists for: the novelty signal (a new organism tied to a common disease) should have prompted a closer look, not a dismissal.
Leverage 3 · Magnitude 4 · Evidence 2 · Novelty 5 · Trajectory 3 · Stakes 4
Climate · 1980 · Landmark
merit 7.6Verdict: MajorLixCoO2 (0<x≤1): A new cathode material for batteries of high energy density
Showed that lithium can be pulled out of layered lithium cobalt oxide (down to about Li0.07CoO2) and put back, giving open-circuit voltages of 4 to 5 volts against lithium, roughly double the sulfide cathodes of the day, and a theoretical energy density of about 1.1 kWh per kilogram.
LessonThe oxide cathode doubled cell voltage and let cells be built without a metallic lithium anode: the road to the lithium-ion battery. It rates MAJOR as published (one lab, no full cell). Energy technologies often start this way; a doubling of a fundamental figure of merit is the signal.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 4
Climate · 1997 · Landmark
merit 6.3Verdict: NotablePhospho-olivines as Positive-Electrode Materials for Rechargeable Lithium Batteries
Reported reversible lithium extraction from olivine LiFePO4 at 3.5 volts, with 100-110 mAh/g despite extracting only about 0.6 lithium per formula unit, and proposed it as a cheap, nontoxic cathode for low-power batteries.
LessonPitched as a low-power niche material with a transport bottleneck, it became the cheapest, safest mainstream cathode once conductive coatings and smaller particles fixed its kinetics. NOTABLE as published. The trajectory signals were abundant elements, a stable framework, and a limitation that looked like engineering rather than physics.
Leverage 4 · Magnitude 2 · Evidence 3 · Novelty 3 · Trajectory 3 · Stakes 4
Climate · 2009 · Landmark
merit 5.1Verdict: RoutineOrganometal halide perovskites as visible-light sensitizers for photovoltaic cells
Used methylammonium lead halide perovskite nanocrystals as light absorbers on porous titania in liquid-electrolyte cells, reaching 3.8% conversion efficiency and a photovoltage of 0.96 volts.
Lesson3.8% in an unstable liquid cell: the rubric rates it ROUTINE as published, and the field largely ignored it for three years. The missed signals were a new absorber class, strong absorption, an unusually high voltage, and simple solution processing. Within a decade perovskite cells passed 25%.
Leverage 4 · Magnitude 2 · Evidence 2 · Novelty 3 · Trajectory 4 · Stakes 4
Climate · 2012 · Landmark
merit 7.2Verdict: MajorEfficient hybrid solar cells based on meso-superstructured organometal halide perovskites
Built a solid-state perovskite cell on an insulating alumina scaffold and reached 10.9% power conversion efficiency with open-circuit voltages above 1.1 volts, showing that the perovskite itself carries the electrons rather than merely sensitizing titania.
LessonThe moment a curiosity became a field: a near-tripling of efficiency, a solid-state device, and a mechanistic surprise (the absorber transports charge) that explained why it would keep improving. Low fundamental energy losses were the trajectory signal.
Leverage 4 · Magnitude 3 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 4
Climate · 2024 · Landmark
merit 6.7Verdict: NotableAchievement of Target Gain Larger than Unity in an Inertial Fusion Experiment
On 5 December 2022 a laser-driven implosion at the National Ignition Facility produced 3.1 MJ of fusion energy from 2.05 MJ of laser light on target, a target gain of 1.5 and the first laboratory fusion experiment past scientific breakeven; later repeats reached gains of 1.9 and 1.0.
LessonStrong evidence of a genuine threshold, but the laser draws roughly 100 times more energy from the grid than it delivers to the target, and the facility was not designed for power. NOTABLE on merit: leverage and the path to electricity are limited. Separate a scientific milestone from a technology pathway and say which one a story is.
Leverage 2 · Magnitude 4 · Evidence 4 · Novelty 2 · Trajectory 3 · Stakes 4
Climate · 1967 · Landmark
merit 8.0Verdict: MajorThermal Equilibrium of the Atmosphere with a Given Distribution of Relative Humidity
A one-dimensional model of the atmosphere that holds relative humidity fixed, so water vapor rises as the air warms, found that doubling carbon dioxide warms the surface by about 2.3 °C while cooling the stratosphere.
LessonA deliberately simple model that captured the physics that matters (convection plus water-vapor feedback) gave a climate-sensitivity estimate still inside today's range. MAJOR as published; civilization-scale stakes and a result that aged well earned its place. Simple models that get the mechanism right beat complex ones that do not.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 5
Climate · 2023 · Landmark
merit 7.8Verdict: MajorLearning skillful medium-range global weather forecasting
A graph neural network trained on decades of reanalysis data produces 10-day global forecasts of hundreds of variables at 0.25° resolution in under a minute, outperforming ECMWF's operational deterministic system on 90% of 1,380 verification targets.
LessonBeating the best physics-based forecast on most targets at a tiny fraction of the compute, with code and weights released and independent operational testing, made machine-learning weather prediction mainstream within a year. Verification against the operational baseline, not a weaker model, was the evidence that mattered.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 4 · Stakes 4
Physics · 2016 · Landmark
merit 8.6Verdict: LandmarkObservation of Gravitational Waves from a Binary Black Hole Merger
Both LIGO detectors recorded the same signal on 14 September 2015 with a signal-to-noise ratio of 24 and a significance above 5.1 sigma: the merger of black holes of about 36 and 29 solar masses, which radiated three solar masses of energy as gravitational waves.
LessonCoincidence in two distant detectors, a waveform matching general relativity, and a false-alarm rate below one per 203,000 years: the evidence standard for a discovery. It opened a new way to observe the universe, which is why leverage and trajectory score at the top.
Leverage 5 · Magnitude 5 · Evidence 4 · Novelty 4 · Trajectory 5 · Stakes 2
Physics · 2004 · Landmark
merit 8.2Verdict: MajorElectric Field Effect in Atomically Thin Carbon Films
Isolated carbon films a few atoms thick by peeling graphite and showed they are stable in air, behave as a two-dimensional semimetal, and respond strongly to an applied gate voltage, with room-temperature mobilities around 10,000 cm²/Vs.
LessonAdhesive tape and a silicon wafer overturned the expectation that free-standing two-dimensional crystals could not exist, and launched the field of 2D materials. MAJOR as published because applications were speculative; a cheap, repeatable way to make a new class of material is itself a platform.
Leverage 5 · Magnitude 4 · Evidence 3 · Novelty 5 · Trajectory 4 · Stakes 3
Physics · 2018 · Landmark
merit 7.2Verdict: MajorUnconventional superconductivity in magic-angle graphene superlattices
Twisting two graphene sheets to about 1.1°, the first 'magic' angle, produced flat electronic bands, correlated insulating states, and, with gate doping, superconductivity up to 1.7 kelvin with a phase diagram resembling the cuprates.
LessonA tunable, all-carbon platform for strongly correlated physics, controlled by a twist angle and a gate voltage. The low critical temperature was not the story; the tunability and the cuprate-like phase diagram were. Score platforms for what they let others measure.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 2
Physics · 2019 · Landmark
merit 6.2Verdict: NotableQuantum supremacy using a programmable superconducting processor
Google's 53-qubit Sycamore processor sampled a random quantum circuit a million times in about 200 seconds; the authors estimated the same task would take a state-of-the-art supercomputer about 10,000 years.
LessonThe quantum hardware was real; the 10,000-year classical estimate was the weak point. IBM argued within days that a better classical method needed about 2.5 days, and later tensor-network simulations did it far faster. Beyond-classical claims are only as strong as the classical baseline they beat.
Leverage 3 · Magnitude 4 · Evidence 3 · Novelty 3 · Trajectory 3 · Stakes 2
Physics · 2025 · Landmark
merit 7.6Verdict: MajorQuantum error correction below the surface code threshold
On Google's Willow processors, surface-code memories showed logical errors falling by a factor of 2.14 each time the code distance grew by two, reaching 0.143% error per cycle with 101 qubits and outliving the best physical qubit by 2.4 times; real-time decoding kept up at distance 5.
LessonExponential error suppression with code size is the prerequisite for useful quantum computers, shown here for the first time in a surface code. Strong evidence (several distances, two processors, open data); the caveat worth printing is the rare correlated error burst that sets a floor.
Leverage 4 · Magnitude 4 · Evidence 4 · Novelty 3 · Trajectory 5 · Stakes 3
Physics · 2012 · Landmark
merit 6.4Verdict: NotableObservation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC
Combining 2011 and 2012 LHC data, ATLAS observed a neutral boson with a mass of 126.0 ± 0.4 (stat) ± 0.4 (sys) GeV at 5.9 standard deviations, consistent with the Standard Model Higgs boson; CMS reported an independent observation at the same time.
LessonThe cleanest evidence in the canon (two independent experiments, blinded analyses, five-sigma thresholds), yet NOTABLE on merit: it confirmed a 48-year-old prediction and changes no one's life directly. The rubric rewards what changes; a confirmation this important still leads on attention and on the news desk.
Leverage 2 · Magnitude 5 · Evidence 5 · Novelty 2 · Trajectory 2 · Stakes 1
Physics · 2024 · Landmark
merit 7.4Verdict: MajorLogical quantum processor based on reconfigurable atom arrays
A neutral-atom processor with up to 280 physical qubits ran algorithms on encoded logical qubits: it improved a logical two-qubit gate by scaling surface-code distance from 3 to 7 and entangled up to 48 logical qubits in error-detected circuits that beat the matching physical-qubit fidelities.
LessonMoving atoms instead of wiring them gave arbitrary connectivity and parallel logical operations, a genuinely different architecture. Much of the headline performance used error detection with post-selection rather than full correction; that is the caveat to carry.
Leverage 4 · Magnitude 4 · Evidence 3 · Novelty 4 · Trajectory 4 · Stakes 3
AI · 2023 · Cautionary
merit 2.4Verdict: SkipExploring the MIT Mathematics and EECS Curriculum Using Large Language Models
Claimed that GPT-4 with prompt engineering solved every question in a test set drawn from 4,550 MIT mathematics and EECS problems. Within days, three MIT students found unsolvable and duplicated questions and a grading loop in which GPT-4, holding the reference answers, decided when to stop retrying.
LessonA perfect score is a red flag, not a headline. Ask who grades (here the model graded itself with the answers in hand), whether retries stop on success, and whether test items leak into prompts. Extraordinary benchmark claims with self-grading earn evidence 1.
Leverage 1 · Magnitude 4 · Evidence 1 · Novelty 2 · Trajectory 2 · Stakes 3
AI · 2024 · Cautionary
merit 4.0Verdict: RoutineArtificial Intelligence, Scientific Discovery, and Product Innovation
Reported that randomly rolling out an AI materials-discovery tool to 1,018 scientists at an unnamed U.S. firm raised materials discoveries by 44%, patent filings by 39%, and product innovation by 17%. In May 2025 MIT said it had no confidence in the provenance, reliability, or validity of the data and asked arXiv to withdraw the paper.
LessonA clean causal design is only as good as data someone else can check. One junior author, one unnamed firm, no data access, and effect sizes unusually crisp for economics should cap evidence at 2 despite the randomized framing. Verifiability is part of evidence.
Leverage 2 · Magnitude 3 · Evidence 2 · Novelty 3 · Trajectory 2 · Stakes 3
Biology · 2011 · Cautionary
merit 4.1Verdict: RoutineA bacterium that can grow by using arsenic instead of phosphorus
Claimed that a bacterium from California's Mono Lake could substitute arsenic for phosphorus in its DNA and proteins. Two independent studies in 2012 found it still depends on phosphorus and does not build arsenic into its DNA, and Science retracted the paper in July 2025.
LessonAn extraordinary claim (a new chemistry of life) rested on ordinary proof and was launched with a NASA press conference. Evidence 1, whatever the venue. A press conference before scrutiny is a red flag, and 'published in Science' is not an evidence score.
Leverage 4 · Magnitude 5 · Evidence 1 · Novelty 5 · Trajectory 3 · Stakes 3
Biology · 2014 · Cautionary
merit 4.2Verdict: RoutineStimulus-triggered fate conversion of somatic cells into pluripotency
Claimed that a simple stress such as a brief acid bath could turn mature mouse cells into stem cells. Other labs could not reproduce it, an institutional investigation found fabricated and manipulated data, and Nature retracted the paper within six months.
LessonA claim that a trivially simple stimulus overturns decades of stem-cell work needs independent replication before it leads anything. Famous co-authors and a top journal are not evidence; watch for results no one outside the lab can repeat.
Leverage 4 · Magnitude 5 · Evidence 1 · Novelty 5 · Trajectory 3 · Stakes 4
Medicine · 2020 · Cautionary
merit 2.6Verdict: SkipHydroxychloroquine or chloroquine with or without a macrolide for treatment of COVID-19: a multinational registry analysis
Reported that among 96,032 hospitalized Covid-19 patients from 671 hospitals on six continents, hydroxychloroquine or chloroquine was associated with higher in-hospital mortality and heart-rhythm problems. The data came from Surgisphere, a small company that would not let auditors see them, and the paper was retracted about two weeks after publication.
LessonImplausibly large, fast, and tidy data from an unknown vendor (its Australian death counts exceeded the national total) are an evidence problem no statistical adjustment fixes. Registry size is not evidence when provenance cannot be checked, and an observational analysis cannot settle a question randomized trials were already testing.
Leverage 1 · Magnitude 3 · Evidence 1 · Novelty 1 · Trajectory 1 · Stakes 4
Medicine · 1998 · Cautionary
merit 2.9Verdict: SkipIleal-lymphoid-nodular hyperplasia, non-specific colitis, and pervasive developmental disorder in children
Described 12 children with bowel inflammation and developmental regression, in eight of whom parents linked the onset of symptoms to the MMR vaccine. Large studies found no link, the data were later shown to be falsified, and The Lancet fully retracted the paper in 2010.
LessonTwelve selected children, no controls, and timing reported by parents cannot support a causal claim about a vaccine given to millions; undisclosed conflicts compounded it. A case series is evidence 1, and a press conference that turns one into a public-health claim is a red flag.
Leverage 1 · Magnitude 3 · Evidence 1 · Novelty 4 · Trajectory 1 · Stakes 4
Climate · 1989 · Cautionary
merit 4.3Verdict: RoutineElectrochemically induced nuclear fusion of deuterium
Reported excess heat from palladium electrodes in heavy water at room temperature and attributed it to deuterium fusion inside the metal. Laboratories worldwide failed to reproduce it, and the neutron evidence was corrected in an erratum.
LessonAnnounced at a press conference before peer review, with calorimetry that lacked proper controls and nuclear signatures far too weak for the claimed heat. An extraordinary claim with ordinary proof: evidence 1. The mismatch between heat and radiation was the tell, and any physicist could check it.
Leverage 4 · Magnitude 5 · Evidence 1 · Novelty 5 · Trajectory 3 · Stakes 5
Physics · 2023 · Cautionary
merit 4.6Verdict: RoutineThe First Room-Temperature Ambient-Pressure Superconductor
Claimed that a copper-doped lead apatite, LK-99, superconducts above 400 K at ambient pressure. Replications within weeks found no superconductivity; the sharp resistance drop was traced to a copper sulfide impurity and the partial levitation to ordinary magnetism.
Lesson'For the first time in the world' and 'proved' in an abstract, a claim that would upend a century of physics, levitation videos, and no independent data. Without the evidence gate it would rank MAJOR (8.0); with it, 4.6. Viral attention on a preprint is not evidence.
Leverage 5 · Magnitude 5 · Evidence 1 · Novelty 5 · Trajectory 4 · Stakes 5
Physics · 2011 · Cautionary
merit 3.6Verdict: RoutineMeasurement of the neutrino velocity with the OPERA detector in the CNGS beam
Reported that neutrinos sent from CERN to Gran Sasso, 730 km away, arrived 60.7 nanoseconds earlier than light would, a 6-sigma anomaly implying faster-than-light travel. In 2012 the effect was traced to a badly connected fiber-optic cable and a clock fault, and the corrected result agreed with the speed of light.
LessonSix sigma measures statistical noise, not unknown systematic errors. The collaboration asked for scrutiny rather than claiming a discovery, and an editor should do the same: a result that contradicts special relativity and supernova neutrino timing needs an independent experiment before anything else.
Leverage 3 · Magnitude 5 · Evidence 1 · Novelty 5 · Trajectory 1 · Stakes 2
Physics · 2020 · Cautionary
merit 3.7Verdict: RoutineRoom-temperature superconductivity in a carbonaceous sulfur hydride
Claimed superconductivity at up to 287.7 K (about 15 °C) in a carbon-sulfur-hydrogen compound squeezed to 267 GPa in a diamond anvil cell. Nature retracted it in 2022 over a nonstandard background subtraction in the magnetic data, and a university investigation later found research misconduct.
LessonRoom-temperature superconductivity at 2.7 million atmospheres was never going to be useful, and the magnetic evidence rested on a background subtraction no one could reproduce from raw data. When raw data are withheld for an extraordinary result, cap evidence at 1.
Leverage 3 · Magnitude 5 · Evidence 1 · Novelty 4 · Trajectory 3 · Stakes 3