BiologybioRxiv
Heuristic editor, no API keyVerdict: NotableEvidence Scaling for Zero-Shot Protein Reasoning with Large Language Models
Large language models (LLMs) show emerging zero-shot capability for protein variant prediction, yet still lag behind specialized protein models.
VerdictWorth a reader's time today.
Abstract
Large language models (LLMs) show emerging zero-shot capability for protein variant prediction, yet still lag behind specialized protein models. We ask whether this gap can be reduced by scaling access to biological evidence rather than adapting model parameters. We introduce BioEvidence, a training-free and model-agnostic interface that converts structural and evolutionary information from standard biological tools into compact evidence for frozen LLMs. On the ProteinGym benchmark, we observe evidence scaling: performance improves as evidence becomes richer. Structural and evolutionary evidence each improve performance, and combining them yields further gains, while mismatching the same evidence to the wrong variants degrades performance below the no-evidence baseline. Notably, BioEvidence enables zero-shot ranking to reach strong specialized protein predictors on matched evaluations, and the improvement persists on post-cutoff data released after the model's knowledge cutoff. Evidence also interacts with conventional scaling: for GPT-5.6 Sol, evidence at low reasoning effort outperforms the no-evidence condition at medium effort, while a six-model analysis associates stronger no-evidence performance with larger margins over evolutionary rank fusion. These results identify external evidence as a complementary scaling axis for scientific prediction alongside model capability and inference effort.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ████░ 4 | 24% | A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing). |
| Magnitude | ███░░ 3 | 16% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 20% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ███░░ 3 | 20% | A genuinely new approach to an open problem. |
| Trajectory | ███░░ 3 | 10% | A clear path to scale. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (zero/few-shot); gains (outperforms); novelty (alternative to status quo, discovery); verification (code released); scale (scalable, improves with scale).
How the score was computed
- Merit
- 6.3 / 10
- Adjusted merit
- 5.0 / 10
- Attention
- 0%
- Freshness
- 88%