BiologybioRxiv
Heuristic editor, no API keyVerdict: NotableCerberus: bidirectional state space blocks improve accuracy and efficiency of regulatory sequence models
Sequence-to-function models predict regulatory activity and variant effects from DNA sequence alone, and leading models such as Borzoi compute long-range interactions with self-attention.
Key numbers
- 5.7-fold faster on 8
Caveats
- Preprint; not yet peer reviewed.
VerdictWorth a reader's time today.
Abstract
Sequence-to-function models predict regulatory activity and variant effects from DNA sequence alone, and leading models such as Borzoi compute long-range interactions with self-attention. The reference genome caps unique training sequences, and new assays add labels, not sequences. This constraint favors architectures that learn more from limited sequence diversity. We compared long-range blocks with the rest of the architecture and the training data held fixed, replacing Borzoi's transformer blocks with Hydra, a bidirectional state space block built on Mamba-2. Hydra blocks ran 5.7-fold faster on 8,192-token inputs, and models built on them predicted binned coverage on held-out sequences and classified fine-mapped eQTLs more accurately. Hydra's linear cost in sequence length let us run the blocks at 32 bp resolution and remove the U-net decoder, leaving a convolution-SSM model smaller and more accurate than the one it replaced. Scaling this architecture up on an augmented track collection produced Cerberus, an ensemble of eight models over 786 kb of input. We evaluated it on GTEx v11 fine-mapped eQTL, sQTL, and paQTL benchmarks with negatives matched on allele frequency, gene expression, and phenotype-specific positional annotations to reduce confounding by these properties. Cerberus is more accurate than Borzoi on all three phenotypes, 0.692 versus 0.668 mean eQTL AUPRC across 48 tissues, and estimates eQTL effect sizes better, 0.379 versus 0.321 Spearman ρ. Against AlphaGenome, it trades leads in eQTL classification, ahead on variants 3 to 100 kb from the transcription start site and behind in coding sequence and within 3 kb. Interpretation of the trained blocks shows interaction range growing with block depth, from a 0.9 kb median half-life in the first block to 79 kb in the seventh, and the deepest heads anchoring on promoters, enhancers, and CTCF peaks. We release the models, the training data, the training and evaluation code, and the benchmark sets, supporting variant scoring and transfer learning.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 16% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 20% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 20% | A new combination of known ideas. |
| Trajectory | ███░░ 3 | 10% | A clear path to scale. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: gains (x-fold, versus baseline); verification (multiple benchmarks, held-out test, code released); scale (scalable, efficient).
How the score was computed
- Merit
- 5.4 / 10
- Adjusted merit
- 4.6 / 10
- Attention
- 0%
- Freshness
- 88%