BiologybioRxiv

Heuristic editor, no API keyVerdict: Notable

EpiZoo: a DNA sequence-aware foundation model for cross-species single-cell epigenomics

Large-scale single-cell epigenomic atlases characterize chromatin regulatory landscapes across diverse biological contexts, spanning cell types, tissues, individuals and species.

By Li, Chen, Jiang +3

Score████░░░░░░4.4

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Large-scale single-cell epigenomic atlases characterize chromatin regulatory landscapes across diverse biological contexts, spanning cell types, tissues, individuals and species. Foundation models provide an opportunity to capture the full spectrum of cellular diversity in these atlases, yet current models remain largely confined to individual species by genomic coordinate dependence and overlook regulatory information encoded in DNA sequences. Here we introduce EpiZoo, a DNA sequence-aware foundation model for cross-species single-cell epigenomics. EpiZoo converts million-dimensional single-cell epigenomic profiles from diverse species into compact cell sentences that integrate DNA-encoded regulatory information, sequence-independent epigenomic context and accessibility-based importance. Built around a mixture-of-experts transformer and containing 2.6 billion parameters in total, EpiZoo is pretrained on our manually curated multi-species Omni-scATAC corpus of approximately 20.9 million cells to learn regulatory programs across species. On external datasets excluded from pretraining, EpiZoo achieves state-of-the-art performance in fundamental single-cell analysis tasks, including feature extraction, cell type annotation and data imputation. Its sequence-aware architecture enables extension to evolutionarily diverse species, and supports comparative analysis of regulatory conservation and divergence during primate evolution. Benefiting from this modeling design, EpiZoo enables context-aware prioritization of somatic mutations in cancer and prediction of cell-type-specific chromatin accessibility from DNA sequences across genomic regions and species.

K. Li, X. Chen, Q. Jiang, Z. Wang, H. Lv, R. Jiang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage███░░ 324%A method or resource many groups across the field will adopt within a year.
Magnitude███░░ 316%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 320%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 220%A new combination of known ideas.
Trajectory███░░ 310%A clear path to scale.
Stakes███░░ 310%Meaningful benefit to many people within a few years.

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (state of the art); verification (independent replication); scale (scalable, larger models); stakes (global scale, major disease). Red flags: derivative (comparative study).

How the score was computed

rank-2026-09-29

Score████░░░░░░4.4

Score = 10 × (75% × adjusted merit / 10 + 15% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.7 / 10
Shrunk toward the desk prior by editor confidence (46%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
87%
Half-life decay since publication.

No attention signals recorded yet.

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • Categories: bioinformatics
  • BRIEF, No.6 in the Biology edition of September 30, 2026.
  • BRIEF, No.3 in the Biology edition of September 29, 2026.