BiologybioRxiv
Heuristic editor, no API keyVerdict: Notablepre-miRBench: a multispecies benchmark and reference model for animal precursor microRNA prediction
Motivation: Precursor microRNA (pre-miRNA) prediction is used to prioritize genomic hairpins for miRNA annotation.
Key numbers
- 95% confidence interval of
VerdictWorth a reader's time today.
Abstract
Motivation: Precursor microRNA (pre-miRNA) prediction is used to prioritize genomic hairpins for miRNA annotation. Published predictors are trained and evaluated on different species collections, negative sets and data partitions, making reported gains difficult to compare. Results: We introduce pre-miRBench, a benchmark of 7,056 200-nt windows centered on curated pre-miRNA loci and 70,560 genomic negatives with predicted hairpin structures from 71 animal species, with a fixed 1:10 positive-to-negative ratio in every partition. Four test sets cross whether the species and miRNA family are represented in training: known species/known family, known species/held-out family, held-out species/known family, and held-out species/held-out family. We retrained six published pre-miRNA predictors (DeepMir, deepMiRGene, dnnPreMiR, mirDNN, miRe2e and MuStARD) under the same benchmark splits and evaluated them alongside the pre-miRBench model, an ensemble of three neural networks developed with Agentomics, an autonomous LLM-agent system for biomedical machine learning. The pre-miRBench model achieved the highest mean average precision (AP) across the four test sets (0.9809; dnnPreMiR, 0.9719) and led three tests. The two held-out species were Gallus gallus and Drosophila melanogaster, and the joint species-and-family holdout contained 69 positives and 690 negatives. dnnPreMiR had the highest AP on the joint species-and-family holdout (0.9677 versus 0.9659). The AP difference (pre-miRBench minus dnnPreMiR) was -0.0018, with a paired bootstrap 95% confidence interval of -0.0445 to 0.0341. pre-miRBench establishes a common evaluation for species and miRNA-family holdouts and provides a reproducible reference model for future comparisons. Availability and implementation: Reproducible source code is available at https://github.com/dimostzim/pre-miRBench. Datasets, trained model weights, record-level predictions, evaluation metrics and species-panel metadata for pre-miRBench and the six retrained predictors are available at https://zenodo.org/records/22813044.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ████░ 4 | 24% | A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing). |
| Magnitude | ███░░ 3 | 16% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ████░ 4 | 20% | Strong: large scale, preregistered, independently replicated, or a well-powered randomized trial. |
| Novelty | ██░░░ 2 | 20% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 10% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (versus baseline); verification (confidence interval, multiple benchmarks, held-out test). Red flags: derivative (comparative study).
How the score was computed
- Merit
- 6.1 / 10
- Adjusted merit
- 4.9 / 10
- Attention
- 0%
- Freshness
- 88%