BiologybioRxiv

Heuristic editor, no API keyVerdict: Notable

Fitting dynamics is not identifying causal edges: a white-box masked ODE benchmark for trans-omics digital twins of drug action mechanisms

Multi-component drug regimens, with traditional Chinese medicine formulas as the hardest case, act across signalling, transcriptional, proteomic and metabolic layers, and elucidating their mechanisms requires dynamic…

By Zhang, Jia, Pan

Score████░░░░░░4.4

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Multi-component drug regimens, with traditional Chinese medicine formulas as the hardest case, act across signalling, transcriptional, proteomic and metabolic layers, and elucidating their mechanisms requires dynamic models that predict molecular trajectories rather than static association networks. Trainable ordinary differential equation (ODE) systems fitted to time-series omics are increasingly used for this purpose; yet their validation remains fit-based: at genomic scale, no ground truth has existed to test whether a good fit implies correct mechanisms. Here we build that ground truth: a white-box masked ODE benchmark on the real trans-omics topology of insulin action in mouse liver (transcriptome GEO GSE166336, proteome ProteomeXchange PXD022728, phosphoproteome PXD022823, metabolome source-publication Tables S2-S3; 2,106 molecular species; 4,912 ground-truth edges), with controllable noise, missingness and sampling budgets. Four instruments quantify identifiability: an oracle-perturbation basin curve, a held-out-layer corruption assay, a saturation audit, and an ideal-budget ceiling test. We find that a static baseline (FD + LASSO) performs at chance (AUROC ~ 0.50, except GE at 0.567); that from-scratch training remains at chance even with noise-free, fully observed, densely sampled data (per-layer AUROC 0.48-0.53); that held out layers act as corruption sinks whose failure decomposes into an information floor, a scale-mismatch amplifier, and an edge-gradient drag, curable only jointly; that tanh saturation silently zeroes entire regulator columns; and that a 12-knockout validation battery decomposes intervention reliability by network distance. We distill these into operational prescriptions. The binding constraint is not fitting but structural identifiability. Benchmark, code and audit tools are planned for open release upon publication.

Z. Zhang, L. Jia, Y. Pan

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 216%Solid incremental gain on a meaningful problem.
Evidence███░░ 320%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 320%A genuinely new approach to an open problem.
Trajectory███░░ 310%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we report); breadth (many tasks); novelty (alternative to status quo); verification (multiple benchmarks, held-out test, code released); scale (scalable).

How the score was computed

rank-2026-09-29

Score████░░░░░░4.4

Score = 10 × (75% × adjusted merit / 10 + 15% × attention + 10% × freshness)

Merit
6.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.8 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
78%
Half-life decay since publication.
  • Citations0 (reference 20, via openalex, Sep 29, 2026, 23:37 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 29, 2026, 23:53 UTC. Paper type: method.
  • Categories: bioinformatics
  • BRIEF, No.7 in the Biology edition of September 29, 2026.