BiologybioRxiv
Heuristic editor, no API keyVerdict: NotablePanVasc Research for AI assisted evidence analysis in panvascular intervention
Panvascular intervention research requires evidence workflows that preserve source identity, outcome definitions and observation windows.
VerdictWorth a reader's time today.
Abstract
Panvascular intervention research requires evidence workflows that preserve source identity, outcome definitions and observation windows. We developed PanVasc Research, an executable research framework, and evaluated a fixed local Qwen3-4B model using two complementary tasks. Fifty ClinicalTrials.gov records from five vascular query strata generated 200 source-fidelity tests with intact evidence or controlled removal of the requested source, primary outcome or timeframe, plus 50 clean controls. A separate 100-statement sample from the official NLI4CT test set assessed clinical-trial entailment and evidence selection against existing expert labels. Generic and checklist prompts used identical evidence and maximum generation budgets; each output was also evaluated with an input-only deterministic contract. Registry exact accuracy was 51/200 (25.5%) with the generic prompt and 77/200 (38.5%) with the checklist; intact-case agreement was 50/50 and 48/50. The rule baseline recovered 200/200 tuples. NLI label accuracy was 48/100 (48.0%) and 50/100 (50.0%), respectively. The source contract retained 35 and 31 incorrect NLI labels in the two arms. Registry references establish fidelity to a registration snapshot, while NLI4CT concerns breast-cancer trials and does not validate vascular expertise. The framework also retains provenance-recorded literature retrieval, structured research planning and local numerical analysis. These experiments support a bounded assessment of source handling and semantic failure, rather than a new foundation model or autonomous scientific discovery. A proposed endpoint representation identifies the additional domain annotation and independent validation required for a panvascular research model.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ██░░░ 2 | 16% | Solid incremental gain on a meaningful problem. |
| Evidence | ████░ 4 | 20% | Strong: large scale, preregistered, independently replicated, or a well-powered randomized trial. |
| Novelty | ███░░ 3 | 20% | A genuinely new approach to an open problem. |
| Trajectory | ██░░░ 2 | 10% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (new method); breadth (many tasks); novelty (alternative to status quo, discovery); design (registered); verification (multiple benchmarks, held-out test, independent replication); stakes (major disease).
How the score was computed
- Merit
- 5.7 / 10
- Adjusted merit
- 4.8 / 10
- Attention
- 0%
- Freshness
- 78%
- Citations0 (reference 20, via openalex, Sep 29, 2026, 23:53 UTC)