BiologybioRxiv

Heuristic editor, no API keyVerdict: Notable

LLM-powered automatic scoring tools for memory recall

Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it.

By Tang, Born, Delarazan +3

Score█████░░░░░4.5

Caveats

  • Preprint; not yet peer reviewed.

VerdictWorth a reader's time today.

Read the originalPDF

Abstract

Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it. Scoring such recall requires matching each recalled utterance to the encoded stimulus, and doing this manually often makes large studies impractical. Existing automated methods mostly return similarity or aggregate scores rather than explicit matches between recalled and stimulus content. Here we introduce an open-source, model-agnostic pipeline that uses a large language model (LLM) to divide the stimulus and the recall into information units and to match units between the two texts. The output is a table of unit-to-unit matches with accuracy labels, from which standard analysis scripts compute recall measures. We validated the pipeline with four LLMs (Claude Sonnet 4.5, GPT 5.4, Gemini 2.5 Pro, and Gemini 2.5 Flash) on three datasets. On typed recall of four short stories, agreement between each model and two trained raters ({phi} = .73-.79) was close to the agreement between the raters themselves ({phi} = .77). A second run reproduced each model's matches ({phi} = .91-.97) more closely than the raters agreed with each other. On spoken recall of 10 films, every model attributed recall to the correct film with high agreement. On uncorrected speech-to-text recall of a television episode, agreement varied more across models. The median cost was under $1 per participant for every model. These results suggest that the pipeline can serve as a reliable, low-cost alternative to manual scoring of free recall.

R. S. Tang, S. J. Born, A. I. Delarazan, P. L. Demarest, A. B. Karagoz, Z. M. Reagh

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude██░░░ 216%Solid incremental gain on a meaningful problem.
Evidence███░░ 320%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty███░░ 320%A genuinely new approach to an open problem.
Trajectory███░░ 310%A clear path to scale.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); novelty (alternative to status quo); verification (multiple benchmarks, experimental validation, independent replication); scale (low cost).

How the score was computed

rank-2026-10-07

Score█████░░░░░4.5

Score = 10 × (75% × adjusted merit / 10 + 15% × attention + 10% × freshness)

Merit
6.0 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.9 / 10
Shrunk toward the desk prior by editor confidence (44%).
Attention
0%
Citations, upvotes, points, mentions.
Freshness
88%
Half-life decay since publication.

No attention signals recorded yet.

The record

  • Reviewed by heuristic-v5 on Oct 9, 2026, 05:48 UTC. Paper type: method.
  • Categories: neuroscience
  • TOP, No.3 in the Front page edition of October 9, 2026.
  • TOP, No.1 in the Biology edition of October 9, 2026.