BiologybioRxiv
Heuristic editor, no API keyVerdict: NotableLLM-powered automatic scoring tools for memory recall
Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it.
Caveats
- Preprint; not yet peer reviewed.
VerdictWorth a reader's time today.
Abstract
Free recall of naturalistic stimuli, such as films and stories, reveals what people remember and how they organize it. Scoring such recall requires matching each recalled utterance to the encoded stimulus, and doing this manually often makes large studies impractical. Existing automated methods mostly return similarity or aggregate scores rather than explicit matches between recalled and stimulus content. Here we introduce an open-source, model-agnostic pipeline that uses a large language model (LLM) to divide the stimulus and the recall into information units and to match units between the two texts. The output is a table of unit-to-unit matches with accuracy labels, from which standard analysis scripts compute recall measures. We validated the pipeline with four LLMs (Claude Sonnet 4.5, GPT 5.4, Gemini 2.5 Pro, and Gemini 2.5 Flash) on three datasets. On typed recall of four short stories, agreement between each model and two trained raters ({phi} = .73-.79) was close to the agreement between the raters themselves ({phi} = .77). A second run reproduced each model's matches ({phi} = .91-.97) more closely than the raters agreed with each other. On spoken recall of 10 films, every model attributed recall to the correct film with high agreement. On uncorrected speech-to-text recall of a television episode, agreement varied more across models. The median cost was under $1 per participant for every model. These results suggest that the pipeline can serve as a reliable, low-cost alternative to manual scoring of free recall.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ████░ 4 | 24% | A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing). |
| Magnitude | ██░░░ 2 | 16% | Solid incremental gain on a meaningful problem. |
| Evidence | ███░░ 3 | 20% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ███░░ 3 | 20% | A genuinely new approach to an open problem. |
| Trajectory | ███░░ 3 | 10% | A clear path to scale. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); novelty (alternative to status quo); verification (multiple benchmarks, experimental validation, independent replication); scale (low cost).
How the score was computed
- Merit
- 6.0 / 10
- Adjusted merit
- 4.9 / 10
- Attention
- 0%
- Freshness
- 88%