AIarXiv

Heuristic editor, no API keyVerdict: Notable

Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies

Answering questions about long videos often requires connecting events involving the same objects across hours or days.

By Ren, Fan, Pao +5

Score██████░░░░5.9

Key numbers

  • 72.0% accuracy

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: different objects may share a description, while observations of the same object remain disconnected across events. Retrieving relevant events therefore does not necessarily recover the "biography" of the particular entity a question concerns. To address this, we introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies while preserving the context of each moment. During question answering, the biography is retrieved alongside episodic evidence, allowing the model to follow an entity through events using identity links established during memory construction. Evaluations across four benchmarks, including day-long and week-long recordings, demonstrate improvements over prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA, GEB achieves 72.0% accuracy, 4.4 percentage points above the best published result. Ablations show that grounded identity association and biography reading both contribute to the gains, which additional descriptions alone do not fully recover.

Hui Ren, Lei Fan, Henry Pao, Han Guo, Zeeshan Zia, Ying Chen, Alexander Schwing, Gang Hua

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (state of the art); verification (multiple benchmarks, ablations).

How the score was computed

rank-2026-09-29

Score██████░░░░5.9

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.6 / 10
Shrunk toward the desk prior by editor confidence (40%).
Attention
86%
Citations, upvotes, points, mentions.
Freshness
69%
Half-life decay since publication.
  • Hugging Face upvotes72 (reference 25, via hf-daily, Oct 1, 2026, 02:16 UTC)
  • GitHub stars43 (reference 250, via hf-daily, Oct 1, 2026, 02:16 UTC)

The record

  • Reviewed by heuristic-v2 on Sep 30, 2026, 11:05 UTC. Paper type: method.
  • Categories: cs.CV, cs.AI, cs.CL, cs.IR, cs.LG
  • TOP, No.2 in the Artificial Intelligence edition of October 1, 2026.
  • TOP, No.4 in the Artificial Intelligence edition of September 30, 2026.