AIarXiv

Heuristic editor, no API keyVerdict: Notable

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available.

By Zeng, Yang, Zhang +21

Score██████░░░░6.0

Key numbers

  • 27.5% of annotated state tokens

VerdictWorth a reader's time today.

Read the originalPDFCode

Abstract

Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.

Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang

The editor's rubric

Heuristic review

DimensionLevelWeightWhat that level means
Leverage████░ 424%A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing).
Magnitude███░░ 318%Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem.
Evidence███░░ 314%Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data.
Novelty██░░░ 216%A new combination of known ideas.
Trajectory██░░░ 218%Some room to improve with obvious engineering.
Stakes██░░░ 210%Benefits a professional community (practitioners, clinicians, engineers).

Editor’s rationale

Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (outperforms); verification (ablations, independent replication, code released); stakes (global scale).

How the score was computed

rank-2026-09-29

Score██████░░░░6.0

Score = 10 × (65% × adjusted merit / 10 + 25% × attention + 10% × freshness)

Merit
5.6 / 10
Weighted rubric, evidence-gated.
Adjusted merit
4.7 / 10
Shrunk toward the desk prior by editor confidence (42%).
Attention
88%
Citations, upvotes, points, mentions.
Freshness
76%
Half-life decay since publication.
  • Hugging Face upvotes96 (reference 25, via hf-daily, Oct 2, 2026, 13:49 UTC)
  • GitHub stars6 (reference 250, via hf-daily, Oct 2, 2026, 13:49 UTC)

The record

  • Reviewed by heuristic-v2 on Oct 2, 2026, 02:05 UTC. Paper type: method.
  • Categories: cs.CV
  • LEAD, No.1 in the Artificial Intelligence edition of October 2, 2026.