AIarXiv
Heuristic editor, no API keyVerdict: NotableOneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available.
Key numbers
- 27.5% of annotated state tokens
VerdictWorth a reader's time today.
Abstract
Streaming video LLMs must retain evidence before its relevance to future tasks is known and respond when sufficient evidence becomes available. The challenge is to form reusable factual memory without compromising real-time perception. We introduce OneStreamer, which jointly learns query-independent evidence recording and task response through a shared proactive generation process. Its Proactive Hierarchical Caption Memory (PHCM) produces time-grounded local-detail captions and summaries of completed events. Streaming caption targets supervise the interpretation of observed video prefixes during training. At inference, model-generated records complement a recent visual window, providing reusable factual context without revisiting historical visual features. Proactive State Transition Learning (PSTL) reduces the dominance of repeated waiting states by preserving supervision at all output anchors and selecting representative state-change and state-persistence tokens. We further develop a streaming data synthesis pipeline that aligns output content and timing with available evidence. Combining the resulting streaming captions and QA with cleaned open-source data yields OneStreamer-1M, a broad-coverage streaming video interaction dataset with over one million records spanning diverse tasks. Our 4B model achieves the best results among the compared methods across all eight evaluated streaming video understanding benchmarks. Ablations show that retaining generated captions improves historical QA without degrading real-time perception. PSTL also outperforms dense state supervision while supervising only 27.5% of annotated state tokens. Together, these results support proactive generation as a shared learning interface connecting perception, memory formation, and timely response in streaming video interaction.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ████░ 4 | 24% | A general-purpose tool used across several fields (Adam, ResNet, LoRA, next-generation sequencing). |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); breadth (many tasks); gains (outperforms); verification (ablations, independent replication, code released); stakes (global scale).
How the score was computed
- Merit
- 5.6 / 10
- Adjusted merit
- 4.7 / 10
- Attention
- 88%
- Freshness
- 76%
- Hugging Face upvotes96 (reference 25, via hf-daily, Oct 2, 2026, 13:49 UTC)
- GitHub stars6 (reference 250, via hf-daily, Oct 2, 2026, 13:49 UTC)