AIarXiv
Heuristic editor, no API keyVerdict: RoutineKandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters).
Key numbers
- 1920 x 1080
VerdictCompetent work. Briefs at most.
Abstract
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920×1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.
The editor's rubric
| Dimension | Level | Weight | What that level means |
|---|---|---|---|
| Leverage | ███░░ 3 | 24% | A method or resource many groups across the field will adopt within a year. |
| Magnitude | ███░░ 3 | 18% | Large gain: roughly 2x, or a clear new state of the art on a hard, unsaturated problem. |
| Evidence | ███░░ 3 | 14% | Solid: multiple benchmarks or cohorts, ablations, fair baselines, released code or data. |
| Novelty | ██░░░ 2 | 16% | A new combination of known ideas. |
| Trajectory | ██░░░ 2 | 18% | Some room to improve with obvious engineering. |
| Stakes | ██░░░ 2 | 10% | Benefits a professional community (practitioners, clinicians, engineers). |
Editor’s rationale
Heuristic triage from title and abstract text only, not a reading of the paper. Cues found: method (we propose); gains (outperforms); verification (human evaluation, code released); scale (scalable).
How the score was computed
- Merit
- 5.1 / 10
- Adjusted merit
- 4.4 / 10
- Attention
- 89%
- Freshness
- 64%
- Hugging Face upvotes103 (reference 25, via hf-daily, Oct 6, 2026, 13:49 UTC)
- GitHub stars1 (reference 250, via hf-daily, Oct 6, 2026, 13:49 UTC)