Section

Artificial Intelligence

Models, methods, and the machinery of intelligence.

Daily at 01:15 and 13:15 UTC

Edition No. 10
Updated UTC

arXiv

Verdict: Routine

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters).

By Team Kandinsky, Agafonova, Akhmatov +85

  • HF ▲ 135
  • Stars ★ 188
Score██████░░░░5.6

Top stories

  1. 1. Recurrent Looped Transformer

    State tracking requires an update at every input, but the depth a Transformer applies to each token is fixed regardless of sequence length.

    ██████░░░░5.5
  2. 2. TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

    Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training.

    █████░░░░░5.4
  3. 3. Long-WAM: Scaling the Context of World-Action Models

    Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action.

    █████░░░░░5.4

More top stories

No.04 to No.08

  1. arXivby Lin, Cao, Luo +3

    From Evidence to Action: How Tool-Using Agents Fail

    Tool-using agents make consequential changes to external state, yet correct outcomes do not guarantee that their actions were supported by evidence established beforehand.

    No.08Verdict: NotableScore█████░░░░░5.2

Briefs

10 more from the desk