Search arXiv⌕ Search

arXiv · 2610.05608

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Team Kandinsky·Julia Agafonova·Bulat Akhmatov·Mikhail Aksyutin·Grigorii Alekseenko·Anastasia Aliaskina·Olga Androsova·Vladimir Arkhipkin·Anna Averchenkova·Alexander Belykh·Serafima Bocharova·Sofiya Bogakovskaya·Anton Bukashkin·Mark Bulygin·Kirill Buzygin·Irina Cheremnykh·Kirill Chernyshev·Mikhail Chernyshov·Vladimir Chernyy·David Chikovani·Georgy Daniltsev·Denis Dimitrov·Anna Dmitrienko·Vladimir Dokholyan·Sergey Emelyanov·Dmitry Ermilov·Georgii Fedorov·Polina Gavrilova·Nikolai Gerasimenko·Aleksandr Gordeev·Andrey Inozemtsev·Andrei Ivaniuta·Alexander Ivanov·Mikhail Karaev·Anastasiia Kargapoltseva·Ivan Kirillov·Nikita Kiselev·Valeria Kobenko·Yury Kolabushin·Denis Koposov·Anatoly Korobov·Vladimir Korviakov·Kirill Kozlov·Denis Krzhivokolskiy·Konstantin Kuklev·Alexander Kunitsyn·Sergey Kuzin·Vladislav Lakhtionov·Alexey Letunovskiy·Maxim Litvinov·Alexander Lyulkov·Georgy Makarov·Kirill Malakhov·Egor Malykh·Mikhail Mamaev·Dmitrii Mikhailov·Polina Mikhailova·Ivan Mikheev·Elizaveta Muromtseva·Nikolai Nazarkin·Tatiana Nikulina·Lev Novitskiy·Stanislav Onuchin·Nikita Osterov·Denis Parkhomenko·Anatoliy Parpara·Vladimir Polovnikov·Konstantin Reznikov·Azat Saginbaev·Nikita Samsonov·Alexander Sentsov·Nikita Shaimov·Artem Sherstyuk·Andrey Shutkin·Egor Silvestrov·Bulat Suleimanov·Matvey Suprunov·Sergey Taranov·Irina Tolstykh·Tatiana Trofimuk·Ilya Trushkin·Aleksandra Tsybina·Olga Varlashina·Viacheslav Vasilev·Ilya Vasiliev·Eugeny Vilisov·Sergey Yakubson·Konstantin Zakharov

Abstract

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin, Grigorii Alekseenko, Anastasia Aliaskina, Olga Androsova, Vladimir Arkhipkin, Anna Averchenkova, Alexander Belykh, Serafima Bocharova, Sofiya Bogakovskaya, Anton Bukashkin, Mark Bulygin, Kirill Buzygin, Irina Cheremnykh, Kirill Chernyshev, Mikhail Chernyshov, Vladimir Chernyy, David Chikovani, Georgy Daniltsev, Denis Dimitrov, Anna Dmitrienko, Vladimir Dokholyan, Sergey Emelyanov, Dmitry Ermilov, Georgii Fedorov, Polina Gavrilova, Nikolai Gerasimenko, Aleksandr Gordeev, Andrey Inozemtsev, Andrei Ivaniuta, Alexander Ivanov, Mikhail Karaev, Anastasiia Kargapoltseva, Ivan Kirillov, Nikita Kiselev, Valeria Kobenko, Yury Kolabushin, Denis Koposov, Anatoly Korobov, Vladimir Korviakov, Kirill Kozlov, Denis Krzhivokolskiy, Konstantin Kuklev, Alexander Kunitsyn, Sergey Kuzin, Vladislav Lakhtionov, Alexey Letunovskiy, Maxim Litvinov, Alexander Lyulkov, Georgy Makarov, Kirill Malakhov, Egor Malykh, Mikhail Mamaev, Dmitrii Mikhailov, Polina Mikhailova, Ivan Mikheev, Elizaveta Muromtseva, Nikolai Nazarkin, Tatiana Nikulina, Lev Novitskiy, Stanislav Onuchin, Nikita Osterov, Denis Parkhomenko, Anatoliy Parpara, Vladimir Polovnikov, Konstantin Reznikov, Azat Saginbaev, Nikita Samsonov, Alexander Sentsov, Nikita Shaimov, Artem Sherstyuk, Andrey Shutkin, Egor Silvestrov, Bulat Suleimanov, Matvey Suprunov, Sergey Taranov, Irina Tolstykh, Tatiana Trofimuk, Ilya Trushkin, Aleksandra Tsybina, Olga Varlashina, Viacheslav Vasilev, Ilya Vasiliev, Eugeny Vilisov, Sergey Yakubson, Konstantin Zakharov. 2026-10-04. Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation. https://arxiv.org/abs/2610.05608

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Large Pretraining Datasets Don't Guarantee Robustness after Fine-Tuning in Image Classification

Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of maintaining the general performance of the model while allowing it to gain new skills. A valuable goal for all such models is robustness: the ability to perform well on out-of-distribution (OOD) tasks. We assess whether fine-tuning preserves the overall robustness of the pretrained model in image classification, and observed that models pretrained on large datasets exhibited strong catastrophic forgetting and loss of OOD generalization. To systematically assess robustness preservation in fine-tuned models, we propose the Robustness Inheritance Benchmark (ImageNet-RIB). The benchmark, which can be applied to any pretrained model, consists of a set of related but distinct OOD (downstream) tasks and involves fine-tuning on one of the OOD tasks in the set then testing on the rest. We find that though continual learning methods help, fine-tuning reduces robustness across pretrained models. Surprisingly, models pretrained on the largest and most diverse datasets (e.g., LAION-2B) exhibit both larger robustness losses and lower absolute robustness after fine-tuning on small datasets, relative to models pretrained on smaller datasets. We observe this collapse in contrastively pretrained (CLIP) models and their fine-tuned variants, where it grows with pretraining scale; the supervised models we test do not exhibit it. These findings suggest that starting with the strongest foundation model is not necessarily the best approach for performance on specialist tasks. https://jd730.github.io/projects/ImageNet-RIB

cs.CV↗

Feature Space Analysis by Guided Diffusion Model

This paper aims to analyse the feature space of a vision-related Deep Neural Network (DNN) by proposing a decoder that can generate an image whose feature closely matches a user-specified feature. Supported by quantitative evidence of its high feature-matching accuracy, our decoder facilitates precise analysis of the DNN's feature space. Our decoder is implemented as a guided diffusion model that guides the image generation of a pre-trained diffusion model to minimise the Euclidean distance between the feature of a clean image estimated at each step and the user-specified feature. The key advantages of our decoder are its training-free applicability to analyse the feature spaces of different DNNs and its practical feasibility on a single COTS GPU. The experiments targeting CLIP's image encoder and ResNet-50 demonstrate the effectiveness of our decoder both as a feature-matching image generator and as a visual feature space analyser. The codes and data are available at https://github.com/ccilab-doshisha/FeatDec

cs.CV↗

Scaling Laws for Deepfake Detection

This paper presents a systematic study of scaling laws for the deepfake detection task. Specifically, we analyze the model performance against the number of real image domains, deepfake generation methods, and training images. Since no existing dataset meets the scale requirements for this research, we construct ScaleDF, the largest dataset to date in this field, which contains over 5.8 million real images from 51 different datasets (domains) and more than 8.8 million fake images generated by 102 deepfake methods. Using ScaleDF, we observe power-law scaling similar to that shown in large language models (LLMs). Specifically, the average detection error follows a predictable power-law decay as either the number of real domains or the number of deepfake methods increases. This key observation not only allows us to forecast the number of additional real domains or deepfake methods required to reach a target performance, but also inspires us to counter the evolving deepfake technology in a data-centric manner. Beyond this, we examine the role of pre-training and data augmentations in deepfake detection under scaling, as well as the limitations of scaling itself.The ScaleDF dataset is available at https://huggingface.co/datasets/WenhaoWang/ScaleDF.

cs.CV↗