Tim Salimans, Ben Poole, Mohammad Norouzi, Diederik Kingma, David Fleet, Alexey Gritsenko, Ruiqi Gao, Jonathan Ho, William Chan, Jay Whang, Chitwan Saharia
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. If it links a repository, it is listed below.
We present Imagen Video, a text-conditional video generation system based on a cascade of video diffusion models. Given a text prompt, Imagen Video generates high definition videos using a base video generation model and a sequence of interleaved spatial and temporal video super-resolution models. We describe how we scale up the system as a high definition text-to-video model including design decisions such as the choice of fully-convolutional temporal and spatial superresolution models at certain resolutions, and the choice of the v-parameterization of diffusion models. In addition, we confirm and transfer findings from previous work on diffusion-based image generation to the video generation setting. Finally, we apply progressive distillation to our video models with classifier-free guidance for fast, high quality sampling. We find Imagen Video not only capable of generating videos of high fidelity, but also having a high degree of controllability and world knowledge, including the ability to generate diverse videos and text animations in various artistic styles and with 3D object understanding. See imagen.research.google/video for samples.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2210.02303")
get_code_for_paper("2210.02303")
have("2210.02303")
Connect an agent — have() is free.