STIV: Scalable Textual content and Picture Conditioned Video Era

The sector of video technology has made exceptional developments, but there stays a urgent want for a transparent, systematic recipe that may information the event of sturdy and scalable fashions. On this work, we current a complete research that systematically explores the interaction of mannequin architectures, coaching recipes, and knowledge curation methods, culminating in a easy and scalable text-image-conditioned video technology technique, named STIV. Our framework integrates picture situation right into a Diffusion Transformer (DiT) by body alternative, whereas incorporating textual content conditioning through a joint image-text conditional classifier-free steerage. This design allows STIV to carry out each text-to-video (T2V) and text-image-to-video (TI2V) duties concurrently. Moreover, STIV might be simply prolonged to numerous purposes, comparable to video prediction, body interpolation, multi-view technology, and lengthy video technology, and many others. With complete ablation research on T2I, T2V, and TI2V, STIV display sturdy efficiency, regardless of its easy design. An 8.7B mannequin with 512 decision achieves 83.1 on VBench T2V, surpassing each main open and closed-source fashions like CogVideoX-5B, Pika, Kling, and Gen-3. The identical-sized mannequin additionally achieves a state-of-the-art results of 90.1 on VBench I2V job at 512 decision. By offering a clear and extensible recipe for constructing cutting-edge video technology fashions, we goal to empower future analysis and speed up progress towards extra versatile and dependable video technology options.

† College of California, Los Angeles
** Work performed whereas at Apple

Main Menu

What's Hot

5 AI Buying and selling Bots That Work With Robinhood

Everest Ransomware Claims Mailchimp as New Sufferer in Comparatively Small Breach

VMware Options 8 Finest Virtualization Options

STIV: Scalable Textual content and Picture Conditioned Video Era

Introducing AWS Batch Assist for Amazon SageMaker Coaching jobs

Greatest Net Scraping Corporations in 2025

Automate the creation of handout notes utilizing Amazon Bedrock Information Automation

5 AI Buying and selling Bots That Work With Robinhood

Evaluating the Finest AI Video Mills for Social Media

Utilizing AI To Repair The Innovation Drawback: The Three Step Resolution

Midjourney V7: Quicker, smarter, extra reasonable

5 AI Buying and selling Bots That Work With Robinhood

Everest Ransomware Claims Mailchimp as New Sufferer in Comparatively Small Breach

VMware Options 8 Finest Virtualization Options

Introducing AWS Batch Assist for Amazon SageMaker Coaching jobs

Main Menu

Subscribe to Updates

What's Hot

STIV: Scalable Textual content and Picture Conditioned Video Era

Related Posts