Close Menu
    Main Menu
    • Home
    • News
    • Tech
    • Robotics
    • ML & Research
    • AI
    • Digital Transformation
    • AI Ethics & Regulation
    • Thought Leadership in AI

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Selfyz AI Video Technology App Evaluate: Key Options

    December 13, 2025

    UK’s ICO Superb LastPass £1.2 Million Over 2022 Safety Breach – Hackread – Cybersecurity Information, Information Breaches, AI, and Extra

    December 13, 2025

    Ai2's new Olmo 3.1 extends reinforcement studying coaching for stronger reasoning benchmarks

    December 13, 2025
    Facebook X (Twitter) Instagram
    UK Tech InsiderUK Tech Insider
    Facebook X (Twitter) Instagram
    UK Tech InsiderUK Tech Insider
    Home»Machine Learning & Research»STIV: Scalable Textual content and Picture Conditioned Video Era
    Machine Learning & Research

    STIV: Scalable Textual content and Picture Conditioned Video Era

    Oliver ChambersBy Oliver ChambersJuly 31, 2025No Comments2 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr Email Reddit
    STIV: Scalable Textual content and Picture Conditioned Video Era
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link


    The sector of video technology has made exceptional developments, but there stays a urgent want for a transparent, systematic recipe that may information the event of sturdy and scalable fashions. On this work, we current a complete research that systematically explores the interaction of mannequin architectures, coaching recipes, and knowledge curation methods, culminating in a easy and scalable text-image-conditioned video technology technique, named STIV. Our framework integrates picture situation right into a Diffusion Transformer (DiT) by body alternative, whereas incorporating textual content conditioning through a joint image-text conditional classifier-free steerage. This design allows STIV to carry out each text-to-video (T2V) and text-image-to-video (TI2V) duties concurrently. Moreover, STIV might be simply prolonged to numerous purposes, comparable to video prediction, body interpolation, multi-view technology, and lengthy video technology, and many others. With complete ablation research on T2I, T2V, and TI2V, STIV display sturdy efficiency, regardless of its easy design. An 8.7B mannequin with 512 decision achieves 83.1 on VBench T2V, surpassing each main open and closed-source fashions like CogVideoX-5B, Pika, Kling, and Gen-3. The identical-sized mannequin additionally achieves a state-of-the-art results of 90.1 on VBench I2V job at 512 decision. By offering a clear and extensible recipe for constructing cutting-edge video technology fashions, we goal to empower future analysis and speed up progress towards extra versatile and dependable video technology options.

    • † College of California, Los Angeles
    • ** Work performed whereas at Apple
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Oliver Chambers
    • Website

    Related Posts

    Constructing a voice-driven AWS assistant with Amazon Nova Sonic

    December 13, 2025

    How you can Write Environment friendly Python Knowledge Lessons

    December 13, 2025

    3 Function Engineering Strategies for Unstructured Textual content Knowledge

    December 13, 2025
    Top Posts

    Evaluating the Finest AI Video Mills for Social Media

    April 18, 2025

    Utilizing AI To Repair The Innovation Drawback: The Three Step Resolution

    April 18, 2025

    Midjourney V7: Quicker, smarter, extra reasonable

    April 18, 2025

    Meta resumes AI coaching utilizing EU person knowledge

    April 18, 2025
    Don't Miss

    Selfyz AI Video Technology App Evaluate: Key Options

    By Amelia Harper JonesDecember 13, 2025

    Selfyz AI is a cellphone app that depends on AI to construct animated clips from…

    UK’s ICO Superb LastPass £1.2 Million Over 2022 Safety Breach – Hackread – Cybersecurity Information, Information Breaches, AI, and Extra

    December 13, 2025

    Ai2's new Olmo 3.1 extends reinforcement studying coaching for stronger reasoning benchmarks

    December 13, 2025

    How Leaders Can Cease “Seeing Ghosts” And Overcome Invisible Threats and Alternatives in Enterprise

    December 13, 2025
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo

    Subscribe to Updates

    Get the latest creative news from SmartMag about art & design.

    UK Tech Insider
    Facebook X (Twitter) Instagram
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms Of Service
    • Our Authors
    © 2025 UK Tech Insider. All rights reserved by UK Tech Insider.

    Type above and press Enter to search. Press Esc to cancel.