16 articles
Kuaishou's Kling Team has introduced HyDRA, a hybrid memory system that keeps AI-generated subjects consistent even when they leave the frame entirely.
A new token-dropping technique accelerates large vision-language models at inference time - no retraining, no architectural changes, and minimal performance loss.
A new system, StreamDiffusionV2, enables real-time AI video generation with ultra-low latency. The pipeline achieves sub-0.5s first-frame response and up to 64 FPS.
15B multimodal AI model generates 5-second video in 2 seconds on H100, unifying text, video, and audio in one Transformer.
Baidu's SAMA open-source model rivals Kling-Omni in video editing with semantic anchoring and motion alignment.
Pruna AI's P-Video has topped a new Artificial Analysis comparison as the fastest and cheapest AI video model available via public API — generating a 720p five-second clip in around 10 seconds while undercutting rivals on price.
Helios, a 14B video generation model developed by ByteDance and Peking University researchers, reportedly reaches 19.5 FPS on a single H100 GPU - achieving near real-time generation without relying on common acceleration techniques.
A new AI framework called CineTrans dramatically improves scene transition quality in AI-generated videos by training on 250,000 film clips and applying a novel masking technique. The result is output that finally starts to feel like something a human editor might actually approve.
A new benchmark reveals leading multimodal AI models struggle to predict future events from audio and video, with top accuracy reaching only 64.8%.
HKUST researchers unveiled One4D, an AI that builds dynamic 3D video scenes from a single image or limited frames. The system's Decoupled LoRA Control method keeps visuals sharp while maintaining stable 3D geometry across time.
Researchers introduced LongVie 2, a controllable AI "world model" capable of generating coherent videos lasting up to five minutes from minimal input.
Kapwing's study shows AI-generated videos now dominate YouTube recommendations, with 278 channels raking in $117M annually from 63 billion views.
Researchers built a new system that lets AI follow tasks one step at a time - watching how objects physically change, instead of merely tagging what action happens.
FlovaAI unveils a comprehensive AI-powered video creation platform that transforms text, images, and audio into fully edited, multi-scene videos through a single streamlined workflow.
Kling O1 launched a unified multi-modal AI system that handles text-to-video generation, keyframe editing, and advanced content creation in one platform, intensifying competition in the AI video space.
Despite creating photorealistic videos, AI models like Sora and VideoPoet struggle with basic physics, according to a new Google DeepMind study. The Physics-IQ benchmark reveals that visual realism doesn't equal actual understanding of how the physical world works.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy