2 articles
15B multimodal AI model generates 5-second video in 2 seconds on H100, unifying text, video, and audio in one Transformer.
A new benchmark reveals leading multimodal AI models struggle to predict future events from audio and video, with top accuracy reaching only 64.8%.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy