2 articles
Meituan introduced LongCat-AudioDiT, a diffusion-based TTS model operating in waveform latent space. The system improves voice cloning and multilingual audio generation by eliminating intermediate representations like mel-spectrograms.
Google rolled out major upgrades to its Gemini 2.5 Flash TTS and Gemini 2.5 Pro TTS models, introducing style controls, smarter pacing, and better multi-speaker performance.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy