3 articles
Microsoft has open-sourced VibeVoice, a voice AI system capable of processing long-form audio and real-time speech tasks. The model supports structured transcription, multilingual capabilities, and low-latency voice generation.
Microsoft's VibeVoice combines long-form speech recognition and synthesis in a single open-source framework - processing up to 60-minute conversations and generating up to 90 minutes of consistent multi-speaker audio.
Microsoft released VibeVoice, an open-source real-time speech generation system that delivers conversational audio with just 300 milliseconds of latency, supports up to 90 minutes of continuous output, and handles four different speakers in a single session.
We use cookies to improve your experience on our site and to show you relevant advertising. To find our more, read our privacy policy and cookie policy