AI news

Microsoft Unveils New AI Speech Models

Microsoft Unveils New AI Speech Models
———————————
Microsoft has launched three new AI models for voice applications: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The transcription model supports 60 languages with real-time, incremental speech-to-text and automatic language detection.

According to Artificial Analysis, MAI-Transcribe-2-Streaming ranked first for both final and partial transcript accuracy in its streaming evaluation. Microsoft reports that the model can begin producing transcription just over 100 milliseconds after receiving audio.

Microsoft also introduced MAI-Voice-2.1 and its Flash variant for text-to-speech. The Flash model supports 23 languages, can generate 45 seconds of audio with about 150 milliseconds of end-to-end latency, and is priced at about 60% less than comparable models, according to Microsoft.

Related Articles

Leave a Reply

Back to top button