MAI Transcribe 2 Streaming, MAI Voice 2.1, and MAI Voice 2.
A voice agent built on three new first-party models from Microsoft can complete a conversational turn in well under one second: about 130 ms for the new streaming model to hear the user, roughly 350 ms for an AI to plan a reply, and 150 ms to start speaking. That is the math Microsoft published on October 1 when it shipped MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash.
The streaming model closes the leg that batch transcription left open. MAI-Transcribe-2-Streaming transcribes as the user talks rather than after, running over a persistent WebSocket connection in 100–200 ms audio packets with an encoder-decoder transformer using chunked attention. It emits both partial and final transcripts. The two Voice-2.1 models are the text-to-speech side. All three ship through Microsoft Foundry and Azure AI Speech.
The catch is the public-preview designation. None carry a service-level agreement, and Microsoft does not recommend them for production workloads until general availability. The 130 ms end-of-speech figure is Microsoft-measured under cited conditions, not independently benchmarked; no cross-vendor latency comparison was published. Streaming-tier pricing is also not posted; the $0.10 per audio hour figure belongs to the September batch predecessor MAI-Transcribe-2.
The same day, Vercel announced AI Gateway support, corroborating the models are live. A real production test will wait for GA.