Microsoft did not announce MAI-Realtime. On August 2, 2026, the model simply appeared as a hidden early-access entry inside Microsoft's MAI Playground, visible to a small group of partners, with no model card, no pricing page, no launch date, and no blog post. It is Microsoft's first native full-duplex voice model, and the quiet way it surfaced says as much about Microsoft's strategy as the model itself.
What full-duplex actually means
Most voice assistants are half-duplex. You talk, they wait for a pause, then they answer. MAI-Realtime is Microsoft's first system built to listen and speak at the same time rather than trading turns. Reporting from TestingCatalog, which surfaced the model, describes low response latency and clean handling of interruptions, the two things that make an overlapping conversation feel natural rather than robotic.
Half-duplex takes turns. Full-duplex overlaps.
The difference MAI-Realtime is chasing: a voice system that listens and speaks at the same moment, the way people actually talk over each other, rather than waiting for a gap.
Each side waits for silence. Interrupting is awkward, and a pause sits between every exchange.
Speech overlaps. The model can back-channel, be interrupted cleanly, and respond without a turn-taking delay.
The model ships with two voices, Victoria and Grant, both described as noticeably more natural than what Copilot's voice mode currently delivers. It supports 16 languages, including English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese, and Thai, and can switch between them mid-conversation. Turn-taking runs through two modes: a "Switchboard" mode using Microsoft's own MAI-Ears endpointing with inline control tokens, and a deterministic setup that combines silence detection with semantic endpointing. Notably, it is strictly conversational. It does not sing or produce non-speech sounds, which marks it as a talking model rather than a general audio generator.
The real story is vertical integration
MAI-Realtime fills a specific hole. Microsoft's earlier speech models, MAI-Voice-2 for synthesis and MAI-Transcribe-1.5 for transcription, run in one direction only, and Azure has leaned on OpenAI's GPT-Realtime for live, bidirectional voice. A native full-duplex model is the piece Microsoft did not yet own.
That matters because of who owns the rest of the stack. Mustafa Suleyman's Microsoft AI group shipped seven in-house models at Build 2026 and has been steadily swapping OpenAI-supplied components out of Copilot, Teams, and Bing. Real-time voice was one of the last capabilities still routed through a partner. MAI-Realtime is how that dependency starts to close.
Swapping OpenAI parts out, one layer at a time
Microsoft AI has been building its own model for each capability its products used to source from OpenAI. MAI-Realtime is the newest piece, aimed at the real-time voice layer.
MAI-Realtime remains an internal preview with no model card, pricing, or launch date. Placement here reflects the capability it targets, not a shipped product.

Microsoft and OpenAI remain deeply intertwined commercially, so it would be a mistake to read this as a clean break. It is better understood as optionality. Owning a full-duplex voice model gives Microsoft a fallback, a bargaining position, and the freedom to tune latency, cost, and voice behavior for its own products rather than inheriting a partner's roadmap.
Real-time voice was one of the last capabilities Microsoft still routed through a partner. MAI-Realtime is how that dependency starts to close.
On the strategy behind the model
A crowded frontier
MAI-Realtime does not arrive into open space. OpenAI's GPT-Realtime line and Google's Gemini Live already offer overlapping, low-latency speech, and independent efforts like Sesame have pushed on naturalness. The competitive question is no longer whether full-duplex voice is possible but whose is cheapest, most natural, and easiest to embed. That is why the voice layer increasingly looks like a swappable component rather than a moat.
For anyone building on top of these systems, the practical lesson is portability. When three platforms ship comparable real-time voice within a year of each other, being locked to a single provider's voice model is a liability. Model-agnostic tools, including how we think about live voice at Metir, treat the underlying speech engine as a choice to be made per task and per price, not a permanent commitment. A preview that appears without so much as a model card is a reminder of how fast that choice can change.
What to watch next
The open questions are the ordinary ones a preview leaves unanswered: pricing, rate limits, launch regions, and whether MAI-Realtime shows up inside Copilot's consumer voice mode or stays an enterprise and developer offering first. Microsoft has said nothing official, and until it does, the benchmarks and the real-world latency remain unproven. What is already clear is the direction. Microsoft wants to own the voice in its products, and MAI-Realtime is the clearest sign yet that it intends to.
Sources:
- Exclusive: Microsoft tests new MAI Realtime voice model | TestingCatalog
- Microsoft previews MAI-Realtime bidirectional voice model | Crypto Briefing
- Microsoft MAI Realtime Appears in Hidden Preview, No Launch Confirmed | Windows Forum
- Introducing MAI-Voice-2 | Microsoft AI
- Microsoft MAI Models at Build 2026: In-House Reasoning, Image, Voice, and Coding | Windows News
Image credits
- Hero: Building 92 of the Microsoft Redmond campus. Photo by Jiaqian AirplaneFan, licensed CC BY 3.0, via Wikimedia Commons.