metir
metir
Download on App StoreGet it on Google PlayF1 FantasyLoginSign Up
Back to Blog
Microsoft
Voice AI
MAI
OpenAI
AI Models

Microsoft's MAI-Realtime: Its First Full-Duplex Voice AI

Microsoft quietly surfaced MAI-Realtime, its first full-duplex voice model that listens and speaks at once. What it signals about Microsoft building AI in-house and reducing its OpenAI dependence.

Metir AI TeamAugust 3, 20266 min read
Microsoft's MAI-Realtime: Its First Full-Duplex Voice AI

Microsoft did not announce MAI-Realtime. On August 2, 2026, the model simply appeared as a hidden early-access entry inside Microsoft's MAI Playground, visible to a small group of partners, with no model card, no pricing page, no launch date, and no blog post. It is Microsoft's first native full-duplex voice model, and the quiet way it surfaced says as much about Microsoft's strategy as the model itself.

Microsoft logoMicrosoft
OpenAI logoOpenAI
Google logoGoogle
The full-duplex voice race now runs across all three of the largest AI platforms.

What full-duplex actually means

Most voice assistants are half-duplex. You talk, they wait for a pause, then they answer. MAI-Realtime is Microsoft's first system built to listen and speak at the same time rather than trading turns. Reporting from TestingCatalog, which surfaced the model, describes low response latency and clean handling of interruptions, the two things that make an overlapping conversation feel natural rather than robotic.

Half-duplex takes turns. Full-duplex overlaps.

The difference MAI-Realtime is chasing: a voice system that listens and speaks at the same moment, the way people actually talk over each other, rather than waiting for a gap.

Turn-based (half-duplex)
You
Model
You

Each side waits for silence. Interrupting is awkward, and a pause sits between every exchange.

Full-duplex
You
Model

Speech overlaps. The model can back-channel, be interrupted cleanly, and respond without a turn-taking delay.

The model ships with two voices, Victoria and Grant, both described as noticeably more natural than what Copilot's voice mode currently delivers. It supports 16 languages, including English, German, Spanish, French, Italian, Portuguese, Japanese, Korean, Chinese, Dutch, Hindi, Indonesian, Arabic, Russian, Turkish, Vietnamese, and Thai, and can switch between them mid-conversation. Turn-taking runs through two modes: a "Switchboard" mode using Microsoft's own MAI-Ears endpointing with inline control tokens, and a deterministic setup that combines silence detection with semantic endpointing. Notably, it is strictly conversational. It does not sing or produce non-speech sounds, which marks it as a talking model rather than a general audio generator.

16Languages supportedSwitches mid-conversation
2Built-in voicesVictoria and Grant
0Public model cardsStill an internal preview
Aug 2, 2026First surfacedHidden entry in MAI Playground

The real story is vertical integration

MAI-Realtime fills a specific hole. Microsoft's earlier speech models, MAI-Voice-2 for synthesis and MAI-Transcribe-1.5 for transcription, run in one direction only, and Azure has leaned on OpenAI's GPT-Realtime for live, bidirectional voice. A native full-duplex model is the piece Microsoft did not yet own.

That matters because of who owns the rest of the stack. Mustafa Suleyman's Microsoft AI group shipped seven in-house models at Build 2026 and has been steadily swapping OpenAI-supplied components out of Copilot, Teams, and Bing. Real-time voice was one of the last capabilities still routed through a partner. MAI-Realtime is how that dependency starts to close.

Swapping OpenAI parts out, one layer at a time

Microsoft AI has been building its own model for each capability its products used to source from OpenAI. MAI-Realtime is the newest piece, aimed at the real-time voice layer.

Capability
In-house MAI model
What it can displace
Text and reasoning
MAI-1 family
OpenAI GPT models
Speech synthesis
MAI-Voice-2
OpenAI voice
Transcription
MAI-Transcribe-1.5
OpenAI Whisper-class
Real-time voice
MAI-Realtime (preview)
OpenAI GPT-Realtime

MAI-Realtime remains an internal preview with no model card, pricing, or launch date. Placement here reflects the capability it targets, not a shipped product.

Building 92 on Microsoft's Redmond headquarters campus, with the Microsoft logo in the foreground
Microsoft's Redmond campus. Mustafa Suleyman's Microsoft AI group has been building an in-house model for each capability its products once sourced from OpenAI. Photo: Jiaqian AirplaneFan, CC BY 3.0.

Microsoft and OpenAI remain deeply intertwined commercially, so it would be a mistake to read this as a clean break. It is better understood as optionality. Owning a full-duplex voice model gives Microsoft a fallback, a bargaining position, and the freedom to tune latency, cost, and voice behavior for its own products rather than inheriting a partner's roadmap.

“

Real-time voice was one of the last capabilities Microsoft still routed through a partner. MAI-Realtime is how that dependency starts to close.

On the strategy behind the model

A crowded frontier

MAI-Realtime does not arrive into open space. OpenAI's GPT-Realtime line and Google's Gemini Live already offer overlapping, low-latency speech, and independent efforts like Sesame have pushed on naturalness. The competitive question is no longer whether full-duplex voice is possible but whose is cheapest, most natural, and easiest to embed. That is why the voice layer increasingly looks like a swappable component rather than a moat.

For anyone building on top of these systems, the practical lesson is portability. When three platforms ship comparable real-time voice within a year of each other, being locked to a single provider's voice model is a liability. Model-agnostic tools, including how we think about live voice at Metir, treat the underlying speech engine as a choice to be made per task and per price, not a permanent commitment. A preview that appears without so much as a model card is a reminder of how fast that choice can change.

What to watch next

The open questions are the ordinary ones a preview leaves unanswered: pricing, rate limits, launch regions, and whether MAI-Realtime shows up inside Copilot's consumer voice mode or stays an enterprise and developer offering first. Microsoft has said nothing official, and until it does, the benchmarks and the real-world latency remain unproven. What is already clear is the direction. Microsoft wants to own the voice in its products, and MAI-Realtime is the clearest sign yet that it intends to.

Sources:

  • Exclusive: Microsoft tests new MAI Realtime voice model | TestingCatalog
  • Microsoft previews MAI-Realtime bidirectional voice model | Crypto Briefing
  • Microsoft MAI Realtime Appears in Hidden Preview, No Launch Confirmed | Windows Forum
  • Introducing MAI-Voice-2 | Microsoft AI
  • Microsoft MAI Models at Build 2026: In-House Reasoning, Image, Voice, and Coding | Windows News

Image credits

  • Hero: Building 92 of the Microsoft Redmond campus. Photo by Jiaqian AirplaneFan, licensed CC BY 3.0, via Wikimedia Commons.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Blog
  • Enterprise

Company

  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational