Microsoft MAI-Transcribe-2: Speed, Price, Benchmarks
Microsoft's MAI-Transcribe-2 tops the FLEURS speech benchmark across 60 languages at $0.10 an hour, undercutting OpenAI, Google and ElevenLabs on price.
Microsoft’s in-house AI group has a new flagship, and this time it is not a chatbot. On September 3, 2026, Microsoft AI released MAI-Transcribe-2, a speech-recognition model the company calls “the fastest, most accurate and cheapest” of its kind. The claim is aggressive, the pricing is aggressive, and the target is unambiguous: the transcription APIs sold by OpenAI, Google, and ElevenLabs.
MAI-Transcribe-2 is the latest model built by the team under Mustafa Suleyman, and it continues Microsoft AI’s push to ship frontier systems trained in-house rather than relying solely on its OpenAI partnership. Where earlier MAI releases targeted chat and image generation, this one goes after a workhorse enterprise task — turning audio into text — that sits underneath call centers, meeting tools, media captioning, and a growing layer of voice-driven AI agents.
What Microsoft is claiming
The headline number is accuracy. Microsoft says MAI-Transcribe-2 ranks first on the FLEURS benchmark across 60 languages, with an average word error rate of 5.2%. FLEURS is a widely used multilingual speech test, and word error rate — the share of words a system gets wrong — is the standard yardstick in the field. The company also says the model defines the Pareto frontier for accuracy versus latency on the independent Artificial Analysis leaderboard, where it ranks second on raw word-error-rate while running far faster than the models above and beside it.
Speed is the second pillar of the pitch. Microsoft says MAI-Transcribe-2 is:
- 10x faster than OpenAI’s GPT-Transcribe
- 7x faster than ElevenLabs’ Scribe v2
- 5x faster than Google’s Gemini 3.5 Transcribe
all while delivering higher accuracy than each. In transcription, latency is not a vanity metric. Live captioning, real-time meeting notes, and voice agents all depend on how quickly audio can be turned into usable text, and a model that is both faster and more accurate collapses a trade-off engineers have lived with for years.
The company says the model beats Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and ScribeV2 across a broad range of real-world audio — not just clean benchmark clips, but the noisy, accented, overlapping speech that breaks weaker systems.
The features that matter to builders
Raw accuracy is table stakes; the surrounding features are what determine whether a transcription model can be dropped into production. MAI-Transcribe-2 ships with a set aimed squarely at enterprise workloads:
- Speaker diarization — labeling who said what, essential for meetings, interviews, and support calls.
- Word-level timestamps — aligning each word to the audio, which underpins searchable transcripts, captioning, and editing tools.
- Automatic language identification and code-switching — handling mixed-language speech within a single utterance, a persistent weak spot for older systems.
- Keyword biasing — nudging the model toward domain-specific terminology like product names, drug names, or legal terms.
- Configurable transcription styles — choosing between a verbatim transcript (every “um” and false start) and a clean one, depending on the use case.
That combination targets the messy middle of real deployments, where a model has to do more than spit out a single best guess at the words.
The price is the weapon
The most disruptive number in the announcement is not on any benchmark. Microsoft priced MAI-Transcribe-2 at $0.10 per hour of audio, with that promotional rate held through the end of 2026. For high-volume transcription — media archives, call-center recordings, compliance logging — pricing is often the deciding factor, and a dime an hour reframes what “cheap” means in the category.
The model is available now in Azure Speech and Microsoft Foundry as a public preview, without a service-level agreement and explicitly not recommended for production workloads yet. That caveat matters: preview status means the price, the availability, and the terms can all shift before general availability. But it also lets Microsoft put the model in developers’ hands immediately and start applying pressure while competitors decide how to respond.
The distribution advantage is real. Microsoft can route MAI-Transcribe-2 straight into Azure, Foundry, Teams, and the broader Copilot stack, the same compute-priority machinery it has used to push its own models to the front of the queue. A transcription model that is cheaper and faster becomes far more threatening when it is one API call away for every existing Azure customer.
Why Microsoft is building its own
Transcription might look like a narrow slice of the AI market, but it is strategically central to where the industry is heading. Voice is becoming a primary interface for AI systems — from full-duplex voice models that hold real conversations to agents that talk fast and think fast. Every one of those systems needs to convert speech to text quickly and cheaply, and owning that layer means Microsoft is not paying a rival for a core dependency.
Building in-house also gives Microsoft leverage. The company’s relationship with OpenAI remains deep, but MAI releases signal a deliberate strategy of reducing single-vendor reliance and controlling its own model roadmap. A speech model that beats GPT-Transcribe on Microsoft’s own benchmarks — and undercuts it on price by an order of magnitude — is a pointed statement about who Microsoft intends to depend on.
The caveats
Vendor benchmarks deserve scrutiny, and MAI-Transcribe-2’s numbers come from Microsoft. The FLEURS ranking and the “fastest, cheapest, most accurate” framing are the company’s own claims; the independent Artificial Analysis placement is more useful precisely because it is third-party, and there the model ranks second on word error rate, not first — fast and near-best, rather than best outright. Real-world accuracy also varies enormously by domain: medical dictation, heavily accented speech, and low-resource languages all stress a model differently than a benchmark suite does.
The preview label is the other asterisk. The $0.10 rate is promotional and time-limited, there is no SLA, and Microsoft explicitly warns against production use. Teams evaluating the model should treat today’s numbers as a ceiling that could move before the model is generally available.
What it means
MAI-Transcribe-2 is a price and speed shock to a market that many had come to treat as commoditized. By pairing a top-tier accuracy claim with $0.10-an-hour pricing and order-of-magnitude speed gains, Microsoft is trying to reset expectations in a category where OpenAI, Google, and ElevenLabs have been comfortable charging more.
The clearest losers, if the claims hold up, are the specialist transcription vendors and the pricing power of incumbent speech APIs. ElevenLabs, whose Scribe line competes directly here, faces a well-capitalized rival willing to undercut it dramatically; OpenAI and Google see one of their utility models leapfrogged on the metrics that enterprise buyers actually compare. The winner, at least for now, is Microsoft — which gets a cheaper internal dependency, a competitive wedge into every Azure account, and another proof point that its in-house AI group can ship frontier systems, not just fine-tunes.
The bigger signal is about where model competition is going. As frontier chat models converge and open-weight systems close the gap, the fight is shifting to the unglamorous infrastructure layers — transcription, embeddings, routing, inference economics — where speed and cost, not headline intelligence, decide who wins. MAI-Transcribe-2 is a bet that Microsoft can compete hardest exactly there.
What to watch next: whether independent evaluations reproduce the accuracy claims outside FLEURS, whether the $0.10 price survives the move from preview to general availability, and how OpenAI, Google, and ElevenLabs respond. In a market this price-sensitive, a credible “fastest and cheapest” claim rarely goes unanswered for long.
Tagged
Keep reading
Chisato · · 6 min read Microsoft Copilot CoSnitch Flaw (CVE-2026-24301)
CoSnitch let one click on a link make Microsoft Copilot exfiltrate a victim's Gmail and Drive data. How the chained flaw worked and why Varonis called it meta-hacking.
Chisato · · 7 min read Microsoft Maia 300: TSMC Order and Nvidia Challenge
Microsoft is in talks with TSMC to build 300,000+ Maia 300 AI chips, aiming for over 1 million units to cut its reliance on Nvidia. The plan and what it means.
Kurumi · · 5 min read Microsoft Stock: Record $450B One-Day Market Cap Gain
Microsoft added about $450 billion in value on July 30, 2026 — the largest single-day gain in market history — as Azure cloud growth accelerated. What drove it.