<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw" -->

---
title: Microsoft&#x27;s MAI-Voice-2.1 and MAI-Transcribe-2 Streaming...
description: Microsoft&#x27;s MAI-Voice-2.1, MAI-Voice-2.1-Flash, and MAI-Transcribe-2 Streaming audio models are now accessible through Vercel&#x27;s AI Gateway, billed at listed...
canonical: https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Microsoft&#x27;s MAI-Voice-2.1 and MAI-Transcribe-2 Streaming audio models arrive on Vercel AI Gateway | daily.dev
og:description: Microsoft&#x27;s MAI-Voice-2.1, MAI-Voice-2.1-Flash, and MAI-Transcribe-2 Streaming audio models are now accessible through Vercel&#x27;s AI Gateway, billed at listed...
og:url: https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw
og:image: https://api.daily.dev/og/posts/riLkf1Ykw.png
og:image:alt: Microsoft&#x27;s MAI-Voice-2.1 and MAI-Transcribe-2 Streaming audio models arrive on Vercel AI Gateway
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Microsoft's MAI-Voice-2.1 and MAI-Transcribe-2 Streaming audio models arrive on Vercel AI Gateway

**[Collections](https://daily.dev/sources/collections)** · 1 min read · 0 upvotes · 0 comments

## Summary

Microsoft's MAI-Voice-2.1, MAI-Voice-2.1-Flash, and MAI-Transcribe-2 Streaming audio models are now accessible through Vercel's AI Gateway, billed at listed rates with no platform markup. The speech models generate audio in 23 languages, while the transcription model streams partial transcripts in real time. Developers can call them in AI SDK 7 using generateSpeech and streamTranscribe with model identifiers microsoft/mai-voice-2.1-flash and microsoft/mai-transcribe-2-streaming. Microsoft also launched a demo voice experience called Chatter in its MAI Playground to showcase the voices.

## Content

Microsoft's new MAI audio models are now available through Vercel's AI Gateway. There are three: MAI-Voice-2.1 and MAI-Voice-2.1-Flash for speech generation, and MAI-Transcribe-2 Streaming for live transcription.

## What's available

- **MAI-Voice-2.1 and MAI-Voice-2.1-Flash** generate speech in 23 languages.
- **MAI-Transcribe-2 Streaming** returns partial transcripts while someone is still talking.

Vercel bills the models at their listed rates, with no platform markup.

## Calling them from AI SDK 7

The models use the identifiers `microsoft/mai-voice-2.1-flash` and `microsoft/mai-transcribe-2-streaming`. In AI SDK 7, you call them with `generateSpeech` for text-to-speech and `streamTranscribe` for live transcription.

## Microsoft's pitch

@testingcatalog flagged the models as coming to the MAI Playground and Microsoft's APIs. Microsoft describes them as "accurate, fast, low cost, and chart-topping audio understanding and generation for building the best conversational voice agents." That is Microsoft's own claim, and I'd want to test it before taking it at face value.

The MAI Playground also has a new demo voice experience called "Chatter," which is a quick way to hear the voices.

If you're building a voice agent, the useful part is having transcription and speech generation behind one gateway, so you can swap between them without changing providers.

## Questions this post answers

### How do I call Microsoft's MAI-Voice-2.1-Flash model using AI SDK 7?

Use the generateSpeech function with the model identifier microsoft/mai-voice-2.1-flash for text-to-speech. For live transcription, use streamTranscribe with microsoft/mai-transcribe-2-streaming. Both models are accessible through Vercel's AI Gateway, which bills at the models' listed rates without adding a platform markup.

_daily.dev surfaces integration details like this for developers wiring new voice models into their stack._

### What audio models did Microsoft add to Vercel AI Gateway?

Microsoft added three models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash for text-to-speech generation across 23 languages, and MAI-Transcribe-2 Streaming for live transcription that returns partial transcripts while someone is still speaking. All three are billed at Microsoft's listed rates with no added markup from Vercel.

_Developers evaluating voice agent providers follow releases like this through daily.dev._

## Similar posts on daily.dev

- [AI Gateway now supports streaming transcription](https://daily.dev/posts/ai-gateway-now-supports-streaming-transcription-xgbwbzesl) · Vercel · 0 upvotes · 0 comments
- [OpenAI launches new voice intelligence features in its API](https://daily.dev/posts/openai-launches-new-voice-intelligence-features-in-its-api-ymtm53lpv) · TechCrunch · 0 upvotes · 0 comments

---

Tags: [#microsoft](https://daily.dev/tags/microsoft), [#speech-recognition](https://daily.dev/tags/speech-recognition), [#text-to-speech](https://daily.dev/tags/text-to-speech), [#ai-sdk](https://daily.dev/tags/ai-sdk)

[View this post on daily.dev](https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Microsoft's MAI-Voice-2.1 and MAI-Transcribe-2 Streaming audio models arrive on Vercel AI Gateway","url":"https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw"},"datePublished":"2026-10-01T18:16:12.948Z","dateModified":"2026-10-01T18:16:44.597Z","description":"Microsoft's MAI-Voice-2.1, MAI-Voice-2.1-Flash, and MAI-Transcribe-2 Streaming audio models are now accessible through Vercel's AI Gateway, billed at listed...","image":"https://pbs.twimg.com/media/HTkIJzwXkAAIxMz.jpg","thumbnailUrl":"https://pbs.twimg.com/media/HTkIJzwXkAAIxMz.jpg","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"microsoft,speech-recognition,text-to-speech,ai-sdk","timeRequired":"PT1M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Microsoft's MAI-Voice-2.1 and MAI-Transcribe-2 Streaming audio models arrive on Vercel AI Gateway"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/microsoft-s-mai-voice-2-1-and-mai-transcribe-2-streaming-audio-models-arrive-on-vercel-ai-gateway-rilkf1ykw#faq","mainEntity":[{"@type":"Question","name":"How do I call Microsoft's MAI-Voice-2.1-Flash model using AI SDK 7?","acceptedAnswer":{"@type":"Answer","text":"Use the generateSpeech function with the model identifier microsoft/mai-voice-2.1-flash for text-to-speech. For live transcription, use streamTranscribe with microsoft/mai-transcribe-2-streaming. Both models are accessible through Vercel's AI Gateway, which bills at the models' listed rates without adding a platform markup. daily.dev surfaces integration details like this for developers wiring new voice models into their stack."}},{"@type":"Question","name":"What audio models did Microsoft add to Vercel AI Gateway?","acceptedAnswer":{"@type":"Answer","text":"Microsoft added three models: MAI-Voice-2.1 and MAI-Voice-2.1-Flash for text-to-speech generation across 23 languages, and MAI-Transcribe-2 Streaming for live transcription that returns partial transcripts while someone is still speaking. All three are billed at Microsoft's listed rates with no added markup from Vercel. Developers evaluating voice agent providers follow releases like this through daily.dev."}}]}
```

