<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/zq6h3l6u4" -->

---
title: Microsoft Releases VibeVoice-ASR: A Unified...
description: Microsoft has released VibeVoice-ASR, a unified speech-to-text model capable of processing 60-minute audio files in a single pass using a 64K token context...
canonical: https://daily.dev/posts/zq6h3l6u4
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass Microsoft VibeVoice ASR is a unified speech to text model for 60 minute… | daily.dev
og:description: Microsoft has released VibeVoice-ASR, a unified speech-to-text model capable of processing 60-minute audio files in a single pass using a 64K token context...
og:url: https://daily.dev/posts/zq6h3l6u4
og:image: https://api.daily.dev/og/posts/zq6H3l6U4.png
og:image:alt: Post cover image
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass Microsoft VibeVoice ASR is a unified speech to text model for 60 minute…

**[Machine Learning News](https://daily.dev/sources/mlnews)** · [@ailover](https://daily.dev/ailover) · 0 upvotes · 0 comments

## Summary

Microsoft has released VibeVoice-ASR, a unified speech-to-text model capable of processing 60-minute audio files in a single pass using a 64K token context window. The model simultaneously performs automatic speech recognition, speaker diarization, and timestamping to produce structured transcripts showing who spoke, when, and what was said. It supports customized hotwords for domain-specific terminology without requiring retraining, targeting meeting and conversational scenarios. The model is evaluated using metrics like DER, cpWER, and tcpWER, and integrates into meeting assistants, analytics tools, and transcription pipelines.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.linkedin.com/posts/asifrazzaq_microsoft-releases-vibevoice-asr-a-unified-share-7420218797598334976-ayvz>

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments
- [CNCF Unveils Schedule for KubeCon \+ CloudNativeCon Europe 2026](https://daily.dev/posts/cncf-unveils-schedule-for-kubecon-cloudnativecon-europe-2026-ikhcoa5cb) · CNCF · 2 upvotes · 0 comments
- [CNCF Debuts KubeCon \+ CloudNativeCon Japan 2026 Schedule](https://daily.dev/posts/cncf-debuts-kubecon-cloudnativecon-japan-2026-schedule-xp5pyudub) · CNCF · 1 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#microsoft](https://daily.dev/tags/microsoft), [#azure](https://daily.dev/tags/azure), [#nlp](https://daily.dev/tags/nlp), [#speech-recognition](https://daily.dev/tags/speech-recognition)

[View this post on daily.dev](https://daily.dev/posts/zq6h3l6u4)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/zq6h3l6u4","headline":"Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass Microsoft VibeVoice ASR is a unified speech to text model for 60 minute…","text":"Shared: Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass Microsoft VibeVoice ASR is a unified speech to text model for 60 minute…","url":"https://daily.dev/posts/zq6h3l6u4","datePublished":"2026-01-22T21:41:39.264Z","dateModified":"2026-01-22T21:42:15.005Z","author":{"@type":"Person","name":"Asif Razzaq","url":"https://daily.dev/ailover","image":"https://media.daily.dev/image/upload/s--fSsSf6QQ--/f_auto/v1754582511/avatars/avatar_2CaKzJAqMuPDI2am3eJjb?_a=BAMClqZW0","description":"Tech Blogger","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":680}},"image":"https://media.daily.dev/image/upload/s--foaA6JGU--/f_auto/v1722860399/public/Placeholder%2004","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/QABpFQWiZ"},"isPartOf":{"@type":"WebPage","url":"https://daily.dev/squads/mlnews","name":"Machine Learning News"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Machine Learning News","item":"https://daily.dev/squads/mlnews"},{"@type":"ListItem","position":3,"name":"Microsoft Releases VibeVoice-ASR: A Unified Speech-to-Text Model Designed to Handle 60-Minute Long-Form Audio in a Single Pass Microsoft VibeVoice ASR is a unified speech to text model for 60 minute…"}]}
```

