<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v" -->

---
title: Grok Voice Transcribe 2.0: $0.10/hour transcription with...
description: xAI (now SpaceXAI after SpaceX&#x27;s acquisition) released Grok Voice Transcribe 2.0, a multilingual speech-to-text model keeping pricing at $0.10/hour batch and...
canonical: https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: Grok Voice Transcribe 2.0: $0.10/hour transcription with a murky data advantage | daily.dev
og:description: xAI (now SpaceXAI after SpaceX&#x27;s acquisition) released Grok Voice Transcribe 2.0, a multilingual speech-to-text model keeping pricing at $0.10/hour batch and...
og:url: https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v
og:image: https://api.daily.dev/og/posts/CbK9XzN2V.png
og:image:alt: Grok Voice Transcribe 2.0: $0.10/hour transcription with a murky data advantage
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Grok Voice Transcribe 2.0: $0.10/hour transcription with a murky data advantage

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

xAI (now SpaceXAI after SpaceX's acquisition) released Grok Voice Transcribe 2.0, a multilingual speech-to-text model keeping pricing at $0.10/hour batch and $0.20/hour streaming, claiming double the accuracy of v1.0 and top ranking among 32 streaming models on the Artificial Analysis leaderboard. Word error rate on multilingual short-phrase commands dropped from 20.6% to 6.8%, relevant to Tesla's in-car assistant. New features include speaker diarization, multichannel audio, key term biasing, and smart turn detection. Atlassian's Loom has already switched to the new model, and v1.0 will be deprecated within weeks with pinning available during transition. The piece raises concerns that xAI's evaluation data comes from production traffic including spoken account codes, phone numbers, and addresses, with no disclosure on consent or retention, framing this data as the company's real competitive moat.

## Content

xAI (now operating as SpaceXAI following SpaceX's February acquisition) has released Grok Voice Transcribe 2.0, a multilingual speech-to-text model priced at $0.10/hour for batch and $0.20/hour for streaming. Pricing is unchanged from version 1.0, but the company claims double the accuracy.

The model ranks first among 32 streaming models on the Artificial Analysis leaderboard. The biggest accuracy jump is on multilingual short-phrase commands, where word error rate dropped from 20.6% to 6.8% - a use case that maps directly to Tesla's in-car voice assistant.

Feature-wise, it supports batch and streaming transcription, speaker diarization, multichannel audio, key term biasing, text formatting, filler word removal, and smart turn detection.

## The training data question

The more interesting part of the announcement is where the evaluation data comes from: production traffic including customer-support calls, Grok conversations, voice commands, and users reading out account codes, phone numbers, and addresses. That last category is worth pausing on. The announcement doesn't explain how that spoken-credentials data is collected, retained, or consented to - which raises real questions under European data protection rules.

At commodity pricing, the model itself isn't the differentiator anymore. Access to diverse, real-world conversational data is. That's harder to replicate than matching a price point.

## Adoption and deprecation

Atlassian has already switched Loom's transcription pipeline to the new model. Version 1.0 will be deprecated within weeks, though API pinning is available during the transition.

## Other AI news from September 19

- **Meta**: Muse connectors are live for developers (bring your own API; Muse handles the agent, browser, and user context). Muse is now available in Canada, and a dedicated Muse Mail tab is in development.
- **OpenAI**: ChatGPT's desktop browser now runs Chrome extensions. Most plugins can connect multiple accounts in one chat, with a profile tool so ChatGPT can label them.
- **Anthropic**: Claude Code 2.1.277 reads AGENTS.md when no CLAUDE.md is present, toggled via /config. Anthropic is also partnering with Accenture on embedded frontier evaluation.
- **Google**: Google Pics is generally available in Workspace for image generation and co-creation. Dreambeans, a daily personalized story collection, is GA from Labs.
- **Mistral**: Investigated a claimed breach and says systems were not compromised.

## Questions this post answers

### What is the pricing for Grok Voice Transcribe 2.0?

Grok Voice Transcribe 2.0 costs $0.10 per hour for batch transcription and $0.20 per hour for streaming transcription, the same pricing as version 1.0. The model claims roughly double the accuracy of version 1.0, ranking first for accuracy among 32 streaming models on the Artificial Analysis leaderboard, with word error rate on multilingual short-phrase commands dropping from 20.6% to 6.8%.

_Comparing speech-to-text providers by price and accuracy gets easier when tracking releases like this on daily.dev._

### When is Grok Voice Transcribe 1.0 being deprecated?

Grok Voice Transcribe 1.0 will be deprecated within weeks of the 2.0 release, though pinning to the older version remains available during the transition period. Teams relying on the original model should plan to migrate to 2.0, which supports the same batch and streaming modes plus new features like speaker diarization and smart turn detection.

_Developers migrating between model versions can follow deprecation timelines like this on daily.dev._

### What privacy concerns exist around xAI's Grok Voice Transcribe training and evaluation data?

xAI builds its evaluation sets from production traffic including customer-support calls, Grok conversations, voice commands, and users reading out account codes, phone numbers, and addresses. The announcement does not disclose how this spoken credentials data is obtained, retained, or consented to, which raises concerns particularly under European data protection rules.

_Engineers weighing AI vendor data practices can keep tabs on disclosures like this via daily.dev._

## Community take

How the wider developer community reacted, aggregated from 2 discussions and 19 comments across x (as of 2026-09-19).

**TL;DR:** Reactions focus mostly on the $0.10/hour price point being reframed as commoditized 'plumbing' rather than a labor replacement, with some pushback that the feature set (diarization, key term biasing) is table stakes rather than a breakthrough and skepticism about the leaderboard methodology.

**Sentiment:** 30% positive · 45% mixed · 25% skeptical

**The case for**

- The ultra-low per-hour price is seen as turning transcription into cheap infrastructure rather than a costly feature.
- Combining multilingual streaming with diarization is viewed as a meaningful, useful capability if it holds up on real audio.

**The pushback**

- Diarization, key term biasing, and filler-word removal are called baseline features already offered by competitors like Deepgram or Whisper wrappers, not a breakthrough.
- Topping a 32-model leaderboard is questioned since benchmark details like language mix, latency budget, and streaming definition aren't clear.
- Some doubt the model's real-world performance on mumbling, noisy, or overlapping audio despite the leaderboard rank.

**By community**

- x (mixed): Replies mix genuine enthusiasm about the price shift with skepticism that the features are anything more than baseline parity, plus open questions about real-world performance.

**Hottest debate:** Whether ranking first on a 32-model leaderboard and shipping standard features constitutes real innovation or just competitive parity.

**Open questions**

- How well does it handle mumbling, noisy, or overlapping speech in practice?
- What exactly was the benchmark setup (language mix, latency budget, streaming definition, diarization scoring)?
- Which specific API integrations or Chrome extensions will still function as agent layers like Muse and ChatGPT desktop expand?

**Highlights**

> @testingcatalog slapping "diarization, key term biasing, and filler word removal" onto an API isn't a breakthrough, it’s literally the baseline feature set required to compete with Deepgram or Whisper wrapper implementations
> — [MoveDecisions on x](https://x.com/MoveDecisions/status/2101092809231192278)

> @testingcatalog Speech-to-text just got priced like storage, not labour. $0.10/hr batch means a year of your meetings costs less than the coffee drunk in them. Transcription stopped being a feature and became plumbing.
> — [BriskFalcon\_284 on x · 1 points, 1 comments](https://x.com/BriskFalcon_284/status/2101247742777475193)

> @testingcatalog Interesting result. The benchmark setup matters as much as the rank here—language mix, latency budget, streaming definition, and whether diarization was scored separately would make the comparison much easier to interpret.
> — [meimu7ns4 on x](https://x.com/meimu7ns4/status/2101178321300074633)

> @testingcatalog 32 models? so 1 in 32 is the only one that actually works?
> — [ADLXBT on x](https://x.com/ADLXBT/status/2101096839835754687)

> @testingcatalog The useful bit is the price-to-plumbing shift: once transcription is cheap, latency, retention and failure handling become the product. The invoice is no longer the scary part.
> — [ImZaneKelly on x](https://x.com/ImZaneKelly/status/2101256979503104362)

**Source threads**

- [x](https://x.com/testingcatalog/status/2101064766814871708) · 0 points · 6 comments
- [x](https://x.com/testingcatalog/status/2101246004255158279) · 0 points · 13 comments

---

Tags: [#data-privacy](https://daily.dev/tags/data-privacy), [#speech-recognition](https://daily.dev/tags/speech-recognition), [#grok](https://daily.dev/tags/grok)

[View this post on daily.dev](https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"Grok Voice Transcribe 2.0: $0.10/hour transcription with a murky data advantage","url":"https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v"},"datePublished":"2026-09-19T08:31:47.230Z","dateModified":"2026-09-19T15:45:40.416Z","description":"xAI (now SpaceXAI after SpaceX's acquisition) released Grok Voice Transcribe 2.0, a multilingual speech-to-text model keeping pricing at $0.10/hour batch and...","image":"https://editorial.thenextweb.com/wp-content/uploads/2026/09/og-grok-voice-transcribe-2-0.webp","thumbnailUrl":"https://editorial.thenextweb.com/wp-content/uploads/2026/09/og-grok-voice-transcribe-2-0.webp","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"data-privacy,speech-recognition,grok","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"Grok Voice Transcribe 2.0: $0.10/hour transcription with a murky data advantage"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/grok-voice-transcribe-2-0-0-10-hour-transcription-with-a-murky-data-advantage-cbk9xzn2v#faq","mainEntity":[{"@type":"Question","name":"What is the pricing for Grok Voice Transcribe 2.0?","acceptedAnswer":{"@type":"Answer","text":"Grok Voice Transcribe 2.0 costs $0.10 per hour for batch transcription and $0.20 per hour for streaming transcription, the same pricing as version 1.0. The model claims roughly double the accuracy of version 1.0, ranking first for accuracy among 32 streaming models on the Artificial Analysis leaderboard, with word error rate on multilingual short-phrase commands dropping from 20.6% to 6.8%. Comparing speech-to-text providers by price and accuracy gets easier when tracking releases like this on daily.dev."}},{"@type":"Question","name":"When is Grok Voice Transcribe 1.0 being deprecated?","acceptedAnswer":{"@type":"Answer","text":"Grok Voice Transcribe 1.0 will be deprecated within weeks of the 2.0 release, though pinning to the older version remains available during the transition period. Teams relying on the original model should plan to migrate to 2.0, which supports the same batch and streaming modes plus new features like speaker diarization and smart turn detection. Developers migrating between model versions can follow deprecation timelines like this on daily.dev."}},{"@type":"Question","name":"What privacy concerns exist around xAI's Grok Voice Transcribe training and evaluation data?","acceptedAnswer":{"@type":"Answer","text":"xAI builds its evaluation sets from production traffic including customer-support calls, Grok conversations, voice commands, and users reading out account codes, phone numbers, and addresses. The announcement does not disclose how this spoken credentials data is obtained, retained, or consented to, which raises concerns particularly under European data protection rules. Engineers weighing AI vendor data practices can keep tabs on disclosures like this via daily.dev."}}]}
```

