Announcing the fastest inference for realtime voice AI agents
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
Together AI has launched a full voice AI infrastructure stack targeting real-time voice agent developers. Key additions include: Streaming Whisper STT via WebSocket with optimized voice activity detection (VAD) completing transcripts up to 35% faster than alternatives; two serverless open-source TTS models — Orpheus (187ms TTFB, high fidelity) and Kokoro (97ms TTFB, ultra-low latency); Voxtral Mini for premium multilingual batch transcription with lower word error rates than Whisper; and automatic speaker diarization. The infrastructure runs on the same GPU clusters as LLMs to minimize cross-provider latency, supports WebSocket multiplexing for high-concurrency deployments, and targets sub-200ms TTS and sub-500ms end-to-end response times for natural conversation flow.