Deepgram's CEO discusses the company's journey from particle physics research to building scalable voice AI systems. The conversation covers technical architecture decisions—combining CNNs, RNNs, and attention mechanisms for end-to-end deep learning—and practical challenges like handling dialects, noisy environments, and reducing costs from $3/hour to under $2/hour. Key topics include synthetic data generation for training, ethical considerations around voice cloning, AWS Bedrock integration for bidirectional streaming, and the vision for a 'Neuroplex' architecture that maintains modularity while passing full context through speech-to-speech systems.

35m read timeFrom stackoverflow.blog
Post cover image
Table of contents
TRANSCRIPT
304 Impressions