Slack's engineering team details their three-year journey scaling LLM infrastructure from AWS SageMaker to a multi-cloud architecture spanning AWS Bedrock and GCP Vertex AI. The post covers four phases: self-managed SageMaker deployments with GPU scarcity challenges, migration to Bedrock for operational simplicity and model access, transitioning to on-demand capacity with hybrid routing and spillover patterns, and finally expanding to multi-cloud with an intelligent routing layer featuring circuit breakers, A/B testing, and API normalization. Key outcomes include ~10% quality improvement for reasoning tasks and ~67% latency reduction for low-token workloads, achieved through provider-agnostic abstraction and cross-functional alignment on security and compliance.