Together AI is releasing a major update to its Dedicated Model Inference platform, designed to give teams full control over open-weight model deployments without building custom infrastructure. Key features include multiple deployments behind a single stable endpoint, canary/blue-green/rolling updates with auto-rollback, A/B and shadow traffic testing, autoscaling across regions, and a Prometheus-compatible observability endpoint. The platform also introduces roughly 4× faster model warm starts via a rebuilt caching and distribution layer. Alongside this, a closed beta for custom training launches, covering full-weight and LoRA reinforcement learning and supervised fine-tuning, with checkpoints deployable directly to production inference without platform switching.

8m read timeFrom together.ai
Post cover image
Table of contents
Bring the model you want to serveChoose how the model runsSpend less time waiting for models to startControl where and how your deployment scalesKeep the endpoint stable while deployments evolveSafely move changes into productionTest against real traffic without putting users at riskMeasure what is happeningJoin our reinforcement learning betaWhat's next
103 Impressions