DeepSeek-V4 Pro is now available on Together AI with a 512K-token context window (upgradeable to 1M on dedicated infrastructure). The model uses a 1.6T-parameter Mixture-of-Experts architecture with 49B activated parameters and supports three controllable reasoning modes: Non-Think, Think High, and Think Max. Pricing is $2.10/1M input tokens, $0.20/1M cached input tokens, and $4.40/1M output tokens — a 90% cost reduction for reused context. Key use cases include code agents, document intelligence, long-context agent traces, and research synthesis. Teams can start with serverless inference and move to dedicated reserved capacity for production workloads.

5m read timeFrom together.ai
Post cover image
Table of contents
At a glanceBuilt for long-context reasoningChoose reasoning effort by workloadMake repeated long-context queries cheaper with cached input pricingWorkload patternsStart serverless, move to reserved capacityTry it nowGet started
1 Impression