DeepSeek-V4 Pro is now available on Together AI with a 512K-token context window (upgradeable to 1M on dedicated infrastructure). The model uses a 1.6T-parameter Mixture-of-Experts architecture with 49B activated parameters and supports three controllable reasoning modes: Non-Think, Think High, and Think Max. Pricing is $2.10/1M input tokens, $0.20/1M cached input tokens, and $4.40/1M output tokens — a 90% cost reduction for reused context. Key use cases include code agents, document intelligence, long-context agent traces, and research synthesis. Teams can start with serverless inference and move to dedicated reserved capacity for production workloads.
Table of contents
At a glanceBuilt for long-context reasoningChoose reasoning effort by workloadMake repeated long-context queries cheaper with cached input pricingWorkload patternsStart serverless, move to reserved capacityTry it nowGet started1 Impression