Nvidia's Nemotron 3 Ultra, an open Mixture-of-Experts reasoning model, is now available on Vercel AI Gateway. The model features a 1M token context window, up to 350 tokens/second throughput, and is optimized for multi-turn agentic workflows including planning, tool use, sub-agent delegation, and error recovery — at up to 30% lower cost on agentic tasks. Access it via the AI SDK using model ID `nvidia/nemotron-3-ultra-550b-a55b`. Vercel AI Gateway provides a unified API with no markup on provider pricing, built-in usage tracking, failover, and Zero Data Retention support.

1m read timeFrom vercel.com
Post cover image
7 Impressions