Alibaba's Qwen3.8-2.4T-A95B open weights model is now available on Modal as a Shared Endpoint with OpenAI-compatible, token-based pricing. Modal partnered with Qwen ahead of launch to provide day-zero support via Auto Endpoints, backed by SGLang and a custom DFlash speculator tuned to the model's architecture. The new model improves substantially over Qwen 3.7 on coding, work, research, and long-horizon tasks, and the speculator was trained on data emphasizing tool-call-heavy sequences to boost token acceptance rates. Text-only access is offered for a limited one-month window.

1m read timeFrom modal.com
Post cover image
Table of contents
Try it now

Questions this post answers

What is Qwen3.8-2.4T-A95B and how does it compare to Qwen 3.7?

Qwen3.8-2.4T-A95B is a new open weights large language model from Alibaba that shows substantial improvement over Qwen 3.7 in coding, work, research, and long-horizon tasks. It is now available on Modal as an OpenAI-compatible Shared Endpoint with token-based pricing, using SGLang and a custom DFlash speculator for faster inference. Keep track of new open weights model releases like this one as they land on daily.dev.

How does Modal speed up inference for the Qwen3.8 model?

Modal uses a custom DFlash speculator model tuned specifically to Qwen3.8's architecture, paired with SGLang for serving. Because Qwen3.8 improved on coding, research, and work tasks that involve more tool calls, the speculator's training data emphasized tool-call-heavy sequences to increase token acceptance rates and boost speed. Developers evaluating inference speedups can follow speculative decoding techniques like this on daily.dev.

4 Impressions