Baseten is now a supported Inference Provider on the Hugging Face Hub, enabling serverless access to popular open-weight LLMs like DeepSeek V4 Flash, Kimi K3, and GLM-5.2. Developers can use Baseten-hosted models via the HF website UI, Python (huggingface_hub >= 1.26.1), JavaScript SDK, or agent harnesses like OpenCode and Hermes Agents. Two billing modes are available: direct requests billed to your Baseten account, or HF-routed requests billed to your HF account at standard provider rates with no markup. HF PRO users receive $2 in monthly inference credits usable across providers.
2 Impressions