Inference-time scaling (ITS) improves language model reliability by generating multiple candidate outputs at runtime and selecting the best one — without retraining or fine-tuning. Red Hat AI's open source `its_hub` framework packages techniques like self-consistency, best-of-N, hierarchical voting, and particle filtering into a Python SDK and OpenAI-compatible HTTP API. Applied to Red Hat OpenShift Lightspeed's agentic tool-calling workflows, ITS improved pass rates from 10% to 20% across 11 troubleshooting scenarios, with the largest gains on the hardest multi-step tasks. The framework works with any OpenAI-compatible endpoint including vLLM, and Red Hat plans to integrate ITS into its AI gateway as an infrastructure-level capability.
Table of contents
What inference-time scaling isThe enterprise case for inference-time scalingHow its_hub brings ITS into practiceTechniques available in its_hubWhy enterprises should careGetting started with its_hubClosing thought53 Impressions