H Company releases Holotron-12B, a 12B-parameter multimodal computer-use agent model post-trained from NVIDIA's Nemotron-Nano-2 VL. The model uses a hybrid State-Space Model (SSM) and attention architecture that avoids the quadratic cost of full attention, enabling a constant memory footprint per layer regardless of sequence length. On a single H100 GPU with vLLM, Holotron-12B achieves over 2x higher throughput than Holo2-8B, reaching 8.9k tokens/s at concurrency 100. On the WebVoyager benchmark, it scores 80.5% versus the base model's 35.1%. Trained in two stages on ~14B tokens with proprietary localization and navigation data, it also improves on OS-World-G, GroundUI, and WebClick benchmarks. The model is available on Hugging Face under an NVIDIA Open Model License, with plans to post-train on the upcoming Nemotron 3 Omni architecture.

4m read timeFrom huggingface.co
Post cover image
Table of contents
Holotron-12B - High Throughput Computer Use AgentWhy We Built Holotron-12BConclusionWhat’s next: Scaling the Future of Agentic Intelligence with Nemotron 3 Omni
49 Impressions