Weibo AI has released VibeThinker 3B, a compact 3-billion parameter model that claims to outperform much larger models like Kimi 2.5, Claude Opus 4.5, and GLM 5 on certain tasks. The model was trained using synthetic data and reinforcement learning from verifiable rewards (RLVR). A local demo on a Dell Pro Max with RTX 6000 shows ~180 tokens/second at full resolution, demonstrating strong reasoning capability for its size. A full paper detailing the training methodology is available.

1m watch time
12 Impressions