An analysis of 10,643 AI-powered pull request reviews run through Kilo's platform between June 22 and July 23, 2026 reveals that open-weight models took two of the top three spots for surfacing critical findings, with Kimi K2.7 Code leading at 0.179 critical findings per review. Open-weight models accounted for 75% of tokens at roughly 16x lower cost per token than closed models. The security gap between open and closed models largely disappeared when a single outlier (GPT 5.6 Sol) was removed. Key takeaways: license category is a weak predictor of model behavior, using separate models for authoring and reviewing reduces blind spots, and 32.3% of reviews already used a different model than the one that wrote the code.
Table of contents
Open weights took two of the top three spotsModels don’t agree on what counts as criticalOne model created most of the security gapThree quarters of the tokens, a sixth of the costUse one model to write, another to reviewNo single model wins every taskThree rules for picking a reviewerWhat this data doesn’t proveWhy this matters past code review432 Impressions