A head-to-head comparison of Claude Opus 4.8 (at four reasoning levels) and MiniMax M3 on a code audit task using a TypeScript/Bun/SQLite webhook delivery service with 17 known bugs. MiniMax M3 found 13 of 17 issues for ~$0.07, matching Claude Opus 4.8 at medium and high settings but costing over 18x less. Claude Opus 4.8 at xhigh found the most issues (15/17) for $2.03 and was the best value among Claude runs. The max setting cost $3.39, found the same 15 issues as xhigh, and dropped one finding. Higher reasoning levels shifted attention rather than uniformly improving coverage, and the most expensive setting was not the most useful.
Table of contents
PricingOur Setup for This ExperimentWhat Each Run FoundCostTimeHow Claude Opus 4.8 Scaled With EffortConclusion15.4K Impressions3 Comments