Token consumption leaderboards at Meta and Disney revealed a flawed proxy for AI productivity — burning more tokens doesn't mean shipping more value. Engineering leaders from LaunchDarkly, Guild, and Arcade independently retired metrics like token spend, seat count, and code volume in favor of production outcomes: did the work ship, stay stable, and avoid rollback? A key bottleneck has emerged downstream: AI-generated PRs wait 5.25x longer for review, run 2.6x larger, and are accepted at less than half the rate of unassisted ones. Leaders seeing real gains had deployment discipline (feature flags, progressive rollouts, fast rollbacks) before AI arrived. The recommended replacement metric is work that reached production and stayed there — cycle time, rework rates, change-failure rate, and rollback frequency. Finance teams are already moving to cost caps if engineering can't produce an outcome model.