Activity metrics like token usage, accepted completions, and generated lines of code prove AI tools are being used but not that they're delivering value. Platform teams should instead measure engineering outcomes: workflow success rates, pull request review burden, guardrail failures, cost per accepted outcome, and developer confidence. The post proposes a structured measurement model broken down by team, repository, workflow, and model — with task-specific success definitions (e.g., a Terraform provider upgrade has different success criteria than an incident summary). Key warnings include avoiding vanity dashboards, not treating completed agent runs as successful outcomes, and recognizing that review effort is part of the real cost of AI-generated code.

Table of contents
Measure engineering workflows, not individual toolsA measurement model to get goingToken usage needs an outcomeReview burden is part of the costTreat guardrail failures as platform feedbackMeasure whether reusable context worksKeep developer confidence in the modelCost per accepted outcome is useful, within limitsDo not make activity the headlineWhat I would build firstKnown limitationsFrequently asked questionsMeasure what helps you improve the system41.7K Impressions2 Comments