AI-assisted coding has broken the traditional correlation between whiteboard performance and on-the-job delivery, making most technical interview loops measure the AI model rather than the candidate. When any competent engineer with an agent can complete a standard CRUD exercise, the task stops differentiating talent. The real signal lies in the 20% of work agents can't handle: ambiguous requirements, planted wrong suggestions, verification habits, and judgment under pressure. The post argues for redesigning interview exercises so agents carry candidates 80% of the way, with scoring focused on behavior in that final 20%. It also makes an economic case: a $20,000 hiring loop protects against a $300,000 mis-hire, and paid work trials (~$6,000/candidate) are a cost-effective alternative that measures actual delivery. A calibration test is proposed: run the new 80/20 scoring alongside the old artefact-based scoring and see if rankings diverge.

7m read timeFrom codegood.co
Post cover image
Table of contents
The Word Doing All the WorkPricing the InstrumentDesigning for the Last 20%The Calibration Test
816 Impressions