Skip to main content

L4: Specification-driven

You define what to build instead of building it: write the spec, set the tests, judge the results. Leave for twelve hours and come back to a green suite.

Check yourself.

One question per step. Take it cold to find where to start, or after reading to see what stuck. Nobody's grading you.

  1. Your agent keeps failing on a task. Where does the handbook say to look first?

  2. Your test suite is green. Why run evals on agent behavior too?

  3. How should prompts be managed once a feature is in production?

Where to next.

Climb to the next rung, or see how the other levels work.

  1. L0: Manual

    The ground floor: AI as a search engine and occasional tab-complete. No real productivity change.

    Everyone starts here. There's nothing to unlock yet.

  2. L1: Discrete task offloading

    You hand the AI small, self-contained jobs: write the unit tests, draft the docstring, generate one well-scoped function. The real work is still yours.

    Chat assistants · Prompting 101 · IDE assistants

    Explore path
  3. L2: Active pairing

    The AI writes alongside you all day, handling the boring parts while you steer and review everything as it lands. Shapiro estimates about 90% of AI-native developers live here.

    CLI agents · AGENTS.md & CLAUDE.md · What an agent actually is

    Explore path
  4. L3: Human-in-the-loop management

    You mostly stop typing code. Agents run several tasks at once and your day becomes reviewing their work. Most people never go past this level.

    Own the outer loop · Agentic code review · Role separation · MCP & tools

    Explore path
  5. L4: Specification-driven

    You define what to build instead of building it: write the spec, set the tests, judge the results. Leave for twelve hours and come back to a green suite.

    The harness as an artifact · Evals-as-tests · LLMOps & guardrails

    You're on this rung

  6. L5: The dark factory

    Agents building the software with humans out of the loop. Shapiro reports it working only for teams under five people, and calls it likely our future.

    Swarms & fleets · The factory model · Human factors

    Explore path