Skip to content
PKResources
Course outline

Agentic Engineering: the hands-on course · Module 6: Agentic engineering in teams

Measuring impact without vanity metrics

Lines generated and prompts sent tell you nothing. Measure outcomes that are hard to game, compare against a baseline, and pair the numbers with what your team says.

Lesson 23 / 24 · ⏱ 7 min

Sooner or later someone asks: is this agent thing actually working? The easy answer is a dashboard of activity. More code, more PRs, more prompts. It looks great and proves nothing.

Agents make activity cheap. So any metric that counts activity will go up whether or not you’re shipping better software.

Activity vs outcomes

measures activitymeasures outcomeseasyto gamehardto gamelines generatedprompts sent% of code “by AI”PRs merged per weekdeploy frequencylead time to productionrework rateescaped defectsreview wait time
Aim for the bottom right: outcomes that are hard to inflate. The top left is where vanity metrics live.

PRs merged sits in the top right on purpose. It’s closer to an outcome, but agents make it trivially easy to split work into more PRs. Watch it only next to rework and lead time.

Swap the vanity metric

Flip each card for what I’d measure instead.

🃏 Vanity metric, better question

Tap a card to flip it

A measurement plan that fits on a page

  1. 1

    Before

    Take a baseline

    Pull a few weeks of lead time, rework and review wait time from your tracker and repo history, before changing how the team works.

  2. 2

    Week 1

    Tag the work

    Mark agent-assisted PRs with a label or template field. Without it, you can't compare anything.

  3. 3

    Monthly

    Compare like with like

    Compare similar kinds of work: bug fixes to bug fixes, small features to small features. A mix shift can fake an improvement.

  4. 4

    Monthly

    Ask the team

    Five minutes, three questions: where did agents help, where did they cost you time, what would you stop doing? Numbers say what, people say why.

  5. 5

    Quarterly

    Decide something

    A metric nobody acts on is decoration. Each review should change a convention, a guardrail or where you use agents.

Before you move on

✅ Key takeaways

0 / 5 completed