Writing · AI Engineering

Making AI's Work Countable

From commit signatures to LLM judgment — building AI-contribution measurement infrastructure


TL;DR

  • To answer "did AI actually help?" with numbers instead of gut feeling, we built a two-layer measurement infrastructure: every AI artifact gets an automatic signature, and an LLM judges the level of contribution.
  • Six months of real data: hundreds of AI-authored merged PRs, 1,000+ AI-driven tasks, request lead time cut to roughly a fifth.
  • One core principle — never depend on human memory. Hooks and CI enforce the signals, and probabilistic judgments get a human-verification safety net.
AI-authored merged PRsHundredssix months after adoption
AI-driven tasks handled1,000+lv2 and above
Request lead time (median)Cut to 1/5measured before and after
PRs by the team lead (me)Hundredsfirst half of 2026

1. The problem — "what AI did" is fundamentally uncountable

"Our team adopted AI." You hear it from every engineering org these days. Yet almost no org can answer the very next question with a number: "So, did it actually work?"

The moment you try to measure, you hit a wall. AI-generated commits and human commits are indistinguishable in git history. A PR that AI drafted and a human polished — whose is it? Surveys aren't the answer: they're one-shot, memory is unreliable, and they don't scale over time. What we wanted was numbers that accumulate on a dashboard every day.

Three requirements

  1. Deterministic — anything that relies on human memory or diligence will inevitably miss data
  2. Level-aware — a binary "used AI / didn't" carries almost no information
  3. Zero workflow changes — if measurement slows development down, you've defeated the purpose

2. The design — two layers: deterministic signals + probabilistic judgment

Layer 1 · Deterministic signatures (left by machines, automatically) Commit trailer auto-attached by prepare-commit-msg hook PR label CI greps commits → auto-label Doc signature label + machine-readable property Layer 2 · Judgment pipeline Automatic transition rules (deterministic) bot triage → lv1 · AI PR linked → promote to lv2 LLM judgment (probabilistic) re-judge resolved issues → lv3 / low confidence → unknown lv1 AI drafts + human finishes lv2 AI executes + human decides lv3 AI autonomous · zero intervention unknown human-verification queue
The full measurement infrastructure — the bottom layer leaves deterministic signals; the top layer interprets them as contribution levels

Layer 1 — an indelible signature on every artifact

Commits get a trailer:

Co-Authored-By: Claude <model> <noreply@anthropic.com>
AI-Run-ID: <session tracking value>
AI-Skill: <which automation skill produced this>

The key point: no human ever types this. A git prepare-commit-msg hook detects the AI tool's environment variables and attaches it automatically. For PRs, CI greps commit messages and applies labels automatically; documents (wiki pages) get the same signals — labels, properties, a signature box. Replacing "memory" with a system — that principle repeats through every design decision here.

Lesson from production — put signals in the canonical source

At first we only left the signal in the PR body, and it got lost on squash merge. Measurement signals must live in git history itself. Auxiliary metadata can vanish at any compression or summarization step.

Layer 2 — classifying the "level" of contribution

LevelDefinitionExample
lv1AI drafts + human reviews and finishesAI triages and guides an issue, a human handles it
lv2AI does most of the execution, human only decidesAI creates the PR, a human reviews and merges
lv3AI runs autonomously, zero human interventionauto-triage → auto-answer → auto-close

Classification is automatic too. The moment a bot triages an issue, it's lv1; the moment a PR carrying an AI signature is linked, it's auto-promoted to lv2; resolved issues are periodically re-judged by an LLM to determine lv3. Two design decisions that mattered:

  • Labels are exclusive — exactly one. Duplicate labels poison aggregation. On promotion, the old label is removed atomically.
  • Levels only move upward. If a probabilistic judge can move levels in both directions, the metric oscillates. On low confidence, instead of demoting, we tag it unknown and send it to a quarterly human-verification queue.

3. Three things we learned in production

  1. Automation's side effects get fixed with automation. Logic that auto-"resolved" issue keys found in PR bodies once closed an issue that was only mentioned as a reference. We fixed the parser and pinned it with a regression test — measurement infrastructure is production, too.
  2. Edge cases distort the metric. Bot PRs and release branches with no linked issue were blocking our validation gate. We defined the exceptions explicitly and let them pass — but also counted how often exceptions fired. Silent exceptions are the most dangerous kind.
  3. Measurement changes the culture — the real payoff. Once "measure before and after" became the default, it spread to everything else the team does. Bot impact is now argued with before/after lead times; a data migration with theoretical vs. measured throughput. The final deliverable isn't the numbers — it's a team that doesn't make claims without measurement.

4. Six months in numbers

The stat tiles at the top are the answer — hundreds of AI-authored merged PRs, 1,000+ AI-driven tasks, lead time cut to roughly a fifth, and me, an engineering leader 18 years in, shipping hundreds of PRs in half a year. The value of these numbers isn't their size — it's their provenance. Every one of them was aggregated from signals the system left automatically, and every one is reproducible with the same query, any time.

5. The minimal setup for teams that want to start

  1. Automated commit trailers — one prepare-commit-msg hook. Takes 30 minutes.
  2. PR-label CI — a reusable workflow that greps commit signatures and applies labels. Each repo just adds a 3-line caller.
  3. A level-definition doc — one page of team consensus on lv1/lv2/lv3. Automation comes after.

Measurement is not surveillance. On our team, these numbers have never once been used for individual performance reviews. They answer exactly one question — "is this change we adopted actually working?" The moment you can answer that with data, AI adoption stops being a fad and becomes engineering.