How to measure AI leverage on your team
AI leverage is measurable, but not by counting licenses or asking people if they use AI. Measure four things per person and per function: a throughput baseline against current output, workflow coverage (what share of recurring work runs through documented AI workflows), verification rate (whether shipped output passed a stated check), and evidence quality. Run the baseline over two to four weeks, place each person on a five-level maturity scale, and treat the spread between your most and least leveraged people as the size of your opportunity.
Most teams cannot answer a simple question: is all this AI actually making us faster, and by how much? Licenses were bought, a lunch-and-learn happened, everyone nods that they use the tools. Meanwhile throughput looks suspiciously like last year’s. If you lead a team, here is the direct answer to how you measure it: track output multiples against a baseline, per person and per function, and audit the workflows and verification behind them. Leverage is a measurable property of how work is engineered, not a vibe about tool adoption.
Why the obvious metrics fail
- License counts and login stats measure procurement and curiosity. A team can be 100 percent “active on AI tools” with zero change in shipped output.
- Self-reported usage surveys measure enthusiasm and social desirability. Everyone says yes now.
- Hours saved estimates are folklore unless anchored to a baseline; people are poor estimators of their own time, in both directions.
The pattern behind all three failures: they measure inputs. Leverage only shows up in outputs and in the structure of the work itself.
The four metrics that work
1. Throughput multiple (against a real baseline)
The anchor metric. For each function, define its unit of shipped work at a stated quality bar: campaigns launched, articles published to standard, sequences live, tickets resolved with quality sampling, dashboards delivered, reports out the door.
- Baseline: units per person per week, measured or reconstructed from the last normal quarter.
- Current: the same units, same quality bar, measured over two to four weeks of normal work.
- Report per deliverable class, not as one number. Honest measurement finds uneven compression: research and first drafts often collapse in cost, judgment-heavy negotiation barely moves. The per-lane detail is what you can act on.
2. Workflow coverage
What fraction of the function’s recurring work runs through a documented, repeatable AI workflow rather than ad hoc effort (with or without a chatbot open in another tab)? Inventory the recurring deliverables, then ask each owner to show the workflow: steps, tools, prompts or agents, quality gates. “Show” is the operative word; a workflow that cannot be shown does not exist organizationally, and it leaves when its owner does. Teams discover their real coverage number this way, and it is usually far lower than the tool-adoption number.
3. Verification rate
Of the AI-assisted output that shipped, how much passed a stated verification step, and what did verification catch? This is the risk side of leverage: unverified AI throughput is speed borrowed against a future quality incident, whether that is a fabricated statistic in a customer-facing document or a confident wrong number in a board deck. Spot-audit shipped work: ask what was checked before this went out and what the check caught this month. A named checklist per workflow is mature; “the person read it over” is not a verification step.
4. Evidence quality
Can each person demonstrate their leverage with artifacts: real deliverables, dated, with the workflow that produced them? This metric keeps the other three honest, and it doubles as your hiring and review instrument. It is the same standard we hold certified operators to (throughput evidence is competency four of the Operator Standard): claims with artifacts, or they are claims.
The maturity scale
Place each person, then each function, on this five-level scale. It converts measurements into decisions:
| Level | Name | What you observe |
|---|---|---|
| 0 | Unaided | Traditional methods; AI absent or incidental |
| 1 | Casual use | Occasional chatbot use; no measurable output change |
| 2 | Habitual use | Daily tool use, real but unmeasured gains; nothing documented, nothing verified systematically |
| 3 | Engineered | Documented workflows on core deliverables; verification steps exist; measurable multiple on some lanes |
| 4 | Operator | Owned stack, workflow coverage across the function, verification rate near 100 percent, evidenced multiple; leverage that would survive handover |
Most teams that “adopted AI” in the last two years sit at level 2, which is precisely why their throughput charts look flat: the gap between level 2 and level 3 is where the multiplier lives, and it is an engineering gap, not an enthusiasm gap. Note also the spread: nearly every team has one quiet level 3 or 4. Find them first; their workflows are your cheapest training material.
Running the measurement: a two-week minimum plan
- Week 0: define units of output and quality bars per function; reconstruct baselines from the last normal quarter; tell the team what is being measured and why (leverage, not surveillance; this matters and should be said plainly).
- Weeks 1 and 2 (to 4): count shipped units; collect workflow inventories by demonstration; spot-audit verification on shipped work.
- Close: score each person on the maturity scale; write per-function findings: throughput multiple by lane, coverage percentage, verification rate, and the spread between the most and least leveraged people.
What the findings drive: promote the level 3 and 4 workflows to team standards, train the level 1 and 2 majority against them, and set a re-measurement date, because a one-time audit is a snapshot and leverage compounds (or decays) quarterly.
If you would rather have your team taught
Everything above is runnable in-house, and this guide is meant to be sufficient. If you would rather have your people taught to do it, the AI course for business trains your team to build and verify AI workflows to the Operator Standard themselves, privately, on the kind of work they already do, so the capability stays with the team after the program ends. Either way, measure first. A team that knows its numbers stops arguing about whether AI is working and starts deciding where to aim it. See the AI course for business.