All insights

Measurement

13 September 2026
6 min read

How to measure whether AI training worked: three numbers at week four

Every AI training session ends with a feedback form, and every feedback form comes back positive. People enjoyed the morning, the trainer was knowledgeable, they would recommend it to a colleague. Four weeks later the manager cannot say whether anyone is doing anything differently, and the finance director, who signed off the licences and the training, has stopped asking.

The fix is not a better feedback form. It is deciding what "worked" means before the session, in numbers small enough to collect without a project, and then actually collecting them. We use three. This article explains what they are, why we chose them over the dashboards, and how to run the review that makes them useful.

What the usual measures miss

Vendor dashboards report active users. The Copilot Dashboard in Viva Insights, the admin analytics in ChatGPT Enterprise and the Gemini usage reports in the Google Admin console all count who touched the tool in a period. That is worth having, and it is not the same as value. A person who asks the assistant one idle question a week is active. A person who has rebuilt how they prepare the monthly board pack is also active. The dashboard cannot tell them apart.

Satisfaction scores measure the session, not the month after it. Training evaluation has known this since Donald Kirkpatrick set out his four levels in the 1950s: reaction, learning, behaviour, results. Most AI training is measured at level one and never gets to level three, which is the only level a manager cares about. Did behaviour change, and did it stay changed.

Hours-saved surveys, on their own, drift upwards. People round in the direction they think you want. The 2023 field experiment at BCG by Dell’Acqua and colleagues, Navigating the Jagged Technological Frontier, is a useful corrective: 758 consultants using AI completed 12.2% more tasks and finished 25.1% faster on tasks inside the tool’s competence, and were 19% less likely to be correct on a task designed to sit outside it. Speed is real. So is the failure mode. A measure that only counts speed will not see the second half.

The three numbers

  1. 1

    Minutes saved per person per week, on named tasks

    Not "how much time does AI save you". Instead: for the five tasks we built workflows for on the day, how long did each take this week, and how long did it take before. Named tasks anchor the estimate and make it comparable across people. Accept that it is self-reported and soft; it is still the number the finance director will ask for, and anchoring it makes it honest enough to use.

  2. 2

    Workflows still in weekly use

    Count the workflows built on the training day, then count how many were used at least once this week by at least one person. This is the hard number for behaviour. If eight were built and two survive, you have learned which two mattered and which six were the trainer’s idea rather than the team’s. Both halves of that are worth knowing.

  3. 3

    People who used the tool for a real task this week

    Not licensed, not logged in: used it to do a piece of work that went somewhere. Ask the question directly in the team meeting or a two-line poll. This is the adoption number, and defining "real task" up front is what stops it inflating.

Set the baseline before the session, not after

The most common measurement mistake is to decide the measures after the training and then try to remember what things were like before. Do it in the survey that precedes the session. For each person: the five tasks they repeat most, roughly how long each takes now, and whether they have used the assistant for anything in the last fortnight. That survey is doing two jobs, shaping the exercises and setting the baseline, and it takes ten minutes per person.

Turn the pre-session survey into a baseline tableAny assistant
Below are survey responses from [number] people on my team. Each person listed the tasks they repeat most, an estimate of how long each takes per week, and whether they have used [assistant] in the last two weeks.

Build a baseline table with one row per named task: the task in plain words, how many people do it, the median weekly minutes across those people, and the range. Then give me one line: the total weekly minutes across the team on these tasks, and the share of people who have used the assistant recently. Do not add tasks nobody mentioned. Flag any estimate that looks like an outlier so I can check it with the person.

[Paste responses]

Good output is a table you can put in a spreadsheet and revisit at week four with the same headings. The outlier flags catch the person who wrote "20 hours" when they meant two.

The week-four review, in thirty minutes

Four weeks is long enough for habits to form or fail, and short enough that people still remember the day. Book the review before the training happens so it cannot slide. The agenda is fixed.

  • Five minutes: the three numbers, against the baseline, on one slide or one sheet. No commentary yet.
  • Ten minutes: the workflows that fell away. For each, one question: was it the tool, the task, or the setup? A workflow that died because SharePoint permissions blocked it is an IT ticket, not a training failure.
  • Ten minutes: the workflows that stuck. Who is using them, what they changed, what they would show a colleague. These become the next entries in the team library and the material for the second cohort.
  • Five minutes: decide. Extend to the next team, fix what blocked the first, or stop. Write the decision down with the numbers beside it.
Prepare the review from the week-four pollAny assistant
Here is our baseline table from before the training and the responses to this week’s poll, which asked each person: which of the [number] workflows they used this week, how long the five named tasks took, and whether they used [assistant] for a real piece of work.

Produce: (1) the three numbers, minutes saved per person per week on the named tasks, workflows used this week out of [number], and people who used the tool for a real task, each shown against the baseline; (2) a short list of workflows nobody used, with the reasons people gave grouped as tool, task or setup; (3) two or three quotes or examples from the responses that show a workflow working, in the person’s own words. Keep the whole thing under one page. Do not editorialise.

[Paste baseline table and poll responses]

Good output is the whole review pack. If it tells you something is a "success", ignore that and look at the numbers; the judgement is yours and the room’s.

Talking to the finance director

Finance directors are not hostile to AI spend. They are hostile to spend nobody can account for. Minutes saved, multiplied by people, multiplied by a loaded hourly rate, gives a number they can compare with the licence and training cost. Present it honestly: self-reported, anchored to named tasks, collected at week four. Put the two hard numbers beside it. A team where seven of eight workflows are still in use and everyone used the tool for real work this week is a team where the soft number is probably true.

Resist the temptation to project. A month of data from one team is a month of data from one team. The case for extending to the next team is that the first one worked, not that the whole organisation will save a thousand hours by Christmas. The projection will be wrong and it will be remembered.

Questions people ask

Agree three numbers before the session: minutes saved per person per week on named tasks, workflows from the day still in weekly use, and people who used the tool for a real task this week. Set the baseline in the pre-session survey and review at week four. Minutes saved times people times a loaded hourly rate gives finance a comparable figure; the two hard numbers tell you whether to believe it.

Active-user percentages from vendor dashboards are a weak guide because they count idle use. A better test is how many people used the tool for a piece of real work this week and how many of the trained workflows are still in use. A team where most workflows survive a month and everyone used the tool for real work is adopting; a team with 90% "active users" and no surviving workflows is not.

Four weeks after the session, in a thirty-minute review booked before the training takes place. That is long enough for habits to form or fail and short enough for people to remember the day. Review again at twelve weeks before deciding on the next cohort.

Further reading, all free

Not ready for a call

Tell us what you are trying to do.

A few lines is enough. We will reply by email with whether we can help and roughly what it would take. If a call would answer it faster, we will say so. Nothing to commit to either way.

Or email [email protected]

Used to reply to you. See our privacy notice.