Learn AI

Measuring whether it worked

Three numbers, measured before and after. Without them you are guessing, and enthusiasm is not evidence.

AI generated for this lesson — machine-made illustration.

What you'll be able to do afterwards: Measure one AI workflow with three numbers and decide honestly whether to keep it.

Most teams never measure this, which is why most teams cannot say whether their AI spending works. It is not hard. It is three numbers, taken before and after.

Number one: time per task, including the review

Not the time to produce the output. The time from starting to the task being finished, including everything a person did.

This is where almost everyone fools themselves. A draft appearing in ten seconds feels like a saving of half an hour. It is only a saving if the reading, correcting and approving takes less than thirty minutes minus ten seconds.

Measure it with a timer on five real instances, before and after. Do not estimate. Estimates of task duration are reliably wrong, and wrong in the flattering direction.

Number two: how often it needs correcting

Of the last twenty outputs, how many did a person have to change?

A workflow that saves four minutes but needs correcting seven times in twenty is not saving four minutes. It is saving four minutes and creating a subtle new job: checking whether the machine was right this time.

There is a threshold that matters here. Below roughly one in ten, corrections are noise. Above roughly one in three, the workflow is not ready — either the instructions are too loose, or the task is not one of the five from lesson 3 and should not have been automated.

Number three: whether anyone is actually using it

Adoption is not a soft metric. It is the test of whether the thing was useful.

A workflow built in March and used twice in April has failed, whatever it measured. Either it did not fit how people work, or it was never trusted, or the person who wanted it was not the person who had to use it. All three are worth knowing.

Take the measurements properly

You need a before and an after, which means measuring before you build. This is the step everyone skips because they are keen to start, and it is the reason the ROI conversation is always vague later.

It does not need to be a project. Five tasks, timed, on a note in your phone. A count of corrections over one week. That is enough.

Then wait two weeks and do it again. Two weeks is long enough for novelty to wear off and short enough that you still remember the baseline.

Reading the result honestly

Clearly better — meaningfully less time, few corrections, people using it. Keep it, and look for the next task of the same shape.

Neutral — the time saved is roughly the time spent reviewing. This is more common than the enthusiastic case, and it is not a failure. The honest options are: improve the instructions, narrow the task, or stop. Neutral is not worth defending out of pride.

Worse — more total time, or heavy correction. Stop. You have learned cheaply that this task is not a fit, and the lesson transfers: it was probably a judgement task wearing a repetitive task's clothing.

The trap of output as a metric

The easiest thing to measure is volume, and it is the least useful. "We produced four times as many posts" tells you nothing unless the volume was the constraint.

Ask instead: did the thing that was actually slow get faster? If the bottleneck was always the review, or the client's decision, or your own judgement, then producing more drafts upstream just fills a queue and adds a backlog to manage.

Measure the bottleneck. Ignore the rest.

When to stop

Stop when any of these is true, and stop without embarrassment:

  • Corrections are frequent and not improving after two attempts at better instructions.
  • Nobody has used it in a month.
  • The cost of reviewing has quietly become someone's standing job.
  • The consequence of a mistake has grown since you built it.

Stopping a workflow that does not work is the cheapest lesson available, and teams that do it quickly end up with two things that work instead of six that half do.

Do this now

Pick your one live workflow. Time five instances of it this week. Write the number down. That single number is the beginning of knowing whether any of this is worth what you are spending.

Next: lesson 10, a plan for the first thirty days.

Read the next one first

One email a day

The day's consequential AI developments with the operational consequence stated, plus every price change we detect. Free, one send a day, one click to leave.

No third parties, no sponsored placements inside the brief, no list rental.

0 comments

No comments yet. If you have run any of this, that is the most useful thing you could add.

Add yours

Comments are read by a person before they appear. No sign-up, no account.