You hire someone new. Within two weeks you know whether they’re pulling their weight. You don’t need a scorecard or a consultant. The phones get answered faster, or they don’t. The backlog shrinks, or it doesn’t. Your evenings get shorter, or they don’t.

Measuring AI should be exactly that simple. Most people make it complicated instead. Dashboards, twelve metrics, reports about reports, and six months later still can’t answer the only question that matters: is this helping?

The whole method

One metric. Thirty days.

Pick the single number that matters for the task you handed over. If AI is handling customer emails, that’s reply time. If it’s producing quotes, it’s how many go out a week. If it’s sorting incoming orders, it’s errors per hundred.

What that looks like

A small accounting firm started using AI to draft client update emails. Before: fifteen updates a week, twenty minutes each to write. They picked one number, emails sent per week, and left it alone for a month.

After thirty days: twenty-two updates a week, five minutes each to finalise. Roughly fifty percent more going out, at a quarter of the time apiece.

They didn’t need a dashboard. They needed one number and a calendar reminder.

Pick the metric before you start

Not after. If you wait until the end, you’ll pick whatever happens to look best, and you’ll have learned nothing. Decide in advance what helping would mean (opens in a new tab). Write it on a sticky note and put it where you’ll see it.

At thirty days the answer is yes, no, or not sure. Yes means keep going. No means cancel it and try something else. Not sure means one more month with a sharper number.

You wouldn’t keep an employee a year without checking their work. Same standard here.