You're reading The AI Directive - practical AI intelligence for business leaders. Subscribe here if someone forwarded this.

THE SIGNAL

What 90 Days Of Using AI Properly Has Actually Shown

Over the last 90 days, alongside the day job, I have been using evenings and weekends to run a personal AI lab: testing how AI behaves when it is put near real business work.

I want to be clear about the setting, because it changes how this should be read. This is not a corporate case study or a write-up of an internal company AI programme.

The point of the lab is different: to do the awkward, practical testing on behalf of business leaders who want the lessons without having to build the whole experiment themselves.

Three months is long enough for the novelty to wear off, which is when AI gets more interesting.

The first week is full of easy wins. A draft appears quickly. A messy document becomes readable. A comparison that would normally take an hour comes back in minutes. A research trail appears before the meeting has even been booked.

Those wins are real. The better test comes later, when AI is asked to help with work that has consequences: recurring reports, content pipelines, operational checks, decision packs, approval prompts, follow-ups and tasks that have to survive interruption.

That is what I have been testing in the lab: not "can AI produce a good answer once?", but "can it support repeatable work without quietly losing the plot?"

After 90 days, the lesson is not "AI can do everything" or "AI cannot be trusted". Both are lazy positions.

The practical lesson is narrower: AI is very good at preparing work, and much weaker as an owner of work.

If a business gets that distinction right, AI reduces drag. If it gets it wrong, AI produces confident movement without enough control around the outcome.

Here is what I would take from the first 90 days.

WHAT WORKED

1. Narrow jobs worked better than broad instructions

The strongest results came from specific jobs with a clear output.

Useful examples:

  • turn this note into a first draft;

  • compare these two versions and show the material differences;

  • check whether this article draft matches the published source;

  • summarise the status of this workflow from the tracker and reports;

  • prepare a decision pack from these known inputs.

The weaker instructions were the broad ones:

  • manage the newsletter;

  • sort the pipeline;

  • keep this moving;

  • find the best opportunity;

  • handle this every week.

Those sound efficient, but they hide too much judgement. The agent can appear busy while the real business outcome remains unfinished.

My recommendation after 90 days in the lab: make AI jobs smaller than your instinct suggests. The first version of a workflow should be easy for a person to inspect and explain.

2. AI saved time when it shortened the route to a human decision

The best work did not remove the human. It got the human to the right decision point faster.

That showed up in a few patterns:

  • first drafts for review;

  • status summaries from messy source material;

  • comparison checks;

  • exception lists;

  • shortlist preparation;

  • evidence packs before approval.

This is where AI is genuinely useful for leaders. It reduces the amount of low-value preparation work between "we need to decide" and "here is the thing we need to decide".

The mistake is to assume that because AI prepared the material, AI has owned the decision. The person still owns whether to send, publish, buy, approve, reject, pause or escalate.

3. Reviewable outputs built trust

Trust did not come from the model sounding impressive.

It came from outputs that could be checked quickly.

A good AI output made it clear:

  • what it was based on;

  • what it produced;

  • what changed;

  • what still needed a person;

  • what action was being requested.

The more reviewable the output, the more useful the workflow became.

This is why first drafts, comparison work and exception summaries are strong early AI uses. The reviewer can tell whether the work is good enough.

WHAT WAS PROBLEMATIC

1. Chat promises were too easy to lose

The weakest pattern was work that lived only in conversation.

Someone asks for a follow-up. The assistant says it will happen. The chat moves on. Later, nobody can easily tell whether the work was done, failed, stalled, or merely discussed.

This is the old business problem of "I'll sort that", with better vocabulary and more confidence.

AI makes it worse because the reply can sound complete before the underlying work exists.

The fix is simple: work that matters needs a durable place to live. A draft. A checklist. A tracker row. A calendar entry. A report. Something outside the conversation.

2. Recurring jobs looked healthy when the business result was missing

This was the sharpest operational lesson from the lab work.

A scheduled task can run. A reminder can fire. A script can exit cleanly. A message can say the right thing.

None of that proves the intended result happened.

For recurring AI work, the test cannot be "did the automation run?" The test has to be "did the expected business output exist afterwards?"

If the job was meant to create a draft, the draft has to exist.

If it was meant to prompt an approval, the prompt has to arrive.

If it was meant to publish only after approval, it has to stop cleanly until approval exists.

That is the difference between an automation that reassures you and one that actually works.

3. AI language drift has to be edited out

AI has a house style if you let it.

Short tidy paragraphs. Balanced contrasts. Phrases that sound sensible but do not give the reader much to do. A habit of turning real experience into abstract advice.

That is a problem in leadership content because senior readers want the point, the implication and the action. If the writing does not change what they would do, it is just noise with better formatting.

The edit has to be harder than "make it sound more human".

For this newsletter, the test should be:

  • would I actually say this to a senior peer?

  • does this sentence change what the reader would do?

  • is this specific enough to be useful next week?

  • have I made a recommendation, or only described a principle?

If a paragraph fails those tests, it should be cut or rewritten until it earns its place.

THE LESSONS

Lesson 1: Start with one real business workflow

Do not start with "we need agents".

Start with a business workflow that already causes friction:

  • drafts are slow;

  • reviews get missed;

  • information sits in too many places;

  • people waste time preparing the same updates;

  • decisions wait because nobody has pulled the evidence together;

  • recurring checks depend on memory.

Then decide what AI should do inside that workflow.

The right first role is usually preparation, comparison or exception-finding. Not full ownership.

Lesson 2: Separate preparation from permission

This is the cleanest rule from the 90 days of testing.

AI can prepare more than it should be allowed to execute.

Let it draft the email. Keep the send human-approved.

Let it prepare the public post. Keep the publish button human-approved.

Let it build the shortlist. Keep the commercial decision human-approved.

Let it inspect the process. Keep the process change human-approved.

That rule keeps the speed without handing over judgement too early.

Lesson 3: Judge AI workflows by the handoff

The best question is not "did AI do the task?"

The better question is "did AI make the next human action easier?"

Good handoffs are obvious. The reviewer can see what happened, what matters and what decision is needed.

Bad handoffs create more checking, more doubt and more work for the person who was supposed to benefit.

That is where a lot of AI projects quietly fail. The output exists, but the next action is still muddy.

Lesson 4: Do notautomate around a weak process

AI exposes weak processes quickly.

If the business has no owner, no source of truth and no decision rule, an agent will not magically create discipline. It will just move the confusion faster.

Before adding automation, write five lines:

  1. What starts the work?

  2. What sources are allowed?

  3. What must exist at the end?

  4. Who approves the next step?

  5. What stops the work?

If those five lines are hard to write, the workflow is not ready for autonomy.

THE SHORTLIST

1. AI works best first as preparation, comparison and exception-finding.

2. Broad instructions create movement without enough accountability.

3. Recurring jobs should be judged by the business output, not the run log.

4. Human approval still belongs around external action, money, customers, suppliers and public claims.

5. AI writing needs a harder edit than most teams give it.

ONE THING

Pick one recurring workflow and write the five-line version:

trigger, sources, output, approval, stop point.

If that is clear, AI can probably help.

If it is not clear, fix the workflow first.

FROM THE EDITOR

The 90-day lab lesson is not that AI needs more ceremony. It needs better jobs: small enough to check, useful enough to matter, and clear enough that a person knows what decision is being asked of them.

See you Tuesday.

  • Toby

Further reading

  • NIST - AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework

  • McKinsey - The agentic organization: https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-agentic-organization-contours-of-the-next-paradigm-for-the-ai-era

  • Harvard Data Science Review / MIT Press - AI agents andhuman-in-the-loop decision making: https://hdsr.mitpress.mit.edu/pub/fdzqkh85

TOOLS I USE & RECOMMEND

These are tools I use personally. Affiliate links marked - I earn a small commission if you sign up, at no extra cost to you.

  • ElevenLabs - AI voice generation. I use this for scripted narration and YouTube production. (affiliate)

  • HeyGen - AI video avatars. I use this for structured video content and repeatable production. (affiliate)

  • beehiiv- The platform this newsletter runs on. If you're starting a seriousnewsletter, this is the stack I'd use again. (affiliate)

Some links in this issue are affiliate links. I only recommend tools I actually use.

Keep reading