1
Nh
net human
Why Net Human Our work Insights
What we do
About Talk to us
Home Why Net Human Our work Insights FAQ
WHAT WE DO
Consultancy Certification The Atlas The Institute How certification works The Guardrails EU AI Act readiness
About Tamsin TALK TO US
← INSIGHTS
PRACTICE · THE AUDIT

What an AI deployment audit actually looks at

Almost none of it is the model. The interesting evidence is in rotas, ticket queues, appraisal forms and the question of who is allowed to say no.

TAMSIN DEASEY-WEINSTEIN12 AUGUST 20267 MIN READ

People expect a model audit: weights, bias metrics, red-team results, a note about hallucination rates. Those matter, and plenty of firms do them well. They also answer a question almost nobody at board level is actually asking, which is why the resulting document tends to be filed rather than used.

The question being asked is simpler and harder. We put AI into this process. What happened to the people?

One: what the deployment was optimised for

Every deployment has an objective function, whether or not anyone wrote it down. Sometimes it is in the vendor contract, sometimes in the success metric on the slide, sometimes only in what the team gets praised for. We find it and read it back to you in plain words.

This is usually the moment the room goes quiet. A deployment sold internally as freeing clinicians from admin turns out to be measured on appointments per hour, which is a different thing pointed in a different direction.

A technology does what it is measured on, not what it was announced as.

Two: the Human Return baseline

Before you can say a deployment gave something back, you need to know what was there before. The baseline covers four things we can actually count: hours, skill, human contact, and agency.

  • Hours. Not hours saved — hours received. Where did the time go? Into learning, care, rest, seeing people? Or straight back into throughput, in which case the saving belongs to the organisation and the staff got a busier day.
  • Skill. Which capabilities has the system taken over, and are the humans meant to oversee it still practising them? Delegation without atrophy is a design requirement, not a hope.
  • Human contact. How many conversations, visits and relationships existed before, and how many exist now? For anything sitting between people, this is the metric — not engagement.
  • Agency. When the system is wrong about somebody's money, health, grade or livelihood, who can reach a person, and can that person actually overrule it?

Those four become the net human score: one figure for what the deployment gives back minus what it takes. It is not a survey result. It is calculated from evidence and verified, because a self-declared number is a marketing asset rather than a governance one.

Three: the eight Guardrails, with evidence

Each Guardrail names the evidence it requires, so the assessment is a document review rather than an opinion. What we ask for is mostly unglamorous and already exists.

WHAT WE ASK TO SEE
01
The business case as it was signed off, including the version with only savings on it.
02
The metric your teams are scored on, and the appraisal form that carries it.
03
The escalation path a real customer or citizen follows to reach a human, timed.
04
The disclosure a first-time user sees, in the words they see it.
05
Rotas, queue data or caseload figures from before and after the deployment.
06
The training plan for the people expected to supervise the system.

If a Guardrail cannot be evidenced, we say so and write down what would satisfy it. An audit that produces nothing failable is a brochure.

Four: what the board gets

A short document a non-technical director can read in ten minutes: the net human score with its workings, the Guardrails met and unmet, the EU AI Act Article 50 position, and a small number of changes ranked by how much they move the score. No maturity matrix, no heat map with forty amber squares.

The findings are usually cheaper to act on than people fear. The expensive discovery is not a broken model — it is a metric aimed at the wrong thing, which costs a sentence to change and eighteen months of trust to recover from.

What it is not

It is not a certification, though the same assessment sits underneath one. It is not a compliance sign-off, though it will tell you where you stand on disclosure and human oversight. And it is not an argument for deploying less AI. Most of the organisations we work with are deploying more of it afterwards, with the ROI and the human goal written on the same page.

Scope, timetable and cost are agreed in writing before an engagement begins — they depend on how many deployments are in scope and how much evidence already exists.

IF THIS IS YOUR PROBLEM

Bring one deployment. We will show you what the evidence says people got back.

READ NEXT