Skip to content

BlogAI assurance

A guide to what is, and might soon be, on the assurance professional's plate

Four kinds of system an audit or assurance function may have to assure, from people and process to autonomous AI agents, and what each one adds to the job.

Author
Henry Zhang, Co-founder, Kaption
Published
Reading time
8 min read

Over the past few months, one topic has come up again and again in my conversations with audit and assurance professionals: how do we deliver AI assurance? I have heard it on conference stages, in LinkedIn threads, and in the quieter chats over coffee afterwards.

The interest is real, and so is the confusion. "AI", "automation", "agents" and "models" are often used as if they mean the same thing. They don't, and the difference matters for how we assure them.

Whether governing AI agents ends up on the assurance profession's plate is still an open question. I am not going to predict the answer. What I can do is set out, in a structured way, what could be coming our way over the next few years.

This piece is a simple starting point. I hope to minimise the jargon and share a clear picture of what might land on the assurance professional's plate.

Four things that might be on your plate

In recent years, new things have been added to what has always been on the assurance professional's plate. Today there could be three or four:

  1. People and process
  2. IT systems
  3. Traditional AI and machine learning
  4. Emerging, autonomous agentic systems

One simple question separates them: who does the work, and who makes the decision? Keep that question in mind as we go through each one.

Four boxes in a row: people and process, IT systems, traditional AI, agentic systems. Beneath each, the new risk area it adds: process and review, logic and access, output quality, actions and authority.
Four types of system, and what each adds to assure.

1. People and process

People do the work. People decide.

This is where the profession started. A finance team approves invoices, a manager signs off a reconciliation, a committee reviews a policy. Assurance looks at whether the process is well designed and whether people actually follow it. This has been on the assurance professional's plate since the very beginning, so I won't go into it further here.

What we look after: segregation of duties, authorisation, evidence that reviews happened.

2. IT systems (deterministic)

Systems record and process. People still do the work and decide.

These are the systems we know well: the ERP, the CRM, the ticketing tool. They store business data and follow rules that an engineer wrote, for example "if an invoice is over £10,000, send it for a second approval". Given the same input, they give the same output every time. This is called deterministic.

This makes assurance relatively straightforward. The output is driven by fixed logic, so once we have controlled the logic, we have controlled the output.

With the output under control in this way, assurance comes down to three things:

  • How the logic is designed. Are the rules right in the first place?
  • How the system is implemented. Was it set up, configured and connected as intended?
  • How people use it. Who has access, could someone misuse it, and is the data protected as it moves around the wider tech estate?

What we look after: IT general controls, such as user access, change management and data protection. The profession has decades of practice here as well.

3. Traditional AI and machine learning (probabilistic)

The system can now make suggestions. A person still decides.

Here, AI is applied to a single, contained task. Someone provides data and a prompt, the model gives an output, and a person decides what to do with it. For example:

  • a control testing engine that assesses evidence against a SOX control
  • a planning assistant that drafts an audit plan from a prompt
  • an anomaly detection engine that flags possible fraud or financial crime

Flow from input, to model, to output, to a person who reviews, to action. Risk sits mainly in whether the output is right. Nothing happens until a person decides.
How traditional AI behaves: a person decides before anything happens.

An IT system is driven by rules: a person decides the logic and writes it into code. An AI or machine learning system is driven by data: it is shown many past examples and works out the patterns for itself.

Take a familiar example: spotting duplicate payments. A rules-based IT system might say "flag any two payments with the same supplier, amount and date". It does exactly that, every time. But it misses anything that is not an exact match: the same invoice paid to "ABC Ltd" and "A.B.C. Limited".

A traditional AI or machine learning model works differently. It is trained on vast amounts of transaction data, including payments that turned out to be duplicates. From that, it learns what duplicates tend to look like, so it can catch these near-misses too.

That makes it more capable. But when it flags a payment, there is no single rule you can point to. The honest answer to "why?" is "because it looks like something it has seen before". That is why these systems are often called probabilistic. It is also why:

  • ask the same question twice and you may get two different answers
  • it cannot always show you where an answer came from

This changes how we assure it. As a recent guide from Palo Alto Networks puts it, standard AI governance is more about output risk: whether a response is "accurate, fair, or compliant". The controls it describes centre on the following:

  • Training data. Is the data the model learned from of good quality and fit for the task?
  • Bias. Is the model checked for unfair patterns it may have picked up from past data?
  • Explainability. Can we understand, at least broadly, why the model gave its answer?
  • Review of the output. Does a person check what the model produces before anyone acts on it, and is the model's performance reviewed regularly?

What we look after: everything from the IT system, plus the quality of the output and of the human review.

One point worth making clear: when people raise concerns about governing AI, this is what many think about. However, the model works on an isolated task and a person stays in control. The risk sits mainly in the output, and it is relatively contained compared with what comes next.

4. Agentic systems

The system can decide and act autonomously.

This is what many recent discussions in the tech community are about. An agent does not simply answer a question. It plans a task, calls tools, moves across systems and takes action, often across many steps and without a person approving every step.

These are not yet common in our profession, but they are easy to picture. Take audit issues. An agent could pick up findings from a report, send each one to the right owner, set up follow-up meetings, track remediation updates and draft the status report for the audit committee in one sequence.

A person sets the goal. An agent plans and decides each step, choosing between allowed actions such as sending findings, booking meetings and tracking fixes, through tools and APIs into external systems like email, calendar and an issue tracker.
How an agentic system behaves: the agent decides.

This is clearly far more capable than traditional AI. But it also challenges the safeguard we relied on before: a human decision on every output. The risk is no longer only "is the answer right?" It becomes "what is this system allowed to do, and who answers for what it did?" Some recent writing on the topic calls this the move from output risk to action risk.

A white paper from Akamai sets out how different the governance task becomes. In short:

Traditional AIAgentic AI
Who starts itA person, each timeThe agent plans and runs multi-step tasks itself
What it doesProduces one outputTakes actions across several systems
What data it usesA defined setLive business data, apps and outside sources, as it goes
Main risksAccuracy, bias, fairnessData leakage, unauthorised actions, excess permissions, manipulated inputs
Who is accountableThe model ownerSpread across developers, deployers and connected systems
How it is monitoredPeriodic reviewsContinuously, in real time

What we look after: everything from IT systems and traditional AI, plus the new risks highlighted:

  • Acting without sign-off. Once given a goal, an agent can keep working through steps without a person approving them.
  • Using the wrong tools. An agent picks its tools as it goes, and can reach ones it was never meant to use.
  • Too much access. Agents often borrow system credentials and can end up with more authority than they need.
  • Data on the move. Sensitive information can pass between systems without anyone clearly seeing it.
  • Agents affecting agents. When agents work together, small errors can combine into bigger problems.
  • Blurred accountability. Responsibility can spread across the model provider, the platform and the organisation using it.
  • Drift. Over time, what an agent does can move away from what it was set up to do, and its authority can quietly grow.

Who holds responsibility is still an open question, but it is worth setting out the current thinking. One view is that whoever deploys an agent should own it, since they decide what it is allowed to do. Another view is that a central team of specialists should take this on, because few business owners will have the expertise to oversee an agent alone. Which model wins out is yet to be seen.

Beyond the systems: what to anticipate

Sorting systems this way is a useful start. But it also raises questions that go beyond the technology itself.

  • Everything at once. Most organisations already run at least three side by side, sometimes within a single process. Each part needs a different kind of assurance. How many audit plans are built to see that?
  • Systems that change category quietly. An AI notetaker is a traditional AI tool. Give it a feature that sends follow-up emails and books meetings on its own, and it becomes an agentic system. So who in the organisation is checking?
  • Where it sits. Is overseeing agents a first-line responsibility, a second-line risk function, or something internal audit should own? Most organisations have not yet decided.

None of these have settled answers yet. They are worth asking now, while agents are still early in our profession.

Why this matters to us at Kaption

At Kaption, we think about these questions for two reasons, even though they go beyond what most of our clients are asking us today.

First, we build and deploy both kinds of systems ourselves, traditional AI and agentic. Understanding their risks is not optional for us.

Second, we see ourselves as the AI partner for audit and assurance professionals. Our focus today is practical: giving them the capacity, coverage and timely insight to spend less time on groundwork. If governing agents does become a key area, we want to be ready to work alongside organisations, and to build what professionals need to keep these systems under control.

Questions this article answers

What is the difference between traditional AI and agentic AI for assurance purposes?
Traditional AI performs a single contained task and produces an output that a person reviews before anything happens, so assurance focuses on the quality of the output and of the human review. An agentic system plans a task, calls tools and takes actions across systems without a person approving every step, so assurance must also cover what the system is allowed to do, the access it has, and who is accountable for what it did.
Why is a deterministic IT system easier to assure than a machine learning model?
A deterministic system follows rules an engineer wrote and gives the same output for the same input every time. Once the logic, the implementation and the access are controlled, the output is controlled. A machine learning model learns patterns from data, can give different answers to the same question, and cannot always show where an answer came from.
What new risks do agentic AI systems introduce?
Acting without sign-off, reaching tools they were not meant to use, holding more access than they need, moving sensitive data between systems unseen, errors compounding when agents work together, blurred accountability across providers and deployers, and drift away from the original purpose over time.
Who should be responsible for governing AI agents in an organisation?
This is not yet settled. One view is that whoever deploys an agent should own it, because they decide what it is allowed to do. Another is that a central team of specialists should take it on, because few business owners have the expertise to oversee an agent alone.

Written by Henry Zhang, Co-founder, Kaption. Kaption is an AI partner for governance, risk and audit functions, based in London.

All posts

Ready to see governance at the speed of growth?

Book a demo