

- Case study
- Internal audit
- Manufacturing
How Celsa Group piloted AI-powered control testing with Kaption
Audit functions worldwide face a widening gap between increasing control volumes and limited team capacity. The result is compromised risk coverage, heavy reliance on outsourcing, and delayed insight for improvement.
To address these challenges, Celsa Group, a leading European manufacturing group, partnered with Kaption to create MAAT (Multi-Agent Audit Technology). A five-month pilot across 60+ manual and IT general controls under International Financial Reporting Standards proved that a human-on-the-loop approach lets AI expand coverage while keeping expert auditors firmly in control.
- Client
- Celsa Group
- Headquarters
- Barcelona, Spain
- Sector
- Steel and manufacturing
- Employees
- c.5,000
- Legal entities
- 30+
- Pilot scope
- 60 ICFR / SCIIF controls

01 The problem
An industry-wide constraint
Many enterprises face the same assurance gap: the volume of controls their audit function is expected to test has grown well beyond what the in-house team can realistically cover.
The result is heavy reliance on outsourcing, slow and costly insight into where controls are failing, and risk coverage that falls short of what regulators and boards expect. Fewer than half of Chief Audit Executives feel confident in their team's ability to assure rapidly expanding risk and compliance demands.
For most multi-entity enterprises, full-population control testing is simply not achievable with existing resources: too many controls, too many of them informal or undocumented, too few auditors, and testing cycles that stretch over months rather than days.
- Controls to test
- In-house capacity
- Untested risk
02 The challenge
Celsa, a textbook case
Celsa Group is a textbook example of this industry-wide challenge: a control universe of more than 800 distinct controls, generating over 4,500 testing instances a year, managed by an internal audit team of fewer than five FTEs.
800+
distinct controls
4,500+
testing instances a year
<5
internal audit FTEs
Operating across a complex, multi-entity group, the gap between what the team could realistically cover and what the organisation required had become significant, and was widening.
As an initial response, Celsa relied on an outsourced testing provider. A single engagement covering approximately 170 controls took three months to complete. At that pace, full-population coverage was operationally and financially out of reach.
The team had already solved coordination with more than 150 control owners by adopting a GRC platform to centralise risk and control data. Testing itself remained entirely manual, which limited the value the platform could deliver and left the core capacity problem unsolved.
As the business evolved and stakeholder demands intensified, Celsa faced a compounding constraint: more coverage, more frequently, with greater rigour, and no scalable way to deliver it.
03 The solution
Introducing MAAT
Kaption worked closely with Celsa's internal audit team to develop MAAT, an AI system that automates control testing at scale and mirrors the way Celsa's auditors really work.
Rather than a generic deployment, MAAT was configured against Celsa's own control standards, audit methodology and evidence evaluation criteria, and also consolidates the expectations of external auditors. It ingests real audit evidence, analyses it against each control's requirements, and delivers a clear assessment of control effectiveness with full reasoning and source-level citations, documenting every step.
Every verdict is logged and traceable, giving audit leadership complete visibility into how each conclusion was reached. Because audit is judgement-based and relies on experience, MAAT was designed and tuned with Celsa's experienced auditors to produce justifiable, defensible and fully traceable reviews.
01 Ingest
Diverse evidence
Structured and unstructured files, such as PDFs, spreadsheets and screenshots, are ingested by the model.
02 Analyse
Multi-agent reasoning
Specialist agents evaluate evidence against Celsa's control standards and methodology.
03 Cite
Source-level citation
Every conclusion links back to the specific evidence and rule that produced it.
04 Verdict
Justified outcome
Effective, Ineffective or Needs Review. Fully logged, traceable and defensible.
04 The approach
Human on the loop
MAAT is built on a human-on-the-loop principle. Unlike human-in-the-loop systems, where the AI stops at every turn to wait for manual intervention, human on the loop shifts the auditor's role from execution to oversight.
That distinction matters. The goal is not an autonomous system that blindly signs off on controls, nor a tool that needs hand-holding at every step. The AI tests independently, at a scale previously impossible without large teams. The expert stays on the loop, supervising the entire population, reviewing the evidence the AI has compiled, and confirming final conclusions rather than performing every step by hand.
A set of safeguards strengthens the results in line with external stakeholders' expectations. MAAT is designed to know the limits of its own confidence: it gives a firm opinion when the evidence supports one, and defers to the auditor whenever it does not.
When confident
It states an opinion
Where the evidence clearly supports a conclusion, MAAT issues a definitive verdict, fully cited and traceable to source.
When uncertain
It suggests, not asserts
Where confidence is lower, MAAT offers a recommendation and flags the item for human review rather than guessing. This keeps hallucination risk low.
Safeguard 01
Evidence-bound reasoning
Every conclusion is grounded in the evidence actually reviewed and linked back to the specific source and rule that produced it.
Safeguard 02
Bias toward caution
Lower-confidence cases are routed to Needs Review rather than passed through, so the system errs on the side of caution, never leniency.
Safeguard 03
Full traceability
Every verdict is logged and auditable, giving leadership complete visibility into how each conclusion was reached, and the ability to challenge it.
Safeguard 04
Expert sign-off
The auditor remains the final authority. MAAT prepares and justifies; the human reviews, confirms and owns the conclusion.
“Far beyond anything similar I've ever seen in the market.”
05 The results
Pilot outcome
The results validated the approach. All control verdicts met expectations, and the system behaved exactly as a high-stakes audit environment requires.
In 85% of cases, MAAT's reasoning not only reached the right verdict, it matched the quality and structure of what an experienced auditor would produce. All outputs were grounded in real evidence, clearly reasoned and fully traceable to source.
MAAT produced fewer than 10% false negatives across the entire pilot. Where confidence was lower, the system flagged the assessment as Needs Review rather than passing it through.
On average, MAAT completed the assessment of a control in under two minutes, with a fully cited explanation ready for auditor review.
Reasoning alignment
85%
MAAT against the auditor baseline85 / 100
False negatives
<10%
- Verdict confirmed
- Flagged for review
Time per control, on average
<2minutes
With fully cited reasoning
Pilot coverage
50controls
- Procure to pay
- Order to cash
- Payroll
- Inventory
- Treasury
- Valuation
“Auditors rarely agree word for word. 85% alignment on reasoning is more consistent than most human hires.”
06 The impact
Benefits
When AI handles first-line control testing at speed and to standard, the coverage gap closes. A fully tested control matrix across the full population, not just a sample, becomes operationally achievable for the first time.
- 01
Higher-frequency testing
Controls can be tested more often, including those currently sampled only once a year.
- 02
Broader entity coverage
Entities previously out of scope because of capacity limits can be brought into the testing population.
- 03
Auditor capacity unlocked
Time previously spent on first-line testing is redirected to higher-value, judgement-led work.
07 Next step
What's next
The pilot results were strong. The next phase takes MAAT from pilot to group-wide rollout.
MAAT's natural evolution expands coverage in two ways: moving from testing to advising, by suggesting to control owners how their controls could improve; and extending to other domains such as tax, cybersecurity, operational and safety controls across the Group.
From c.60 controls piloted to a full-group rollout across finance, tax and operational controls.
1,000+control tests a year