Interview
Why AI assurance must go deeper than testing
A Fellow of the Society on benchmarks, blind spots and what competent assurance actually requires.
Interview by Priya Raghunathan
Editor, ACS Insights
July 2026 · 7 min read

Dr. Alan Okonkwo FACS has spent twelve years assuring automated decision systems in health and finance. He argues that the profession's reliance on benchmark scores has produced confident engineers and fragile systems.
You've said benchmarks are the profession's comfort blanket. What do you mean?
A benchmark answers a question you already thought to ask, on data you already had. That is genuinely useful, and it is also the smallest part of assurance. The failures I have investigated almost never involved a model that scored badly. They involved a model that scored well on a distribution that stopped describing the world, in an organization where nobody owned the job of noticing.
So when a team shows me a table of evaluation results, my first question is not about the numbers. It is: what would have to be true for these numbers to mislead you, and how would you find out?
“A committee cannot be struck off a register. A person can — and knowing that changes how carefully you read your own reasoning.”
What does competent assurance look like in practice?
It looks like an argument, not a score. A competent practitioner can state the claim being made about the system, the evidence supporting that claim, the assumptions the evidence rests on, and the conditions under which the claim would no longer hold. That structure comes from safety engineering, and computing has been slow to adopt it because our failures were historically recoverable. They are not recoverable now.
The other half is monitoring. An assurance argument has a shelf life. If nobody is watching for the assumptions to expire, you have written a document, not an assurance case.
- State the claim: what exactly is being asserted about the system's behavior?
- Marshal evidence: tests, evaluations, red-teaming, operational data, formal analysis.
- Expose assumptions: what must remain true for the evidence to mean anything?
- Define withdrawal conditions: what observation would cause you to stop the system?
Where do practitioners most often go wrong?
Three places. They assure the model rather than the system, ignoring the humans, interfaces and downstream processes that determine actual outcomes. They treat red-teaming as a one-time event before launch. And they allow accountability to dissolve into a committee, which means no individual ever has to be uncomfortable.
That last one is why professional registration matters. A committee cannot be struck off a register. A person can, and knowing that changes how carefully a person reads their own reasoning.
What should a member do tomorrow morning?
Take the most consequential system you touch and write one page: the claim, the evidence, the assumptions, the withdrawal conditions. If you cannot fill in a section, you have found the work. Most engineers I know discover the gap is in assumptions — they have tested a great deal and articulated almost nothing about the world in which those tests are valid.
Join the professional body behind this work
ACS members receive our research first, free CPD and ethics modules every year, and a route to professional registration assessed by their peers.
Become a memberMore from ACS Insights
Optimus: a humanoid robot from prototype to production line
Four years from an AI Day slide to a converted Fremont assembly line — and still no commercial sale.
AnalysisGrok, Colossus and the compute arms race
xAI built a 100,000-GPU cluster in 122 days, doubled it, and merged twice. The externalities arrived with the electricity.
ArticleA national consortium to build trust in AI
ACS joins federal partners, universities and industry to strengthen assurance practice for high-impact AI systems.