Skip to main content

The AI policy was signed by the board. The AI committee did not exist.

By Harry Sidhu — ISO 27001 Lead Implementer · Director and Principal Consultant, Aegentra · Published 16 August 2026

Eight AI mistakes, one firm, nine months to fix them — and why none of the failures were unusual.

This is an illustrative composite scenario, not a named client engagement. It is constructed to demonstrate Aegentra’s methodology in ISO/IEC 42001 implementations. The organisation profile and figures below do not describe any single identifiable client, and no client is named or implied. Aegentra’s audited engagement records are published separately — see the ISO 27001 internal audit case study.

The organisation believed it was ahead of the market on AI governance. It had a board-approved AI policy, a certified model provider, and a documented human-in-the-loop control. All three were, on inspection, worth very little. None of the failures were unusual — which is exactly why this engagement is worth reading.

At a glance

MeasureDetail
OrganisationUnderwriting agency and claims administration provider, operating on behalf of two APRA-regulated insurers
Size~220 staff, Australian offices in two states, hybrid workforce
AI footprintClaims triage recommendation engine, fraud propensity scoring, document summarisation, internal policy chatbot, plus a long tail of unsanctioned tools
Existing postureISO/IEC 27001 certified; board-approved AI policy adopted eleven months earlier
EngagementISO/IEC 42001 gap assessment, AIMS implementation and certification readiness — integrated with the existing ISMS
TriggerAn insurer principal’s operational risk review questioning AI use in claims decisioning affecting members

The situation

The firm administers claims on behalf of insurers. That means its decisions land on people — a claim is approved, delayed, escalated for investigation, or declined — and it means its insurer principals carry regulatory accountability for how those decisions are made. When one principal ran an operational risk review of its material service providers, it asked a question the firm could not answer in writing: where is AI used in your claims process, and what happens when it is wrong?

The firm’s initial response was to send its AI policy. The policy was two pages, adopted by the board eleven months earlier, and stated that all AI use must be approved by the AI Governance Committee. There was no AI Governance Committee. Nobody had noticed, because nobody had ever tried to get anything approved.

Why this is common, not embarrassing. AI arrived inside organisations faster than any technology in the last twenty years, and it arrived through the business rather than through IT procurement. Most governance responses were written quickly, by people who had not yet seen how the tools were actually being used. The firms getting this wrong are not careless — they are early.

Eight mistakes we found — and see repeatedly

  1. Buying a policy instead of building a practice. A two-page template policy, board-approved, referencing a committee that had never been convened and an approval process nobody had used. A policy that describes an organisation that does not exist is worse than no policy — it becomes evidence of a control failure rather than a control
  2. Treating ISO 42001 as an IT project. Ownership had been handed to the CTO. That put the two highest-consequence systems under the governance of the team that built them, and left the executives accountable for claims outcomes with no say in how those outcomes were produced. Clause 5 leadership requirements are not ceremonial; they exist to prevent exactly this
  3. Assuming the vendor’s certification covers you. Leadership believed that because their model provider held recognised certifications, their own use was covered. Supplier assurance tells you the model was built responsibly. It says nothing about whether your deployment of it is appropriate, monitored, or lawful
  4. Confusing ISO 27001 with AI governance. The firm expected its ISMS certificate to answer AI questions. ISO/IEC 27001 governs confidentiality, integrity and availability of information. It is silent on bias, fitness for purpose, explainability, human oversight and foreseeable misuse — which is the entire subject of an insurer’s question
  5. Documenting intended use, never checking actual use. The document summariser was approved for internal file notes. It was being used to draft decision letters sent to claimants, including declines. Nobody had evaluated it for that purpose, and nobody had prohibited it, because nobody had looked
  6. Never defining what counts as AI. The policy governed “AI systems” without defining the term. Six rules engines and two long-standing statistical models sat in an argument nobody could settle, so the scope of the programme changed depending on who was in the room. Undefined scope is unauditable scope
  7. Human-in-the-loop that nobody measured. Assessors reviewed every triage recommendation — the control existed, was documented, and was described to the insurer. Concurrence with the model was 94%, with a median review time of under four minutes on complex claims. The oversight was real on the org chart and fictional in the workflow
  8. No definition of an AI incident. The incident response plan covered security events. A model producing a wrong, unfair or unexplainable output was not an “incident”, so in two years of operation exactly zero had been recorded. Leadership read that as evidence nothing had gone wrong

The finding that changed the engagement

We told the executive team that their strongest control was their weakest one. Human oversight was the answer they had given their insurer principal, their board and their complaints team — and we could show it was not happening. A 94% concurrence rate at under four minutes a claim is not review, it is confirmation. We also told them that measuring it would produce a number they would not enjoy sharing. They chose to measure it anyway, and that decision is the reason the certification audit went the way it did.

What Aegentra did, in order

The sequence matters more than the content. Organisations that write the policy first end up governing an imaginary version of themselves; we did not draft a single policy line until the inventory existed.

  1. Settled the definition question first. A determination test for what constitutes an AI system in this organisation, grounded in the ISO/IEC 22989 concepts of adaptiveness and inference rather than in marketing language. Twenty-three candidate systems were assessed against it: nine were in scope, fourteen were not — including the six rules engines, documented as out of scope with reasoning. The argument was settled once, in writing, and never had to be relitigated in front of an auditor
  2. Moved accountability out of the technology function. Top management ownership was established under Clause 5, and an AI Governance Forum was actually convened — chaired by the Chief Risk Officer, with claims operations, legal, people and culture, and technology represented. Approval authority for new AI use was placed with the people who carry the consequence of the decision, not the people who build the capability
  3. Built the inventory before the policy. Three weeks of discovery across procurement records, expense claims, SaaS spend, browser telemetry and twenty-eight interviews. The long tail surfaced here: staff drafting claimant correspondence in consumer chatbots, a team using an unapproved transcription tool on recorded calls, and a spreadsheet-based scoring model that met the firm’s own definition of AI and had never been reviewed
  4. Ran impact assessments on the systems that touch people. Claims triage and fraud propensity scoring received full AI system impact assessments under Clause 6.1.4, informed by ISO/IEC 42005. The affected parties were identified as claimants — not “customers” in the abstract — and the assessments documented foreseeable misuse, not just intended use. Low-impact tools received proportionate treatment: a register entry, a named owner and a written use boundary
  5. Redesigned human oversight so that disagreement was possible. For claims above a value and complexity threshold, the workflow was changed: the assessor records an independent initial view before the model’s recommendation is displayed. Concurrence, review time and override reasons became monitored metrics with thresholds. Concurrence fell from 94% to 71% within two months — meaning roughly a quarter of those claims had previously been decided by a model and countersigned by a person
  6. Established performance baselines before the next model update. Accuracy and outcome distribution were measured by claim type and by claimant cohort, so that drift after a provider-side model change could be detected rather than inferred. A rollback trigger and an owner were defined for each threshold. Without a baseline, “the model got worse” is an opinion
  7. Defined what an AI incident is, and wired it into the existing process. Three new categories were added to the existing incident taxonomy: erroneous output with claimant impact, unexpected behaviour following a model or prompt change, and use outside an approved boundary. Eleven incidents were logged in the first month — none of them new, all of them previously invisible
  8. Drew a boundary around the summariser, in writing. Use for internal file notes: permitted. Use for correspondence communicating an adverse decision to a claimant: prohibited, on the grounds that the system had never been evaluated for that purpose and the firm could not explain an output it had not tested. The prohibition was enforced in the tool, not just in the policy
  9. Fixed the supply chain with contract language. Provider assurance obtained and assessed rather than assumed; thirty days’ notice of material model change; committed processing region; prohibition on training against claims data; and a documented split of which controls the firm relies on the provider for and which it must operate itself
  10. Integrated the AIMS into what already existed. One risk register with an AI risk category, one management review, one internal audit programme, one corrective action process, one supplier register. Two parallel management systems produce two different answers to the same auditor question, and roughly double the ongoing overhead for no additional assurance

The outcome

MeasureResult
Gap assessment to ISO/IEC 42001 certification, no major nonconformities9 months
Candidate systems assessed; nine in scope, fourteen excluded with documented reasoning23 → 9
Model concurrence rate once assessors recorded an independent view first94% → 71%
AI incidents recorded in month one; visibility gained, not performance lost0 → 11

The insurer principal’s operational risk review closed without a finding against the firm. The AIMS now answers AI due diligence questionnaires from a single evidence set, and the firm has since used its certificate in two competitive tenders where AI governance was a scored criterion. Internally, the most valuable output was not the certificate: it was a claims process in which a person can now disagree with the model and have that disagreement recorded.

The lesson you can apply today

If your AI cannot be overruled in practice, you do not have human oversight — you have a signature. Test it yourself before an auditor does: take any AI-assisted decision in your organisation and ask three questions. How often does the human disagree with the system? How long do they spend before agreeing? And what happens, procedurally, when they do disagree? If you cannot answer the first two with a number, the third question has never been asked.

Five questions, before you spend anything

  • Can you produce a list of every AI system in use, dated within the last month? If not, every other AI control you have is being applied to an unknown population. The inventory is the foundation, and it is the artefact every due diligence questionnaire is really asking for.
  • Does your organisation have a written definition of what counts as an AI system? Without one, your scope will move every time a new person joins the conversation, and an auditor will set the boundary for you.
  • Who approves a new AI use — by name, not by committee title? If the answer is a body that has never met, or a technology executive who also builds the systems, accountability has not been established.
  • What is your override rate, and do you measure it? This single metric costs almost nothing to collect and tells you more about your AI risk than any policy document you will ever write.
  • What is an AI incident in your organisation, and how many have you logged? If the answer is zero, that is a reporting failure, not a safety record.

Frequently asked questions

Does ISO 27001 certification cover AI governance?

No. ISO/IEC 27001 governs the confidentiality, integrity and availability of information. It is silent on bias, fitness for purpose, explainability, human oversight and foreseeable misuse — which is what a regulator, an insurer or a customer is asking about when they ask about AI. ISO/IEC 42001 is the management system standard that addresses those, and it integrates with an existing ISMS rather than replacing it.

Does our AI vendor’s certification cover our use of the tool?

No. Supplier assurance tells you the model was built and operated responsibly by the provider. It says nothing about whether your deployment is appropriate for its purpose, monitored in production, or lawful in your jurisdiction. Under ISO/IEC 42001 the deploying organisation carries its own obligations, and a certification body will ask which controls you rely on the provider for and which you operate yourself.

What counts as an AI system for ISO 42001 scope?

The standard does not hand you a list, which is why the determination has to be written down before scoping begins. ISO/IEC 22989 supplies the concepts — adaptiveness and inference — and a written determination test applied consistently is what makes scope defensible. Rules engines and long-standing statistical models frequently fall outside, but they must be excluded with recorded reasoning rather than by assumption.

How long does ISO 42001 certification take in Australia?

For an organisation with an existing certified management system to build on, a realistic range is six to twelve months from gap assessment to certification audit, depending on the number of in-scope systems and how much of the inventory already exists. Organisations starting without an ISMS, or with a large unsanctioned AI footprint, should expect longer — the discovery work, not the documentation, is usually what sets the timeline.

What is an AI system impact assessment, and is it mandatory?

ISO/IEC 42001 Clause 6.1.4 requires an assessment of the potential consequences of an AI system for individuals and groups. ISO/IEC 42005 provides guidance on conducting one. It is not optional for in-scope systems, and the part organisations most often miss is that it must consider foreseeable misuse, not only intended use — which is precisely where a tool approved for one purpose and used for another gets caught.

How do we prove human oversight is real rather than nominal?

With numbers. Record how often the human disagrees with the system, how long they spend before agreeing, and what happens procedurally when they do disagree. If a reviewer sees the model output before forming a view, concurrence rates tend to be very high and the control is closer to confirmation than review. Recording an independent view first, then measuring concurrence and override reasons, is what turns a documented control into a demonstrable one.

Start with the inventory, not the policy

Aegentra runs a fixed-scope AI discovery and gap assessment against ISO/IEC 42001 that tells you what you actually have, what it touches, and what a certification body would ask about it. See ISO 42001 implementation, the ISO 42001 certification guide for Australia, or the ISO 42001 Lead Implementer course if you would rather build the capability in-house. Book an ISO 42001 discovery call.

About this case study

The engagement described here is an illustrative composite scenario, constructed to demonstrate Aegentra’s methodology and decision-making in ISO/IEC 42001 implementations. Organisation profile, figures and details do not describe any single identifiable client, and no client is named or implied. Named references are available on request under a mutual non-disclosure agreement. Nothing in this document constitutes legal or regulatory advice; certification outcomes depend on the accredited certification body’s independent assessment, and obligations under APRA prudential standards and the Privacy Act should be confirmed with qualified advisers.