The gap between using AI and governing it
Most organisations crossed the line from experimenting with AI to depending on it somewhere in the last two years, and very few noticed the moment it happened. Models now sit inside hiring shortlists, credit decisions, pricing, customer service and internal reporting, and a good share of them arrived as a feature in software bought for something else entirely. What has not kept pace is the ability to say who authorised any of it.
ISO/IEC 42001:2023 is the first international standard for an artificial intelligence management system, and it exists to close exactly that gap. It is routinely misread as a documentation exercise, which is the fastest available way to spend a year on it and finish with no more control than you started with. The standard is not asking you to describe how you would govern AI. It is asking you to show the decisions you actually made.
This guide is written for the executive who has to fund the programme and for the compliance, security or data leader who has to run it. It covers what the standard contains, the operating model that makes it function, the three patterns that predict whether a programme succeeds, and where all of this now sits against the EU AI Act after the timetable moved in July 2026.
The four dates that frame this programme
ISO/IEC 42001:2023 published in December 2023. ISO/IEC 42006:2025, which governs the bodies allowed to certify you, published 7 July 2025. ISO 19011:2026, the auditing guidelines, published 27 May 2026 with no transition period. Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force 27 July 2026 and moved full high-risk obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I embedded systems.
What the standard actually contains
ISO/IEC 42001 follows the same Harmonised Structure as ISO/IEC 27001 and ISO 9001, running from clause 4 on context through leadership, planning, support, operation, performance evaluation and improvement in clause 10. If you already operate a certified management system, that structure is familiar and it transfers almost intact, which is the single biggest accelerator available to you.
Attached to those clauses is Annex A, which carries 38 controls grouped into nine control objectives numbered A.2 to A.10, covering AI policy, internal organisation, resources, impact assessment, the AI life cycle, data, information for interested parties, use of AI systems, and third-party relationships. Unlike the clauses, the Annex A controls are applied selectively: you decide which ones apply inside your scope and you justify every inclusion and exclusion in a Statement of Applicability.
The substantive difference from ISO/IEC 27001 sits in one place. An information security risk assessment reasons about confidentiality, integrity and availability, so its subject is the organisation and its information. An AI impact assessment reasons about effects on individuals and groups, including foreseeable misuse, so its subject is the people on the receiving end of an automated decision. Those are different analytical objects, they need different participants around the table, and this is the requirement that mature security teams most often assume they have already satisfied when they have not.
It is worth being equally clear about what the standard is not. It does not assess whether a given model is accurate or safe, because that is a product-level question that belongs in technical documentation. It is not a conformity assessment under the EU AI Act, and as of August 2026 it confers no presumption of conformity, since no harmonised standard has yet been cited in the Official Journal. A certified organisation can still deploy a poor model, and the audit will only catch that if the governance process which should have caught it also failed.
ISO 27001 to ISO 42001: what transfers and what does not
Dimension: Clause structure
Dimension: Management review and internal audit
Dimension: Risk assessment object
Dimension: Annex A controls
Dimension: Asset inventory
Dimension: Certification body accreditation
Dimension: What the certificate proves
Why implementations stall
Programmes that lose momentum tend to fail in one of three recognisable ways, and none of them looks like failure while it is happening. Each produces visible artefacts, which is precisely what makes them dangerous, because a steering committee sees output and concludes that progress is being made.
The first is documentation-first. A policy set arrives quickly because policies are easy to write and easy to show, and then the programme discovers at the certification audit that policies without operating records prove nothing. An AI impact assessment procedure that exists, has been signed off and has never been applied to a live model is not evidence of governance, it is evidence of intent.
The second is technical-only ownership. The programme lands with the data science or platform team because they understand the models, and it never acquires the legal, privacy and business participation that impact assessment requires. The output is competent, well structured and consistently silent on the affected-person dimension the standard is actually asking about.
The third is the parallel structure. A new AI governance committee, a new risk register and a new control library are stood up alongside the ones that already exist, and within two quarters they have drifted apart. Now every change has to be made twice, nobody is sure which register is authoritative, and the duplication itself becomes the reason people stop maintaining either.
The tell that a programme is drifting
Ask for the date of the last AI impact assessment that changed a decision. Not the last one completed, the last one that caused a system to be modified, delayed or rejected. If nobody can name one, the assessments are being produced as artefacts rather than used as controls, and that gap will surface at Stage 2 whatever the documentation looks like.
The operating model that works
A working implementation rests on five components that connect to each other. Treated separately they become five workstreams that never converge, so it is worth being explicit that the inventory feeds the classification, the classification determines the controls, the controls need owners, and the monitoring closes the loop back to the inventory.
An inventory that stays true
Everything else depends on the inventory, and it is the document most often produced once and never touched again. A register that was accurate on the day the scope was written is worse than no register, because it creates confidence that is not warranted. Each entry needs enough to make the system governable rather than merely listed:
- The purpose and the business process it serves, written so a non-technical reader understands what decision it affects
- Lifecycle stage, from proposed through pilot and production to retired, with the date of the last change
- A named business owner accountable for the outcome and a named technical owner accountable for performance
- The risk classification and, more importantly, the basis on which it was assigned
- Data sources, the rights you hold to use them, and for third-party models whatever the supplier discloses about the base model
- Whether it was built, bought, or arrived embedded in a product you already owned
- The model version currently in production, so evidence can be tied to something specific
The harder problem is discovery rather than format. Start from the list your teams volunteer and you will get the systems somebody built on purpose. Do the arithmetic instead: walk the admin console of every major SaaS platform in the estate and count the AI features already switched on, then add the coding assistants in engineering, the summarisation in the service desk, the scoring in the recruitment tool and the drafting inside the productivity suite. Almost none of those were built by anyone in the room, and every one of them meets the standard's definition of an AI system.
Shadow AI, defined usefully
Not employees pasting data into a chatbot, which is a security problem you already have a control for. Shadow AI in the governance sense is AI capability that a vendor enabled inside a product you already licensed, without a purchase decision, without a risk assessment and often without a release note anyone read. It is in scope, it is growing, and no procurement gate currently catches it.
Risk classification that maps to something real
Resist the urge to invent a bespoke five-point scale. The classification is more useful when it maps onto decisions somebody else will also make about the same system, which in practice means anchoring it to the EU AI Act tiers, since a system that is high-risk under Annex III will attract regulatory obligations on a known timetable regardless of how you score it internally.
On top of that regulatory anchor, four internal criteria do most of the work: the severity of effect on an individual if the system is wrong, whether a human meaningfully reviews the output before it takes effect, the sensitivity of the data involved, and how reversible the resulting decision is. A pricing model that can be corrected next week is a different proposition from a screening model that quietly removes a candidate from a shortlist.
What the classification buys you is proportionality, and that is the point. It tells you which systems need full impact assessment, independent validation and active monitoring, and which need registration and an annual look. Without it every AI system attracts the same weight of process, the process becomes unaffordable, and teams route around it.
Ownership that survives contact with reality
Three roles need naming per system, and the failure mode is collapsing them into one. The business owner is accountable for the outcome and has the authority to stop the system, which means it cannot be a job title from the technology function. The technical owner is accountable for performance, drift and the integrity of the model in production. The risk or compliance owner is accountable for oversight, and specifically for the fact that the impact assessment happened and was acted on.
Above the individual systems, most organisations create an AI committee, and most AI committees become theatre within a year. The distinction between the ones that work and the ones that do not is whether the committee holds a decision right. A body that reviews and advises will be skipped the moment a deadline gets tight. A body that must approve deployment of a high-risk system before it goes live will not be, because the deployment cannot happen without it.
Controls placed across the life cycle
Annex A spreads its controls across the life cycle rather than concentrating them at deployment, which is deliberate. Data governance and provenance decisions are made long before a model exists. Impact assessment belongs at design, when changing the answer is still cheap. Validation and testing, including whatever fairness objective you have set for the system, belong before release. Transparency artefacts have to be produced for each audience you identified, and tied to a specific model version. Human oversight and incident handling only start to matter once the system is live.
The implementation instinct that pays off is to embed rather than to layer. If your organisation already has a change advisory board, a procurement gate and a release process, the AI controls belong inside those, expressed as additional criteria. Build them as a separate track and you have created the parallel structure described earlier, with all the drift that follows.
Monitoring that notices the right things
AI incidents do not look like security incidents, and this is where most incident processes quietly fail. Model drift, degraded performance for one subgroup, an unexpected output pattern and a hallucination reaching a customer are all incidents in governance terms, yet an incident process built for outages and breaches has no category for any of them and no threshold that would trigger.
Fixing that is mostly definitional. Extend the incident taxonomy to cover behaviour and performance alongside availability, set monitoring thresholds per system rather than one global rule, and make sure at least one AI incident has travelled the full route from detection through triage to documented resolution before an audit asks. A process that has never been used is indistinguishable from a process that does not work.

Where governance actually happens
Governance is not exercised in a policy, it is exercised at the handful of moments where something can be stopped. If you want to know whether a management system is real, find those moments and check whether the gate is actually load-bearing.
- Intake, where a proposed use case is registered and given a provisional classification before work starts
- Procurement, where a vendor is asked whether the product contains AI and what it does with your data
- Approval, where a high-risk system cannot reach production without a named person accepting the residual risk
- Material change, where retraining or a purpose change triggers reassessment rather than passing silently
- Incident escalation and periodic review, where what you learn feeds back into the classification
Five gates is enough. If those exist and hold, the management system is doing its job, and the documentation becomes a record of it rather than a substitute for it. If they do not exist, everything else you build is a library.
Three implementation patterns
Programmes tend to fall into one of three shapes depending on where the organisation is starting from, and the shape predicts the difficulty more reliably than size or sector does. The first of these is public and verifiable. The other two are reasoned from what happens when the standard meets a particular kind of estate.
Governance first, certification second
Anthropic achieved accredited ISO/IEC 42001 certification for its AI management system, among the first frontier AI labs to do so, with the certificate issued by Schellman Compliance, LLC under ANSI National Accreditation Board accreditation and effective from January 2025. Read the scope line rather than the headline: it covers AI research and development and AI services, not the organisation entire, which is the distinction every buyer should be making when a supplier waves a certificate.
The instructive part is the order of operations. The certification rested on governance that already existed and was already public, including a published responsible scaling framework, deliberate work on model behaviour and alignment, and ongoing safety research. The audit did not create the governance, it made existing governance legible to a third party. That is the right way round, and it is the order most enterprise programmes reverse when they commission a policy set in order to become certifiable.
The certified incumbent
An organisation with a mature certified ISMS approaches ISO/IEC 42001 as a delta exercise, and on the structural side it is right to. The clause architecture, the management review rhythm, the internal audit machinery and the corrective action process all carry over, which typically takes months off the timeline. The trap is assuming the risk work carries over with them, because the impact assessment is genuinely new analysis requiring participants the security function does not usually convene. Programmes in this shape succeed or fail on whether legal, privacy and a business owner are in the room for the first assessment.
The AI-native company with no management system
A company whose product is AI usually has excellent model documentation, real evaluation discipline and no management system at all, so the gap runs the other way. Nobody needs convincing that testing matters. What is missing is the clause 5 leadership commitment, the clause 9 performance evaluation rhythm and the habit of recording that a named person accepted a residual risk on a given date. For this shape the work is administrative rather than technical, which teams find unglamorous, and it is usually the fastest of the three to complete once someone accepts that the paperwork is the deliverable.
Ask me what predicts whether a programme will pass, and it comes down to a single question: can you name the person who approved this model going into production, and can you show what they were told at the time? Everything the standard asks for is downstream of that. An organisation that answers it easily will pass whatever its documentation looks like, and one that cannot will struggle no matter how complete the policy set is, because an audit tests the decision trail rather than the prose.
Integration with what you already run
The integration decision is the one that determines the running cost of the programme for the next five years, and it is usually made implicitly in the first month by whoever opens the first spreadsheet. Made deliberately, it comes down to extending four things you already have rather than building beside them.
Your ISMS risk assessment gains AI-specific threats such as model manipulation, data poisoning and prompt injection, which sit naturally alongside the threats already there. Enterprise risk management gains AI as a named category with its own appetite statement, rather than leaving it buried inside technology risk. The compliance function picks up the AI Act obligations on the same register that already carries GDPR and sector regulation. Data governance extends to cover model and dataset oversight, since the lineage questions an AI audit asks are the ones a data office is already best placed to answer.
Tooling matters less than the discipline, but it does decide whether the discipline survives year two. The requirement is a single control library where one control can carry several regulatory readings at once, so that changing it once updates every framework it serves. Some organisations achieve this inside an existing GRC platform, others use a multi-framework tool such as Acuna, which models ISO 27001, ISO 42001 and EU AI Act controls as one auditable set rather than three parallel documents. A spreadsheet can hold the mapping but cannot show you what a change breaks elsewhere, and that is the property you are actually buying.
Disclosure
Acuna is the tool I work with directly, and I am its CEO. Acuna and Abilene Academy are both part of Abilene Group. The requirement described above, a single control library carrying several regulatory readings at once, is the actual point of this section and can be met with any platform that supports it, so weigh the named example accordingly.
Practical tip
Do not stand up a separate AI governance structure. Extend the ISO 27001 and enterprise risk frameworks you already operate, and give the AI committee one decision right it genuinely owns. A governance body with an approval gate outlives a governance body with an agenda.
The EU AI Act connection after the Omnibus
Any plan written before August 2026 needs re-baselining, because Regulation (EU) 2026/1744, the Digital Omnibus on AI, entered into force on 27 July 2026 and deferred full high-risk obligations to 2 December 2027 for Annex III systems and 2 August 2028 for Annex I embedded systems. The Commission's stated reasons included the absence of harmonised standards and the incomplete designation of conformity assessment bodies, which is an unusually direct admission that the assessment machinery was not ready.
Two things did not move, and both catch people out. Transparency obligations under Article 50 applied from 2 August 2026, with machine-readable marking for generative systems already on the market required by 2 December 2026. Prohibited practices have applied since February 2025 and are unaffected. A deferral of the high-risk regime is not a general amnesty, and it does nothing at all about what your customers put in their procurement questionnaires.
Where ISO 42001 genuinely helps is the overlap with Article 17, which requires providers of high-risk systems to operate a quality management system. Much of what you build for clause 4 through clause 10 answers that requirement directly, and the risk management system in Article 9 maps onto the impact assessment discipline. That overlap is real and worth exploiting, but it is preparation rather than substitution, because Article 43 conformity assessment remains a separate procedure and a management system certificate is not one.
No presumption of conformity, yet
As of August 2026 no AI Act harmonised standard has been cited in the Official Journal, so nothing currently grants the Article 40 presumption of conformity. EN 18286:2026 on quality management systems was approved in July 2026 and is the first AI Act standard published, but publication is not citation and only citation carries the legal effect. Treat any vendor claim that ISO 42001 certification makes you AI Act compliant as a reason to check their other claims.
The first ninety days
Spend the first month on visibility rather than policy. Run the inventory sweep properly, including the admin-console walk through your SaaS estate, and resist the temptation to start drafting until you know what you are governing. The inventory sets the scope, the scope sets the Statement of Applicability, and the Statement of Applicability sets everything you will eventually have to evidence, so getting it wrong first is expensive to unwind.
Use the second month to establish the decision gates and the ownership map, because these are what make the programme real to the rest of the organisation. Pick the two or three highest-classification systems and run a full impact assessment on each with legal and a business owner present, partly to produce the assessment and partly to calibrate how long one actually takes before you commit to a schedule.
In the third month, write the Statement of Applicability against what you now know, and set the internal audit programme up against ISO 19011:2026 rather than the 2018 edition it replaced. Then test yourself the way an audit will: pick a model at random and try to trace it from impact assessment through data decisions and test results to approval and monitoring. Whatever breaks in that trace is your actual backlog.
- The AI inventory covers built, bought and embedded systems, and includes an admin-console sweep of every major SaaS platform
- Every entry has a named business owner and a named technical owner, not a team name
- Risk classification criteria are written down and anchored to the EU AI Act tiers as well as internal impact
- The Statement of Applicability justifies every included and excluded Annex A control against the risk assessment
- Impact assessments for the highest-classification systems are complete, with legal and a business owner among the named participants
- Five decision gates exist and are load-bearing: intake, procurement, approval, material change, and periodic review
- The AI committee holds at least one decision right that cannot be bypassed
- AI controls are embedded in the existing change, procurement and release processes rather than run in parallel
- The incident taxonomy covers performance and behaviour, and one AI incident has been through it end to end
- The internal audit programme has been updated to ISO 19011:2026
- A random model can be traced from impact assessment to approval to monitoring in under an hour
- Any chosen certification body is accredited against ISO/IEC 42006:2025
Pitfalls, described as symptoms
Most failure modes are easier to spot as symptoms than as principles. If the risk model has more than about five factors, it has stopped being a decision aid and become an exercise, and the people meant to use it will start guessing the output and working backwards. If the inventory has not changed in a quarter while the business has shipped new products, it is not being maintained. If no AI initiative has ever been delayed or rejected by the governance process, the process is decorative.
Two more are worth naming. Third-party AI is the most commonly excluded category and usually the largest part of the estate, so a scope that quietly omits it is a scope that will not survive a customer's due diligence. And treating the programme as a project with an end date guarantees decay, because the certificate arrives at exactly the moment attention moves elsewhere, which is also when the first surveillance audit starts accumulating the findings nobody is watching for.
What good looks like after twelve months
A year in, the observable difference is speed rather than paperwork. New AI use cases move faster because the intake path is known and the classification tells everyone immediately how much process applies, so teams stop negotiating the rules case by case. Customer security questionnaires get answered from existing records instead of triggering a scramble. And when something goes wrong, which it will, the organisation can show what it knew, when it knew it and who decided to proceed.
That last capability is the one that turns governance from a cost into a commercial asset, because it is precisely what an enterprise buyer, a regulator and an insurer each want to see, and none of them can be satisfied retrospectively.
Next step
If you are building the management system, the ISO 42001 Lead Implementer programme covers the clause-by-clause work and the Statement of Applicability in the detail this guide compresses. If your next question is what an audit will actually ask for, the companion piece on the ISO 42001 evidence playbook goes through the file control area by control area, and the ISO 42001 Lead Auditor certification is the route for anyone who will lead or commission those audits. Both sit in our AI governance training hub.




