ISO/IEC 42001 asks a company to run its AI under a management system that an auditor can check, and for most companies today the AI that needs managing is not a product they sell but the agents their developers run. Claude Code, Cursor and Codex sit on developer laptops with the developer's access, read text all day that they may treat as instructions, and load MCP servers, skills and extensions that nobody reviewed. This article is for the CISO or compliance lead who has been asked about ISO 42001, and it goes control by control through the Annex A areas that touch those agents: what each one asks, what goes wrong on a real machine, what evidence satisfies it, and which controls still need a person. It is the third in a series on frameworks, after the EU AI Act and the OWASP LLM Top 10.
What is ISO/IEC 42001, and who gets asked for it?
ISO/IEC 42001:2023 is the first international standard for an AI management system, published in December 2023 by the joint ISO and IEC committee for artificial intelligence (source: Schellman, a certification body, on preparing for ISO 42001). It is a management system standard in the same family as ISO 27001: it sets out how an organisation governs its AI, and an accredited certification body can audit the organisation and certify it. BSI, one of those bodies, describes the outcome as independent assessment and certification for your AI management system (source: BSI). The standard applies to organisations that provide, develop, deploy or use AI systems, and its Annex A lists 38 controls grouped into control areas from A.2 to A.10.
Who gets asked for it is the practical question. It is voluntary, so the ask arrives from outside: a customer's security questionnaire, a procurement process, a tender in a regulated sector, or an internal programme that wants one framework to hang the EU AI Act work on. Once a company answers "yes, we are working towards it", the auditor's first question is which AI systems are in scope.
Why the agents on developer machines are in scope
Because "use AI systems" is in the definition, and a company whose developers run coding agents is using AI systems under its own management, whether or not anyone wrote that down. A management system that lists the chatbot in the product and forgets the agents on three hundred laptops has drawn its scope boundary around the smaller risk.
The agents are the larger risk because of where they run. Each one has the developer's access: every cloned repository, every .env file, the cloud CLI profiles, the SSH keys, the package registry tokens. One real laptop had forty-six credentials its agents could read. Each one loads components the developer installed without a ticket, and each one reads text all day that it may treat as instructions. An AI management system that does not cover that estate has a gap that shows up in the first interview, when the auditor asks how the company knows which AI systems it operates.
The Annex A controls, on a developer laptop
The ZYBE compliance view assesses nine Annex A control areas against what the endpoints report. Seven are assessed from evidence. Two are assessed by attestation, because no endpoint can prove them, and the report says so by name. Here they are in order.
A.3 Internal organisation
The control asks that roles and responsibilities for AI are defined and assigned. On an estate of agents, the failure is an endpoint nobody owns, or an endpoint that still belongs to somebody who has left, or one person with a dozen machines and no idea what runs on them. Evidence is the ownership record per endpoint, and the findings Endpoint has no owner, Endpoint belongs to somebody who has left and Owner with unusually many endpoints.
A.4 Resources for AI systems
The control asks that the components, tools and data an AI system uses are identified and documented. For agents, the resources are the agents themselves and everything they load: MCP servers, skills, extensions, hooks and local models. The failure is an agent nobody shipped, a component that arrived without review, an agent configured on a machine but not installed, or a local model server open to the network. Evidence is the inventory per endpoint and the findings Agent nobody shipped, Unreviewed components, Configured but not installed and Local model server reachable from the network. The setting that matters is holding new MCP servers and skills for review, so the list stops growing while you document it.
A.5 Assessing impacts
The control asks that the impacts of AI systems on individuals and society are assessed. This is judgement, not telemetry. No endpoint can produce an impact assessment, so the compliance view marks A.5 as manual and records who attested to it and when.
A.6 AI system life cycle
The control asks that systems are developed, deployed and operated under defined, verified controls. On a laptop the life cycle failure is drift: a component that passed review and then changed, a hook that appeared outside the managed layer, an agent that wrote to a startup location, a component that tells the agent to bypass approvals, or a safety control the agent ships with that somebody switched off. Evidence is the change record and the findings Approved component changed, Unapproved hook, Agent wrote to a startup or hook location, Component tells the agent to bypass approvals and Agent safety control turned off.
A.6.2.6 Operation and monitoring
The control asks that AI systems are monitored in operation and that events are logged. The failure mode is silence: sessions running with nothing observing them, an endpoint that stopped reporting, a machine where the hook reports but no inventory scan has run. Then the behavioural signals that something changed: a destination the endpoint has never used, more tool calls than the agent normally makes, activity outside the agent's usual hours, a long run with nobody at the keyboard. Evidence is the session and tool call record, the setting Token usage is recorded, and the findings Sessions running unobserved, Endpoint has stopped reporting, Hook reporting but no inventory scan, A destination this endpoint has never used, More tool calls than this agent normally makes, Agent active outside its usual hours and Long run with nobody at the keyboard.
A.7 Data for AI systems
The control asks that the data the system uses is controlled, including what it must not see. For agents, the data it must not see is credentials. The failures are a credential in an agent's own configuration, credentials readable in a project the agent works in, an agent that read a credential file, a secret pasted into the chat, a secret that came back in tool output, and the same credential on several machines. Evidence is the credential findings, reported as names and hashes, never values: Credential in an MCP configuration, Credentials an agent can read, Agent read a credential file, Secret pasted into the agent chat, Secret in tool output and Same credential on several endpoints.
A.8 Information for interested parties
The control asks that users and affected parties are informed about the AI system. That is communication, and the endpoint cannot see whether it happened. Manual, attested by name.
A.9 Use of AI systems
The control asks that systems are used as intended, by people who are allowed to, with human oversight. This is the standing permissions control, and the one a CISO should read twice. The failures are an agent that acts without asking, a shell permission that removes the person from the loop, a standing permission granted in a session months ago and never revisited, an agent running unattended, an agent signed in as somebody who has left, or signed in with an account nobody recognises. Evidence is the permission set per agent per machine and the findings Agent acts without asking, Agent may run shell commands without asking, Standing permission granted, Agent runs unattended, Agent signed in as somebody who has left and Agent signed in with an unknown account. The rules a company sets over what an agent may do without asking are enforced on the same endpoint, and new subagents can be held for review.
A.10 Third-party and customer relationships
The control asks that suppliers of AI components are managed and their products checked. Every MCP server, skill and extension is a supplier. The failures are a package or an agent with a known advisory, a component judged malicious or suspicious, an MCP server not pinned to a version, a package published in the last days, and an agent far behind its current release. Evidence is the component inventory matched against advisories and the findings Package with a known advisory, Agent with a known advisory, Package flagged as malicious, Component judged malicious, Component judged suspicious, MCP server package not pinned to a version, Package published in the last days and Agent far behind its current release. The MCP server piece goes into what a supplier can do once loaded.
Seven of these nine controls are answered by what is already on your developers' machines today. Join the launch list and get the framework guides first, starting with the evidence maps for ISO 42001 and the EU AI Act.
A worked example: one laptop, three controls
A senior developer's laptop runs Claude Code and Cursor. Cursor has a standing permission to run any shell command. A project on the machine has a .env with a production database URL in the agent's working directory. Last week the developer added an MCP server from a repository README, and it has not been reviewed.
Three findings: Agent may run shell commands without asking under A.9, Credentials an agent can read under A.7, and Unreviewed components under A.4. The fix is one conversation with one developer, one file moved out of the working directory, and a review of the server. The compliance view marks A.4, A.7 and A.9 as present on that endpoint with the fix beside each, and when the findings close, the record shows the date they closed.
Why one laptop is not a management system
The worked example is one machine, and a careful engineer could have found all three by hand in an hour. A management system is not a finding. It is the ability to answer, on any day the auditor asks, which of the three hundred laptops has a readable production credential, which developer removed a permission prompt last Tuesday, which MCP server appeared yesterday and whether anyone reviewed it, and what the dated record of all that looks like for the surveillance period. The configuration changes daily, by developers and by the agents themselves. Anything done once, by hand, is out of date before the stage 1 audit.
What the auditor gets
A dated report per framework, with each Annex A control assessed from evidence where the endpoints can produce it and attested by name where they cannot. For ISO/IEC 42001 that is seven control areas from evidence and two from attestation, with the count of endpoints in scope, the findings open and closed in the period, and the exceptions with who granted them. The same evidence produces the EU AI Act view, so the two reports never disagree with each other, and a company working towards certification has its period of evidence running from the day the agents were first inventoried.
Where to start
- Draw the scope honestly. If developers run agents, the agents are in the AI management system. The shadow estate is the part nobody wrote down.
- Get the inventory, so A.4 and A.10 have a list to be assessed against.
- Read the A.9 findings first. Standing permissions are the fastest fix with the largest effect.
- Hold new MCP servers and skills for review, so A.4 and A.6 stop getting worse while you work.
- Assign A.5 and A.8 to someone by name and have them attest this quarter.
- Keep the record running. Certification takes months, and the auditor wants a period of evidence, not a snapshot.
ISO/IEC 42001 for a company that runs coding agents is not a new programme. It is the inventory, the permissions, the component review and the monitoring a security team would want anyway, organised the way an auditor reads it.
How ZYBE handles this
ZYBE reads each agent's own configuration and session activity on every developer machine, the moment it is installed, and reports names, counts and hashes, never credential values or conversation text. The rules named in this post are its rules, grouped in the console by cause with the fix beside each. The compliance view assesses ISO/IEC 42001 Annex A from that evidence, marks A.5 and A.8 as manual with the name of the person who attested, and produces the dated report per framework. Seven frameworks are assessed from the same evidence, so the ISO/IEC 42001 view never disagrees with the EU AI Act or OWASP one.
Frequently asked questions
Is ISO 42001 mandatory? No. ISO/IEC 42001 is a voluntary standard, and certification against it is a choice a company makes. It becomes a requirement in practice when a customer, a procurement questionnaire or a tender asks for it, which is how most companies first meet it. The EU AI Act is law; ISO/IEC 42001 is a way to organise the work the law asks for.
What is the difference between ISO 42001 and the EU AI Act? The EU AI Act is a regulation with legal obligations and fines, aimed at providers and deployers of AI systems in the EU. ISO/IEC 42001 is a management system standard: it describes how an organisation runs its AI governance and can be audited on it, in any country. ISO 42001 is voluntary. It can support the quality management, risk management and monitoring processes required for high-risk AI systems under the Act. The same evidence from an endpoint serves both.
Does ISO 42001 cover AI coding agents on developer machines? Yes, if the company uses them. The standard applies to organisations that provide, develop, deploy or use AI systems, and a company whose developers run Claude Code, Cursor or Codex is using AI systems under its own management. An AI management system that lists only the AI products the company sells, and not the agents its developers run, has a gap an auditor will find in the first interview.