Insights
Building a Model Risk Taxonomy for AI Governance

A model risk taxonomy is where AI governance becomes operational. Without a common classification system, every team describes risk differently: engineering focuses on model performance, legal on obligations, security on data exposure, and executives on business impact. The result is a collection of disconnected reviews that cannot show a clear enterprise risk posture.
For organizations operating AI in production, the taxonomy should do more than label systems as high, medium, or low risk. It should create a consistent way to identify what a model does, where it is used, what can go wrong, which controls apply, who owns the decision, and what evidence proves those controls are working.
What Is a Model Risk Taxonomy?
A model risk taxonomy is a structured classification of the risks associated with developing, procuring, deploying, and operating models. In an AI context, it covers traditional quantitative models, machine learning systems, generative AI applications, vendor-hosted models, and automated decision systems.
The purpose is not to create another policy document. Its purpose is to give risk, compliance, product, engineering, security, and finance teams a shared operational language. A well-designed taxonomy connects abstract policy requirements to actual production systems, workflows, vendors, data sources, users, and business outcomes.
That distinction matters. A policy may state that AI systems handling sensitive data require enhanced review. A taxonomy makes that requirement executable by defining sensitive-data exposure, identifying affected systems, assigning the required controls, and retaining evidence that those controls were completed.
Why AI Requires a Different Approach to Model Risk
Traditional model risk management was often built around models used for credit, pricing, forecasting, fraud, or capital planning. The core concerns were validity, documentation, assumptions, performance, and independent review. Those remain relevant, but production AI introduces additional layers of exposure.
Generative AI can create inaccurate or harmful outputs at scale. A third-party model may change its behavior without notice. A retrieval system can expose confidential information through poor access controls. An agent may take actions in downstream systems rather than simply generate text. Usage-based pricing can also make an apparently low-risk pilot a material cost-control issue once adoption spreads.
A useful taxonomy recognizes that risk is contextual. The same foundation model may be low risk when it summarizes public meeting notes and high risk when it drafts customer eligibility decisions, processes employee information, or triggers transactions. Risk is determined by the use case, data, autonomy, audience, and consequence of failure - not by the model name alone.
The Core Dimensions of a Model Risk Taxonomy
An enterprise taxonomy should be detailed enough to drive differentiated controls, but simple enough that operators can apply it consistently. Most organizations need to classify AI systems across several dimensions rather than rely on one overall score.
Business purpose and decision impact
Start with what the system is intended to do and how its output is used. Is it assisting an employee, recommending an action, making a decision, or executing an action automatically? A model used for internal research requires a different control posture than one that influences hiring, lending, healthcare, customer access, or financial reporting.
Decision impact should also account for reversibility. A recommendation reviewed by a trained employee is not equivalent to an automated denial or transaction. Human oversight can reduce risk, but only when the reviewer has enough context, authority, time, and capability to challenge the output.
Data sensitivity and access
Classify the data that enters, trains, grounds, or is generated by the system. Relevant categories often include public, internal, confidential, personal, sensitive personal, regulated, and highly restricted data. The taxonomy should also capture whether data is retained by a provider, used for model training, transferred across jurisdictions, or accessible through connected tools.
This dimension should extend beyond prompts. Retrieval sources, vector stores, system logs, tool outputs, and conversation histories can each create exposure. Treating prompt content as the only data concern leaves material gaps in the control design.
Model and supplier characteristics
Teams need visibility into whether a model is internally developed, open source, accessed through an API, embedded in a SaaS product, or supplied as part of a managed service. Each arrangement changes the organization's ability to test, monitor, document, and control the system.
For vendor models, the taxonomy should capture supplier criticality, contractual restrictions, model versioning practices, geographic processing, incident notification commitments, and available audit artifacts. A provider's reputation is not a substitute for a documented control assessment.
Operational behavior and autonomy
Classify how the AI system behaves in production. Does it generate content, rank options, predict an outcome, retrieve information, call external tools, or initiate actions? Does it operate in a closed environment, or can it interact with customers, employees, code repositories, payment systems, or production databases?
Autonomy is especially consequential. A system that drafts an email for approval has a different failure path than an agent that sends the email, updates a record, or initiates a refund. The taxonomy should distinguish assistance, recommendation, supervised execution, and autonomous execution.
Harm, compliance, and financial exposure
The taxonomy must address the consequences of a failure. These may include consumer harm, discrimination, privacy violations, intellectual property concerns, security incidents, regulatory breach, contractual exposure, reputational damage, or material financial loss.
Not every risk category applies to every system. A sensible structure allows teams to mark categories as not applicable while documenting why. Forced, generic assessments create noise and make genuine exceptions harder to see.
From Classification to Controls
A taxonomy has limited value if it ends in a register. Its real value comes from a defined relationship between risk classes and operational requirements.
For example, an AI system classified as customer-facing, high-impact, and sensitive-data enabled might require pre-deployment approval, documented testing, access controls, vendor due diligence, human escalation paths, output monitoring, periodic review, and executive reporting. A low-impact internal writing assistant may require a lighter process focused on approved providers, data-use restrictions, user guidance, and basic monitoring.
This is not about applying maximum controls everywhere. Over-control encourages teams to bypass governance, especially when AI experimentation moves faster than review cycles. The better goal is proportionate governance: stronger controls where consequences are higher, faster pathways where exposure is limited, and a clear escalation route when a use case changes.
Design the Taxonomy Around Evidence
Audit readiness should influence taxonomy design from the beginning. If a risk category cannot be tied to evidence, it will be difficult to defend under scrutiny.
For every classification and required control, define the evidence artifact, owner, review frequency, system of record, and approval status. Evidence might include data-flow assessments, test results, provider assessments, model cards, approval records, incident logs, monitoring results, user attestations, or documented exceptions.
The critical question is not whether a control was described in a policy. It is whether the organization can demonstrate that the control applied to a particular production use case at a particular point in time. This is where spreadsheets and static questionnaires begin to fail. They are quickly outdated when models, providers, prompts, integrations, users, and policies change.
An operational governance layer can connect taxonomy classifications to live inventories, control workflows, alerts, and reporting. Platforms such as Onaro Meridian are designed to turn those connections into continuous oversight rather than a periodic documentation exercise.
Common Taxonomy Failures
The most common failure is using a single risk score as the entire taxonomy. An overall score can be useful for prioritization, but it conceals the reasons behind the rating. Two systems with the same score may require entirely different controls because one is data-sensitive and the other is highly autonomous.
Another failure is treating the taxonomy as an annual compliance artifact. Production AI changes through new prompts, features, connected data, model updates, user groups, and vendors. Classification must be revisited when those changes alter the system's risk profile.
Organizations also struggle when ownership is vague. Product teams may own use-case documentation, engineering may own technical controls, security may own access requirements, legal may own regulatory interpretation, and risk may own classification standards. Those responsibilities can coexist, but the taxonomy must make accountability explicit.
Finally, avoid importing a framework wholesale without adapting it to the business. Regulatory expectations, industry obligations, internal risk appetite, and the actual ways teams use AI should determine the level of detail. A taxonomy that no one can apply reliably is not more mature because it has more categories.
A Practical Starting Point
Begin by inventorying AI systems that are already operating or being evaluated. Include shadow usage where possible, especially AI capabilities embedded in existing SaaS tools. Then define a small number of classification dimensions, establish control requirements for meaningful combinations of risk, and pilot the approach with several contrasting use cases.
Test whether different reviewers reach the same classification. Test whether the assigned controls can actually be completed within delivery timelines. Test whether leadership can understand the resulting posture without reading technical documentation. These are stronger signs of taxonomy quality than the length of the policy behind it.
A model risk taxonomy should give the organization a repeatable way to make defensible decisions as AI use expands. When it is connected to real systems, accountable owners, and current evidence, governance stops being a gate at the edge of innovation and becomes part of how the enterprise operates safely at scale.

About Brian Diamond
Brian Diamond is a fractional Chief AI Officer who works with mid-market and enterprise organizations on AI strategy, governance, and operations. In 2001 he founded LanStatus, a managed services provider based in Trumbull, Connecticut, with named partnerships across Microsoft, HPE, Citrix, and VMware. He brings 25 years of infrastructure operations to AI leadership and publishes the CAIO Brief.
Also publishes at: day9.coffee · ChiliStation · PlotLuck · Beacon
Subscribe to the CAIO Brief for practical AI leadership every week.
Request an Onaro demo