
Blog Post
Benchmarking AI Risk When Peers Disclose Nothing
September 3, 2026
Cyber risk benchmarking works for a reason that is easy to overlook. Incident disclosure is mandatory in a great many jurisdictions, so decades of events have been reported, aggregated and priced. The comparison rests on a population of observed losses.
No equivalent population exists for AI. There is no notification obligation attaching to an AI-specific failure as such, no established taxonomy for classifying one, and very little voluntary disclosure. The method does not transfer, and the shortfall is structural rather than a matter of the data being young.
Activity Is Not Exposure
The available workaround infers a peer's position from what can be observed externally. Cloud footprint, hiring volume for AI roles, which vendor platforms they run and what litigation they have attracted. All of that is measurable and it answers a different question.
A peer hiring twenty data scientists is doing more AI. Whether they are more exposed depends on what their systems touch and what their business would lose, neither of which appears in a job posting. A competitor running AI across a dozen internal workflows with nothing sensitive behind them is less exposed than one running a single model inside a lending decision, and the external signal ranks them the other way.
The Same Trap Appears Inside a Portfolio
The pattern recurs whenever exposure is inferred from posture rather than modeled. The organization with the weakest controls is frequently not the one carrying the most exposure, because exposure scales with what an entity has to lose, and aggregating across a portfolio shows the same inversion.
Benchmark the Consequence, Not the AI
The useful move splits the calculation and benchmarks the half where data genuinely exists.

Severity is not AI-specific. A business interruption at a manufacturer costs what a business interruption costs, whether the cause was ransomware, a failed migration or a model making decisions nobody checked. Regulated data leaving an organization carries notification and remediation costs determined by the data and the jurisdiction rather than by what moved it. Those figures exist in cyber loss records and industry loss data, and they transfer cleanly, which is why cyber peer benchmarking rests on firmer ground.
Frequency Is the Term That Does Not Transfer
How often an AI-specific failure occurs has no comparable base rate, and pretending otherwise is where benchmarking exercises lose credibility. Some mechanisms borrow legitimately, since credential abuse against a service account behaves like credential abuse generally and supply chain compromise has precedent in package and registry attacks. Others, including goal hijack and memory poisoning, rest on a handful of demonstrations and support a range at best, which pricing an unobservable frequency addresses differently.
What Can Be Compared Without Overclaiming
Four comparisons hold up, in descending order of how much weight they can carry.

- Your Own Position Over Time: The cleanest comparison available, since the cohort is identical by construction and only the environment changed.
- Structural Peers on Severity: Revenue, sector, data volume and regulatory footprint determine what a failure costs, and those are comparable.
- Conventional Mechanisms on Frequency: Where an AI failure runs through credential abuse or a supply chain path, existing base rates apply with adjustment.
Control implementation rates complete the list and carries the least weight, since the available figures come from surveys with self-selected respondents reporting on themselves. Useful for establishing whether you are unusual, not for establishing whether you are exposed.
Self-Comparison Is Underrated
The first item deserves more attention than it gets, because it answers the question executives usually mean. Whether the position improved is a cleaner signal than whether it beats a cohort whose membership somebody chose, and it requires no external data at all. It also exposes methodology changes, provided the model version is recorded beside each figure, since model stability determines how much movement is signal.
What a Defensible Comparison States
Four disclosures make a peer comparison usable rather than decorative, and their absence is what makes most benchmarking presentations impossible to challenge.
The cohort definition, naming how peers were selected and how many there are, since a comparison against four organizations and against four hundred are different claims. Which term is benchmarked and which is modeled, so the audience knows severity came from loss data and frequency from judgment. The vintage of the underlying data, because a base rate assembled three years ago describes a different threat environment. A range rather than a point completes it, since a figure quoted to the nearest thousand implies precision the method does not have.
Refuse the Single Peer Score
The artifact to decline is a composite blending observed loss data, self-reported survey responses and externally inferred activity into one number. It has no interpretable unit, cannot be challenged on any specific input, and moves for reasons nobody can explain. A board asking how the organization compares is better served by two or three separate statements with their sources attached, which board reporting on quantified exposure already establishes for cyber.
Where Peer Data Will Come From
The picture improves over the next several years and it is worth understanding through which channel, because that determines what becomes available.
Regulatory reporting will produce the first structured population, since serious incident obligations attaching to high-risk systems generate records where none existed. Insurance claims data follows as AI-related claims accumulate under existing policies and any affirmative coverage that emerges. Litigation produces a third channel, unrepresentative but documented in detail. None of those is available in volume today, and a benchmarking exercise claiming otherwise is describing something else.
Which Argues for Building the Internal Series Now
An organization recording its own modeled exposure quarterly has a series to compare against when peer data arrives, and a program that has never quantified will start from zero at the point everyone else has three years of history. The asymmetry is the practical argument for beginning before the external data justifies it.
What to Tell the Board Instead
Executives asking for a peer comparison are usually asking whether the organization is doing enough, and there are better answers to that question than a manufactured percentile.
Modeled exposure against the stated appetite threshold, which is an internal comparison and the one that should drive decisions. The direction of travel across the last several assessments. The proportion of the AI estate that has been assessed at all, which is frequently the most revealing figure available. Finally, where the exposure concentrates, since a small number of systems usually carry most of it. Those four say more than a peer ranking and none requires data the organization does not have.
Say Which Half You Benchmarked
Cyber benchmarking rests on a disclosure regime that has no AI equivalent, so the method transfers only in part. Severity benchmarks well, because what a failure costs depends on the business rather than on the cause. Frequency does not, except where the mechanism is conventional enough to borrow a base rate. External activity signals measure how much AI a peer runs rather than how exposed they are, and treating the two as interchangeable produces confident rankings that invert under examination. Stating the cohort, the vintage and which term came from where is what makes the comparison defensible. Kovrr's AI risk quantification reports peer figures alongside a robustness indicator for exactly that reason.
To see modeled AI exposure for your own environment with the data behind each figure, book a demo mapped to your own estate.
AI Risk Benchmarking FAQs
Speak to an ExpertWhy doesn't cyber benchmarking transfer to AI risk?
Because cyber benchmarking rests on mandatory incident disclosure across many jurisdictions, which produced decades of reported, aggregated and priced events. No equivalent population exists for AI, since no notification obligation attaches to an AI-specific failure as such, no established taxonomy exists for classifying one, and voluntary disclosure is minimal. The shortfall is structural rather than a matter of the data being young, so a method that depends on observed loss populations cannot simply be applied.
Can external signals substitute for peer disclosure?
They measure activity rather than exposure. Cloud footprint, hiring volume for AI roles, vendor platform usage and litigation history are all observable and answer a different question. A peer hiring twenty data scientists is doing more AI, while whether they are more exposed depends on what their systems touch and what their business would lose. A competitor running AI across a dozen internal workflows with nothing sensitive behind them is less exposed than one running a single model inside a lending decision.
What part of AI risk can be benchmarked reliably?
Severity, because it is not AI-specific. A business interruption at a manufacturer costs what a business interruption costs whether the cause was ransomware, a failed migration or a model making unchecked decisions. Regulated data leaving an organization carries notification and remediation costs determined by the data and jurisdiction rather than by what moved it. Those figures exist in cyber loss records and industry loss data and transfer cleanly. Frequency is the term that does not transfer, except where the mechanism is conventional.
What comparisons hold up when peer data is thin?
Four, in descending order of weight. Your own position over time, which is the cleanest available since the cohort is identical by construction. Structural peers on severity, where revenue, sector, data volume and regulatory footprint determine what a failure costs. Conventional mechanisms on frequency, where an AI failure runs through credential abuse or a supply chain path that has existing base rates. And control implementation rates, which carry least weight since the figures come from surveys with self-selected respondents reporting on themselves.
What should a defensible peer comparison disclose?
Four things whose absence makes most benchmarking presentations impossible to challenge. The cohort definition, naming how peers were selected and how many there are, since a comparison against four organizations and four hundred are different claims. Which term is benchmarked and which is modeled. The vintage of the underlying data, because a base rate assembled years ago describes a different environment. And a range rather than a point, since a figure quoted precisely implies confidence the method does not support.
What should a board be told instead of a peer ranking?
Executives asking for a comparison usually want to know whether the organization is doing enough, and four internal figures answer that better than a manufactured percentile. Modeled exposure against the stated appetite threshold. The direction of travel across recent assessments. The proportion of the AI estate assessed at all, which is frequently the most revealing number available. And where the exposure concentrates, since a small number of systems usually carry most of it. None requires data the organization does not have.




