Blog Post

Reconciling an AI Risk Estimate Against What Truly Happened

October 2, 2026

Table of Contents

A model produces a figure, an event happens, and somebody asks whether the figure was right. It is the obvious question and it has almost no published answer, because the comparison is harder than it looks.

‍

A single realized loss cannot falsify a distribution. If a model puts a one percent chance on exceeding a threshold and the threshold is exceeded, the one percent case occurred, which is what the model said would sometimes happen.

‍

Why Doesn't One Loss Settle It?

‍

Because a prospective estimate is a distribution and an outcome is a single draw from it, so the two are different kinds of object.

‍

Any single observation is consistent with a wide range of underlying distributions. A loss well above the modeled median is consistent with a correct model and an unlucky year, with a model that understated severity, and with a model that understated frequency. One number cannot distinguish them.

‍

Zero Losses Prove Nothing Either

‍

The reverse fallacy is more common and less noticed. An organization with no AI loss event and a substantial modeled exposure has not been shown that the model is wrong, since it may simply not have observed the tail yet. Zero observations is not evidence of over-estimation, and a one-in-hundred figure describes an annual probability rather than a schedule.

‍

What Would a Real Reconciliation Require?

‍

Many observations, and a method that reads the whole distribution rather than one threshold.

‍

Exposure trended across successive assessments alongside a change log recording what moved between them and when
A series of dated estimates is the minimum structure a reconciliation needs, since the exercise compares a sequence rather than a single figure.

The established technique converts each realized loss into its cumulative probability under the forecast that preceded it, then asks whether those probabilities are spread out the way they should be if the model were correct. It evaluates the shape of the distribution rather than only whether a level was breached.

‍

Which Is Why the Standard Tests Fail Here

‍

The conventional exception-counting tests need enough observations to have statistical power, and with a short window they cannot distinguish a well-calibrated model from one that is somewhat off. An organization with two AI incidents has no test available, and saying so is more useful than running one anyway. The simulation behind the figure produces a full distribution, so a test that only counts threshold breaches discards most of what was modeled.

‍

What Can Be Reconciled With Few Events?

‍

The inputs, which is the available exercise and the one nobody runs.

‍

A model rests on assertions about the environment. Forty agents hold write access, eleven systems process regulated data, the deviation rate is three percent. Each is checkable against observation, there are many of them, and they accrue continuously rather than waiting for an incident.

‍

Which Turns One Test Into Dozens

‍

An output reconciliation offers one observation per loss event. An input reconciliation offers one per assertion per assessment cycle, so a model with thirty inputs reviewed quarterly produces a hundred and twenty comparisons a year. A hundred and twenty is a sample size worth having, and what a modeled figure cannot establish sets out the limits it does not remove.

‍

What Else Accumulates Without an Incident?

‍

Near misses, which are frequency observations that most programs discard as noise.

‍

Portfolio view showing inherent and residual annual loss, annual likelihood and the reduction attributable to controls, with an aggregate exceedance curve
A stated likelihood is the term near-miss counts are able to test, since each refused action is an observation about how often the condition arises.

Every blocked attempt, refused action and triggered constraint is a record of the modeled condition arising. A policy layer refusing forty actions a month says something about frequency that no absence of losses can say, and the count is available from the enforcement point rather than from an investigation.

‍

Which Term Does That Test?

‍

Frequency rather than magnitude, and only frequency. What a refused action would have cost is unobserved by construction, so the severity side continues to rest on judgment even where the likelihood side has real data behind it.

‍

Which Misses Get Discovered?

‍

Only one direction, and the asymmetry produces a drift nobody detects.

‍

An under-estimate is discovered by the loss that exceeds it, and somebody asks why the figure was low. An over-estimate is never discovered at all, because no event prompts the question and a year without incident is read as a good year rather than as evidence about the model.

‍

Which Argues for a Scheduled Challenge

‍

If over-estimation has no natural detection mechanism, it needs a deliberate one. Reviewing the assumptions behind the largest scenarios on a cycle, regardless of whether anything happened, is the only route by which a conservative figure gets questioned, and auditing the inputs rather than the method is where that review has the most traction.

‍

What Does a Single Event Legitimately Tell You?

‍

Something about magnitude and nothing about frequency, which is a narrower conclusion than most post-incident reviews reach.

‍

Given one loss, you can ask where its cost falls in the modeled severity distribution for that scenario. A realized cost at the ninety-ninth percentile of the modeled range is a weak signal that severity was understated. A cost near the median tells you the magnitude model was plausible. Neither says anything about whether the annual likelihood was right.

‍

What Should Be Recorded at the Time?

‍

Which scenario the event corresponded to, and what the model had said about that scenario before it happened, which a figure that moves in steps makes harder to reconstruct than a smooth series would. Without the prior figure the comparison cannot be made later, and reconstructing what the model would have said is not the same as knowing what it did say.

‍

What Is the Reconciliation That Matters Most?

‍

Not calibration at all, which is worth saying because the calibration question absorbs the effort and answers a narrower thing.

‍

A figure exists to support a decision. The useful test is whether the ranking it produced was right, meaning whether the controls it argued for were the ones that mattered when something happened. An event arriving through a scenario the model ranked twelfth is a finding about the ordering rather than about the magnitude.

‍

An Ordering Check Works on One Event

‍

Unlike calibration, an ordering check works on a single observation. Where the event came through a low-ranked path, either the ranking was wrong or an unlikely thing happened, and the diagnostic is whether the inputs feeding that scenario's rank were accurate at the time.

‍

What Does That Change About Post-Incident Review?

‍

It adds one question to a process that already exists. Alongside what happened and what failed, asking where the model had placed this scenario and why produces a reconciliation from the one event available, and which lens decides what gets funded is what the ranking determined in the first place.

‍

What Should Be Established?

‍

Four things, and none requires waiting for an incident.

‍

A dated record of each estimate, since a reconciliation compares a sequence rather than a figure. A list of the model's factual assertions about the environment, which are checkable now. A count of refused and blocked actions, which is frequency evidence accumulating already. Then a scheduled challenge of the largest scenarios, because over-estimation has no other route to discovery. AI risk quantification, or AIRQ, that keeps each figure dated with its inputs is what makes the first two possible.

‍

Reconcile the Inputs, Not the Output

‍

A single realized loss cannot falsify a distribution, since the modeled tail event occurring is what the model said would sometimes happen, and zero losses prove nothing either because the tail may simply not have arrived. A real reconciliation converts each outcome into its probability under the forecast that preceded it and needs many observations to say anything, which most organizations do not have for AI. What they do have is the model's factual assertions about the environment, checkable now and numerous, and a count of refused actions that tests frequency without any loss occurring. Under-estimates are discovered by the loss and over-estimates are never discovered at all, so the conservative direction needs a scheduled challenge. Kovrr's AIRQ keeps each figure dated alongside the inputs it rested on.

‍

To see exposure figures held with the inputs and dates a reconciliation needs, book a demo mapped to your own estate.

Yakir Golan

CEO

Model Reconciliation FAQs

Speak to an Expert

Does one large loss mean a risk model was wrong?

Do zero losses mean a risk model was too conservative?

How do you validate a risk model against realized losses?

Why do standard backtesting tests fail with few observations?

What can be reconciled when there are few incidents?

Why are over-estimates never discovered?