
Blog Post
Reconciling an AI Risk Estimate Against What Truly Happened
October 2, 2026
A model produces a figure, an event happens, and somebody asks whether the figure was right. It is the obvious question and it has almost no published answer, because the comparison is harder than it looks.
A single realized loss cannot falsify a distribution. If a model puts a one percent chance on exceeding a threshold and the threshold is exceeded, the one percent case occurred, which is what the model said would sometimes happen.
Why Doesn't One Loss Settle It?
Because a prospective estimate is a distribution and an outcome is a single draw from it, so the two are different kinds of object.
Any single observation is consistent with a wide range of underlying distributions. A loss well above the modeled median is consistent with a correct model and an unlucky year, with a model that understated severity, and with a model that understated frequency. One number cannot distinguish them.
Zero Losses Prove Nothing Either
The reverse fallacy is more common and less noticed. An organization with no AI loss event and a substantial modeled exposure has not been shown that the model is wrong, since it may simply not have observed the tail yet. Zero observations is not evidence of over-estimation, and a one-in-hundred figure describes an annual probability rather than a schedule.
What Would a Real Reconciliation Require?
Many observations, and a method that reads the whole distribution rather than one threshold.

The established technique converts each realized loss into its cumulative probability under the forecast that preceded it, then asks whether those probabilities are spread out the way they should be if the model were correct. It evaluates the shape of the distribution rather than only whether a level was breached.
Which Is Why the Standard Tests Fail Here
The conventional exception-counting tests need enough observations to have statistical power, and with a short window they cannot distinguish a well-calibrated model from one that is somewhat off. An organization with two AI incidents has no test available, and saying so is more useful than running one anyway. The simulation behind the figure produces a full distribution, so a test that only counts threshold breaches discards most of what was modeled.
What Can Be Reconciled With Few Events?
The inputs, which is the available exercise and the one nobody runs.
A model rests on assertions about the environment. Forty agents hold write access, eleven systems process regulated data, the deviation rate is three percent. Each is checkable against observation, there are many of them, and they accrue continuously rather than waiting for an incident.
Which Turns One Test Into Dozens
An output reconciliation offers one observation per loss event. An input reconciliation offers one per assertion per assessment cycle, so a model with thirty inputs reviewed quarterly produces a hundred and twenty comparisons a year. A hundred and twenty is a sample size worth having, and what a modeled figure cannot establish sets out the limits it does not remove.
What Else Accumulates Without an Incident?
Near misses, which are frequency observations that most programs discard as noise.

Every blocked attempt, refused action and triggered constraint is a record of the modeled condition arising. A policy layer refusing forty actions a month says something about frequency that no absence of losses can say, and the count is available from the enforcement point rather than from an investigation.
Which Term Does That Test?
Frequency rather than magnitude, and only frequency. What a refused action would have cost is unobserved by construction, so the severity side continues to rest on judgment even where the likelihood side has real data behind it.
Which Misses Get Discovered?
Only one direction, and the asymmetry produces a drift nobody detects.
An under-estimate is discovered by the loss that exceeds it, and somebody asks why the figure was low. An over-estimate is never discovered at all, because no event prompts the question and a year without incident is read as a good year rather than as evidence about the model.
Which Argues for a Scheduled Challenge
If over-estimation has no natural detection mechanism, it needs a deliberate one. Reviewing the assumptions behind the largest scenarios on a cycle, regardless of whether anything happened, is the only route by which a conservative figure gets questioned, and auditing the inputs rather than the method is where that review has the most traction.
What Does a Single Event Legitimately Tell You?
Something about magnitude and nothing about frequency, which is a narrower conclusion than most post-incident reviews reach.
Given one loss, you can ask where its cost falls in the modeled severity distribution for that scenario. A realized cost at the ninety-ninth percentile of the modeled range is a weak signal that severity was understated. A cost near the median tells you the magnitude model was plausible. Neither says anything about whether the annual likelihood was right.
What Should Be Recorded at the Time?
Which scenario the event corresponded to, and what the model had said about that scenario before it happened, which a figure that moves in steps makes harder to reconstruct than a smooth series would. Without the prior figure the comparison cannot be made later, and reconstructing what the model would have said is not the same as knowing what it did say.
What Is the Reconciliation That Matters Most?
Not calibration at all, which is worth saying because the calibration question absorbs the effort and answers a narrower thing.
A figure exists to support a decision. The useful test is whether the ranking it produced was right, meaning whether the controls it argued for were the ones that mattered when something happened. An event arriving through a scenario the model ranked twelfth is a finding about the ordering rather than about the magnitude.
An Ordering Check Works on One Event
Unlike calibration, an ordering check works on a single observation. Where the event came through a low-ranked path, either the ranking was wrong or an unlikely thing happened, and the diagnostic is whether the inputs feeding that scenario's rank were accurate at the time.
What Does That Change About Post-Incident Review?
It adds one question to a process that already exists. Alongside what happened and what failed, asking where the model had placed this scenario and why produces a reconciliation from the one event available, and which lens decides what gets funded is what the ranking determined in the first place.
What Should Be Established?
Four things, and none requires waiting for an incident.
A dated record of each estimate, since a reconciliation compares a sequence rather than a figure. A list of the model's factual assertions about the environment, which are checkable now. A count of refused and blocked actions, which is frequency evidence accumulating already. Then a scheduled challenge of the largest scenarios, because over-estimation has no other route to discovery. AI risk quantification, or AIRQ, that keeps each figure dated with its inputs is what makes the first two possible.
Reconcile the Inputs, Not the Output
A single realized loss cannot falsify a distribution, since the modeled tail event occurring is what the model said would sometimes happen, and zero losses prove nothing either because the tail may simply not have arrived. A real reconciliation converts each outcome into its probability under the forecast that preceded it and needs many observations to say anything, which most organizations do not have for AI. What they do have is the model's factual assertions about the environment, checkable now and numerous, and a count of refused actions that tests frequency without any loss occurring. Under-estimates are discovered by the loss and over-estimates are never discovered at all, so the conservative direction needs a scheduled challenge. Kovrr's AIRQ keeps each figure dated alongside the inputs it rested on.
To see exposure figures held with the inputs and dates a reconciliation needs, book a demo mapped to your own estate.
Model Reconciliation FAQs
Speak to an ExpertDoes one large loss mean a risk model was wrong?
No. A prospective estimate is a distribution and an outcome is a single draw from it, so the two are different kinds of object. If a model puts a one percent chance on exceeding a threshold and the threshold is exceeded, the one percent case occurred, which is what the model said would sometimes happen. Any single observation is consistent with a correct model and an unlucky year, with understated severity, and with understated frequency.
Do zero losses mean a risk model was too conservative?
No, and this is the more common fallacy. An organization with no AI loss event and a substantial modeled exposure has not been shown the model is wrong, since it may simply not have observed the tail yet. Zero observations is not evidence of over-estimation, and a year without incident is evidence about the year rather than about the model.
How do you validate a risk model against realized losses?
By converting each realized loss into its cumulative probability under the forecast that preceded it, then asking whether those probabilities are spread out the way they should be if the model were correct. That evaluates the shape of the whole distribution rather than only whether a level was breached, and it needs many observations to say anything, which most organizations do not have for AI exposure.
Why do standard backtesting tests fail with few observations?
Because conventional exception-counting tests need enough observations to have statistical power, and with a short window they cannot distinguish a well-calibrated model from one that is somewhat off. An organization with two AI incidents has no test available, and saying so is more useful than running one anyway and treating the result as informative.
What can be reconciled when there are few incidents?
The model's inputs rather than its output. A model rests on assertions about the environment, such as how many agents hold write access, how many systems process regulated data, and what the deviation rate is. Each is checkable against observation, there are many of them, and a model with thirty inputs reviewed quarterly produces a hundred and twenty comparisons a year rather than one per loss event.
Why are over-estimates never discovered?
Because no event prompts the question. An under-estimate is discovered by the loss that exceeds it and somebody asks why the figure was low, while an over-estimate has no natural detection mechanism, since a year without incident is read as a good year. That asymmetry produces a drift toward conservatism, which is why the largest scenarios need a scheduled challenge regardless of whether anything happened.




