Blog Post

When the Model Disagrees With Your Security Team

August 27, 2026

Table of Contents

A model ranks phishing sixth. The security team has spent three years on phishing and knows how often people click. Somebody in the room concludes the model is wrong, or that the security team is attached to its own program, and the meeting stops being useful.

Most of these disagreements are not about risk. They are about which question each side answered, and establishing that first resolves a surprising proportion of them without anyone conceding anything.

Exploitability and Loss Are Different Questions

Security teams are usually estimating how readily something can be exploited. A model is usually estimating how much would be lost. Those come apart cleanly and both answers can be correct at once.

A trivially exploitable weakness on a system holding nothing of consequence scores high on the first question and low on the second. A well-defended system carrying the majority of revenue scores the reverse. Neither party is mistaken, and the argument continues until somebody notices they are measuring different things.

Ask What Each Side Was Ranking

The opening question in any of these discussions is what the ranking was ordered by. Where the security view is ordered by likelihood of compromise and the model output by expected loss, the disagreement dissolves into two lists that serve different purposes, both worth having. Where both claim to rank the same thing, there is a genuine dispute worth investigating, and risk-focused prioritization depends on settling it.

The Direction Tells You Where to Look

Once both sides are answering the same question, which way the disagreement runs narrows the investigation considerably.

Initial access techniques ranked by annual loss, with valid accounts carrying the largest share and phishing appearing considerably further down the list
Ranking access techniques by the loss each carries produces an order that frequently differs from the one a security team would give from experience.

Team High, Model Low

Look first for a dependency the model does not know about. Practitioners frequently hold information about internal connections, a process that depends on one undocumented system, or a recovery path that has never been tested, none of which appears in an asset register. The second place to look is asset valuation, since a model working from recorded values will understate anything the business has grown around without updating the record.

Model High, Team Low

Look first for a compensating control the model has not been told about. Teams routinely operate mitigations that never made it into a documented control set, particularly architectural ones such as isolation or manual approval steps. The second place to look is a tail the team has never experienced, since practitioners estimate from what they have seen and the events that dominate a loss distribution are ones most people have not lived through.

Frequency Disputes and Severity Disputes Resolve Differently

Separating which half of the calculation is contested changes who should settle it, and programs regularly send the question to the wrong people.

Scenario metrics comparing modeled annual event likelihood and average financial loss against peer figures, with a robustness indicator for the underlying data
Comparing a contested figure against peer data and a robustness indicator turns a disagreement into something checkable.

Frequency Belongs With Evidence

A dispute about how often something happens is settled against the organization's own incident history first and peer data second. Where the model sits well below a peer base rate, that difference needs an explanation somebody can state, and where it sits above, the same applies. A team asserting a frequency without either source is offering an impression, which is worth listening to and not worth substituting for the number.

Severity Belongs With Finance

This is the one that surprises people. Security teams are frequently poor estimators of business loss, not through any deficiency but because they do not own the revenue line and have never had to account for what an outage costs. A severity dispute should be taken to the people who do, and their answer usually differs from both prior estimates. Measuring exposure depends on that input being sourced correctly.

The Model Should Sometimes Lose

A quantification function that never revises a figure after challenge is not operating a model. It is defending a position, and everyone involved works that out quickly.

Revision after a substantive challenge is the strongest available evidence that the process functions. It also changes how the output is received, since a team that has successfully corrected the model once engages with subsequent figures rather than dismissing them. The alternative produces a number nobody argues with because nobody believes it matters.

Record What Changed and Why

Each revision should leave a note stating what was contested, what evidence was produced and whether the figure moved. The note answers a later question about why an assumption is what it is, and it distinguishes a model that was corrected from one that was negotiated. A register built for decisions holds that exchange alongside the conclusion.

When the Practitioners Are Wrong

Expert judgment carries known failure modes, and naming them is fairer than treating every objection as equally weighted.

  • Recency: A painful incident eighteen months ago inflates the estimate for that category well past what the evidence supports.
  • Attachment: A control somebody designed and defended is difficult to score objectively, particularly where the model implies it removes little.
  • Availability: Threats with vivid narratives feel more probable than quiet ones that carry more loss.

The last of these explains the phishing case at the top of this piece. Credential abuse produces more modeled loss than phishing in most portfolios, and it is far less memorable because nothing dramatic happens at the moment of compromise. A team estimating from experience will rank the memorable one higher, and the ranking by modeled exposure consistently disagrees.

What to Do With an Unresolved Disagreement

Some disputes do not settle, usually because the contested input has no available evidence on either side. Forcing a resolution produces false confidence.

The workable response models both positions and reports the range, stating which assumption drives the difference. A decision robust across both estimates can proceed without settling the argument, which is frequently the case. Model stability determines how much difference between the two is worth arguing about at all. Where the decision flips between them, the disagreement has been converted into a specific question worth spending money to answer, which is considerably more useful than a compromise figure neither party believes.

Disagreement Is Information

A model and a security team producing the same ranking would mean one of them was unnecessary. The divergence is where the useful conversation lives, provided it starts by establishing which question each side answered and then separates frequency disputes from severity disputes, since those resolve through different evidence and different people. Revising the model when the challenge is sound is what makes the next challenge worth having. Kovrr's cyber risk quantification exposes the assumptions and peer comparisons behind each figure, which is what allows a disagreement to be about something specific.

To see the assumptions and peer data behind a modeled figure rather than the figure alone, book a demo with our cyber risk experts.

Shalom Bublil

Kovrr Co-founder & Chief Product Officer

Model Disagreement FAQs

Speak to an Expert

Why does a risk model often disagree with the security team?

What does the direction of the disagreement indicate?

Who should settle a frequency dispute?

Who should settle a severity dispute?

Should a model ever be revised after a challenge?

What if the disagreement cannot be resolved?