
Blog Post
Auditing a Cyber Risk Model You Did Not Build
September 6, 2026
Somebody in second line or internal audit is asked to validate a quantified cyber exposure figure. They did not build the model, do not have access to its internals, and in most cases have no statistical training. The request is reasonable and the standard guidance is not much help.
Model validation practice says to recalculate outputs independently, challenge the distribution family, evaluate the correlation structure and run sensitivity analysis. Three of those four require capability the reviewer does not have, and a function attempting them anyway ends up testing documentation because documentation is what it can test.
Which Audit Are You Running?
Two are available and they answer different questions. Choosing deliberately is better than drifting into the second while describing the first.
A model validation examines the arithmetic. Whether the coded logic matches the written methodology, whether the distribution family suits a domain with rare severe events, whether independence was assumed where dependency exists. It requires somebody who can evaluate those choices and it is the right exercise where that person is available.
An input and boundary audit examines everything the model consumes and everything it claims. What was included, where each figure came from, how old it is, which inputs are measured and which are judgment. It requires no statistics and it catches more real errors, because most errors in practice are in the inputs rather than in the mathematics.
Why Are the Errors Mostly Upstream?
Because the engine is usually a competent implementation of a published method and the inputs are assembled locally under time pressure. Asset values carried forward from an older finance cycle, control state taken from a self-assessment, a scope that quietly excludes an acquired business. None of those is a modeling error and all of them move the answer more than a distribution choice would, and what a modeled figure cannot tell you covers the limits the inputs impose.
What Does the Input Audit Test?
Six things, and each is checkable by somebody with no specialist background.

- Scope: What is in and out, and whether the excluded portion was a decision or an omission nobody noticed.
- Asset values: Their vintage and source, and whether they reconcile to what finance currently reports.
- Control state: Where it came from, and whether each control was verified or attested by the team that owns it.
The remaining three concern provenance. Which frequency inputs are drawn from observed data and which are expert judgment, since both usually appear in the same notation. Whether a correlation assumption is stated at all, which requires no ability to evaluate it. Then whether the run is reproducible, meaning the same inputs and the same model version produce the same figure, which the number of simulation trials affects directly.
Which Single Test Yields the Most?
Sensitivity analysis, and it is the one item from the statistical list a non-specialist can run. Change one input, rerun, and see how far the output moves.
The interpretation needs no theory. Where a modest change in a minor input swings the result substantially, the model is unstable on that input and the figure should not be quoted to a precision it cannot support. Where the output barely moves against a large change in an input somebody was worried about, that worry can be set aside. Neither conclusion requires understanding why.
Which Inputs Should Be Tested First?
The ones with the weakest provenance rather than the ones that look most important. An asset value reconciled to audited accounts is unlikely to be wrong. A frequency estimate somebody provided in a workshop is the candidate, and discovering the output is insensitive to it is a genuinely reassuring finding.
How Do You Test Reproducibility Without Access?
Ask for the same run twice and compare, which requires no internal access at all and catches a specific and common problem.

Simulation-based models vary between runs by design, so two runs will not match exactly and should land close. A large difference indicates too few iterations for the figure being quoted. No difference at all indicates a cached result rather than a rerun. A difference nobody can explain because the model version changed between runs is the finding that matters most, since it means the series cannot be read as a trend.
What Should Be Recorded Against Every Figure?
The model version, the run date and the input vintages. A figure without those cannot be compared to a later one, so a program producing quarterly numbers with no version record has produced four unrelated results rather than a trend, and statistical significance determines how much movement between two runs counts as real.
What Can the Opinion Conclude?
Less than people expect, and stating the limit is what makes the rest credible.
An audit can conclude that the inputs are sourced and current, that the method is documented and followed, that assumptions are stated, that the output is reproducible and that sensitivity is acceptable on the weakest inputs. Each of those is a checkable proposition.
It cannot conclude that the number is right. There is no ground truth to compare against, since the organization experiences one outcome and a distribution describes a population. An opinion claiming the figure is accurate has overreached, and one concluding the process producing it is sound has said the strongest available thing.
What Should the Reviewer Ask For?
Four artifacts, requested before fieldwork rather than during it, because assembling them takes the model owner time and the delay is where audits lose weeks.
The scope statement naming what is included and excluded. The input register listing each figure, its source and its date. The assumption list, particularly correlation and any frequency input that is judgment. Then the change log showing what moved in the model since the last assessment. A model owner who can produce those four quickly is running a program somebody else could pick up, and one who cannot has told you something before the testing starts.
What Does It Mean If Those Do Not Exist?
The finding becomes one about the program rather than about the model. A figure nobody can reproduce, from inputs nobody recorded, under assumptions nobody wrote down, may well be correct and cannot be relied upon, and that conclusion is available without evaluating any mathematics. A reviewing function that cannot evaluate what it reviews can still evaluate whether the evidence exists.
How Should the Findings Be Written?
As statements about evidence rather than about the number, since a finding disputing the figure invites a defense the reviewer cannot win and a finding about a missing input does not.
Saying the exposure appears overstated puts the auditor in an argument about modeling with somebody who models for a living. Saying asset values were drawn from a finance cycle two years old, and that the output moves substantially when they are updated, states two verifiable facts and leaves the conclusion to follow. The second is harder to rebut and considerably more likely to change something.
Which Findings Carry the Most Weight?
Those naming a specific input, its provenance and its effect. A finding stating that control effectiveness was self-attested by the teams operating the controls, with no independent verification, and that the output falls by a stated proportion when those ratings are reduced, is a complete argument. Continuously verified control state is what closes it.
What About Findings the Model Owner Agrees With?
Worth recording anyway, and frequently the most useful part of the exercise. Model owners usually know their weakest inputs and lack the standing to get them fixed, so an audit finding naming the same weakness supplies leverage somebody inside the team could not generate. An audit that surfaces nothing the owner disputes has not failed.
When Is Full Model Validation Worth Commissioning?
Where the figure drives a capital decision, a regulatory submission or an insurance placement large enough that the arithmetic itself needs independent examination, which sizing a program from a distribution depends on.
It is a specialist engagement rather than an internal audit exercise, and treating it as the default sets internal audit an impossible task. The proportionate arrangement runs the input and boundary audit internally each cycle and commissions a full validation when the stakes or the methodology change materially. Cyber risk quantification that records its own inputs, versions and assumptions makes the first affordable and the second rarer.
Audit the Inputs, Commission the Arithmetic
Standard model validation guidance asks a reviewer to recalculate outputs, challenge distribution choices and evaluate correlation structure, which requires capability most internal audit functions do not hold. An input and boundary audit needs none of it and catches more, because the errors in practice are stale asset values, self-attested control state and scope nobody examined. Sensitivity analysis is the one statistical test a non-specialist can run and interpret. The strongest available opinion is also that the process is sound rather than that the number is correct, because no ground truth exists to check it against. Kovrr's CRQ records the model version, input dates and assumptions against each run, which is what makes that audit a retrieval rather than an investigation.
To see exposure figures with their inputs, versions and assumptions recorded against each run, book a demo with our cyber risk experts.
Model Audit FAQs
Speak to an ExpertWhat are the two audits available on a cyber risk model?
A model validation examines the arithmetic, covering whether coded logic matches written methodology, whether the distribution family suits a domain with rare severe events, and whether independence was assumed where dependency exists. It requires somebody who can evaluate those choices. An input and boundary audit examines everything the model consumes and claims, covering what was included, where each figure came from, how old it is and which inputs are judgment. It requires no statistics and catches more real errors.
Why are most errors in the inputs rather than the model?
Because the engine is usually a competent implementation of a published method while the inputs are assembled locally under time pressure. Asset values carried forward from an older finance cycle, control state taken from a self-assessment, and a scope that quietly excludes an acquired business are all common. None is a modeling error and each moves the answer more than a distribution choice would, which is why an audit aimed at the inputs finds more than one aimed at the mathematics.
What does an input and boundary audit test?
Six things, each checkable without a specialist background. Scope, meaning what is in and out and whether exclusions were decisions or oversights. Asset values, their vintage and source and whether they reconcile to what finance reports. Control state, where it came from and whether each control was verified or attested. Which frequency inputs are observed data and which are judgment, since both appear in the same notation. Whether a correlation assumption is stated. Then whether the run is reproducible.
Which single test yields the most?
Sensitivity analysis, the one item from the statistical list a non-specialist can run. Change one input, rerun and see how far the output moves. Where a modest change in a minor input swings the result substantially, the model is unstable on that input and the figure should not be quoted to a precision it cannot support. Test the inputs with the weakest provenance first rather than the ones that look most important.
How do you test reproducibility without internal access?
By asking for the same run twice and comparing. Simulation-based models vary between runs by design, so two runs should land close without matching exactly. A large difference indicates too few iterations for the figure being quoted. No difference at all indicates a cached result rather than a rerun. A difference nobody can explain because the model version changed between runs means the series cannot be read as a trend.
What can the audit opinion conclude?
That the inputs are sourced and current, the method is documented and followed, assumptions are stated, the output is reproducible and sensitivity is acceptable on the weakest inputs. Each is a checkable proposition. It cannot conclude the number is right, since no ground truth exists to compare against because the organization experiences one outcome while a distribution describes a population. An opinion claiming the figure is accurate has overreached.




