Blog Post

A Production Line Does Not Restore From Backup

September 15, 2026

Table of Contents

Recovery time objectives in manufacturing are usually set by information technology. Restore the servers, rebuild the domain, bring the applications back, declare recovery. Those milestones are measurable and a plan built around them looks complete.

None of them produces a part. A production line does not resume when the systems are clean, it resumes when the process has been brought back under control and, in regulated manufacturing, when somebody has certified that what comes off it can be sold.

Why Can't a Line Be Restored From Backup?

Because the thing being restored is a physical process rather than a dataset, and the controllers driving it hold state that a file-level backup does not represent.

Programmable controllers, operator interfaces and supervisory nodes carry logic and configuration that may exist only on the devices themselves. Where those are compromised, each has to be treated as suspect, wiped and reflashed from vendor media rather than restored, and the control logic verified against engineering documentation that in older plants may be a paper drawing. Quantifying operational technology exposure runs into the same evidence problem.

The Recovery Is a Reconstruction

An information system restore is a copy operation. Rebuilding a control environment is an engineering exercise conducted by people who understand the process, in an isolated environment, device by device. The two differ in kind rather than in duration, and a plan estimating the second from experience of the first will be wrong by a large factor.

What Sits Between Clean Systems and Sellable Output?

Three stages, none of which appears in an information technology recovery plan.

Outage duration exceedance curve showing the likelihood of an event exceeding a working day, twelve hours, twenty-four hours and forty-eight hours
A duration curve built on system restore times describes the wrong event where output resumes considerably later than the systems do.

Recommissioning brings equipment back to a known operating state, with verified logic, calibrated instruments and a controlled restart sequence. Qualification re-establishes that the process performs as validated, which in regulated sectors is a documented exercise rather than an engineering judgment. Release then requires somebody to certify that material produced during and after the event meets specification.

Which Stage Dominates the Duration?

Qualification, where it applies, because it is procedural rather than technical. The equipment may be ready in hours and the documented requalification takes as long as the protocol takes, with a quality function rather than an engineering one controlling the pace, and testing the recovery on a cadence is what produces a credible estimate of it. A plant that has restored everything and cannot ship is in the most expensive state available.

What Happens to Material Made During the Event?

It becomes questionable, which is a loss category with no information security analogue and one that scales with how long the compromise went unnoticed.

Where an attacker had access to control systems, the integrity of what was produced under that access is uncertain. Product made during the dwell period may need quarantine, retesting or disposal, and the decision depends on evidence about what was altered rather than on whether anything visibly went wrong. Dwell time therefore converts directly into inventory at risk.

Which Changes What Detection Is Worth

In most sectors, faster detection reduces the data exfiltrated. Here it reduces the volume of product whose provenance cannot be established, which is a physical quantity with a unit cost. The value of detection becomes calculable in a way it rarely is, and choosing an indicator that leads rather than lags matters more where the consequence accrues per hour.

How Should the Duration Curve Be Built?

From the restart, not from the restore, and the two need modeling as separate stages with different drivers.

Annual loss broken down by event type, impact scenario and damage type showing which categories contribute most
Separating the stages of an outage in the model is what prevents a recovery estimate built on system restore times from setting the whole figure.

The first stage runs from compromise to clean systems, driven by the response capability and by whether backups survived. The second runs from clean systems to first sellable output, driven by recommissioning effort, qualification requirements and inspection availability. Summing an estimate for the first with an assumption for the second produces the figure most plans carry.

Where Do the Second-Stage Inputs Come From?

Operations and quality, neither of which is usually in the room. How long a controlled restart takes from cold, what the requalification protocol requires, and whether the people or external bodies who sign it off are available at short notice. Those are known internally and absent from every security assessment, which modeling interruption rather than data loss covers as a general pattern.

Is the Lost Output Recoverable?

Partly, and the answer determines the shape of the cost curve rather than a detail within it.

Where a line can run extra hours afterward, the loss is the cost of catching up, which is overtime and expedited logistics rather than the value of what was not made. Where the line was already at capacity, or the product is perishable, or the delivery window has closed, the output is gone and the loss is its full value. Most plants are a mixture, and the proportion is a known operational fact.

What Determines Which Applies?

Spare capacity and commitment structure. A plant running below capacity against forecast demand can recover most of it. One running at capacity against committed delivery schedules cannot, and the consequence propagates to customers as late delivery with contractual effect rather than staying inside the plant.

What Does the Downstream Effect Add?

A category that lands on somebody else's balance sheet and returns as a claim or a lost contract.

A component maker supplying an assembly line elsewhere causes that line to stop. The loss is the customer's production, and whether it returns as a penalty, a claim or a non-renewal depends on the contract and the relationship. The exposure is invisible in a model built from the plant's own revenue, and it is frequently larger, which concentration in a customer base compounds where a few customers take most of the output.

How Is It Estimated?

From contractual terms rather than from loss data. Delivery commitments, penalty clauses and consequential loss provisions in the largest customer agreements state the exposure directly, and reading the top five contracts produces a better figure than any benchmark. Cyber risk quantification built from those inputs prices what an outage costs past the fence.

Does This Change Where the Money Goes?

It changes what the money buys, since the two stages respond to entirely different investments and most spending addresses only the first.

Detection and response capability shortens the compromise-to-clean stage, which is where security budgets go and where the improvement is real. Nothing in that budget shortens the clean-to-output stage, which is bounded by engineering effort and by protocol. Halving the first stage on a plant where the second dominates produces a smaller reduction in total outage than the investment case assumed.

What Shortens the Second Stage?

Three things, none of them a security control. Offline copies of controller logic and configuration held somewhere other than the devices, which converts a reconstruction into a restore. A rehearsed cold-start procedure, since a restart nobody has rehearsed takes longer than one that has been. Then a pre-agreed requalification approach with whoever signs it, since negotiating the protocol during an outage adds days.

Which Investment Ranks Higher?

Whichever addresses the dominant stage, and that ordering is establishable rather than assumed. A plant that can state both durations knows which half of the outage it is paying for, and ranking by what each measure removes then puts the offline logic copy and the detection upgrade in the same list.

What Should Be Established Before an Incident?

Four things, all knowable now and none available during a response.

Whether controller logic and configuration exist anywhere other than on the devices, since that single fact separates a reconstruction from a restore. How long a controlled restart takes from cold, which somebody in operations can state. What the requalification requirement is and who signs it. Then what proportion of output is recoverable through spare capacity, which sets the shape of the cost curve. Those four convert an outage from an unknown duration into a modeled one.

Model From First Part, Not First Login

A control environment cannot be restored the way an information system can, because the logic and configuration driving a physical process may exist only on devices that have to be wiped and rebuilt individually. Between clean systems and sellable output sit recommissioning, requalification and release, and where qualification applies it frequently dominates the duration while a quality function controls the pace. Material produced during the dwell period may be unsellable, so detection time converts into inventory at risk rather than into records exposed. The largest loss frequently lands on a customer whose line stopped. Kovrr's cyber risk quantification models duration as its own distribution, which is what allows the restart to be modeled separately from the restore.

To see outage exposure modeled from restart and requalification rather than from system recovery time, book a demo with our risk experts.

Tomer Shoolman

Product Manager

Manufacturing Recovery FAQs

Speak to an Expert

Why can't a production line be restored from backup?

What sits between clean systems and sellable output?

Which stage takes longest?

What happens to product made during the incident?

How should the duration curve be built?

Is the lost output recoverable?