Create a free Industrial Equipment News account to continue

The Cloud Recovery Trap

Manufacturing resilience isn't about the backup - it's about how quickly operations are restored.

Cloud Phattharachai Rattanachaiwong
istock.com/phattharachai Rattanachaiwong

Cloud backup quickly became the default answer to resilience. Virtually unlimited storage, off-site protection and reduced reliance on physical media made it an obvious choice for many organizations. For office-based IT environments, that approach has worked well. Recovery times measured in hours are often acceptable, and cloud storage provides a practical, scalable way to protect business data.

OT Environments Operate Differently

In manufacturing, energy, utilities and other operational environments, downtime has immediate consequences. According to one study, manufacturers in the US incur up to $207 million every week due to avoidable downtime. When production stops, resilience isn't measured by whether a backup exists. It's measured by how quickly operations can be restored.

When a critical workstation or control system becomes unavailable, nobody asks where the backup is stored. They ask one question: How quickly can operations resume?

This is where many resilience strategies begin to fall short. Protecting data and restoring operations are two fundamentally different challenges. Cloud remains an excellent destination for protecting IT data, but assuming it is also the best place to recover critical infrastructure introduces operational dependencies that many organizations cannot afford. Increasingly, manufacturers are separating storage decisions from recovery decisions.

What Works for IT Won't Always Work for OT

One of the biggest assumptions to emerge during the cloud era was that if data could be backed up anywhere, it could also be recovered from anywhere just as effectively. In critical infrastructure, the reality is very different.

Many manufacturing sites operate with limited bandwidth, while OT environments are often deliberately segmented or air-gapped for security and operational reasons. Recovering large system images from cloud infrastructure isn't always practical when every minute of downtime matters, particularly when recovery also depends on internet connectivity, authentication services, cloud platform availability and other external infrastructure.

Recovering a spreadsheet and recovering a manufacturing workstation are fundamentally different exercises. Production environments depend on operating systems, applications, drivers, machine configurations, recipes and legacy systems working together exactly as expected. Restoring these environments is about far more than recovering data. It's about returning an entire operational system to service quickly, predictably and with as few dependencies as possible.

Cloud is One Layer, Not the Whole Strategy 

None of this means manufacturers should abandon the cloud. In fact, cloud storage still plays an important role in modern backup strategies and is often a key part of approaches such as the 3-2-1-1-0 methodology, providing off-site protection, long-term retention and an additional layer of resilience.

What's changing is how organizations think about recovery. Increasingly, manufacturers are recognizing that the best place to store backups isn't always the best place to recover critical operational systems. 

Rather than relying on a single recovery strategy, they're using the cloud as a safety net while ensuring their most critical systems can be restored through fast, local recovery capabilities when operational downtime is the priority. 

For many manufacturers, the challenge starts with scale. OT environments often consist of large system images that generate significant volumes of data. Storing every recovery image in the cloud can quickly become expensive, leading organizations to make compromises around retention policies or the systems they choose to protect.

More importantly, storage is only one side of the equation. Recovery is the other.

When downtime can cost hundreds of thousands of dollars per hour, the cost of waiting for critical systems to be restored can quickly outweigh any savings made on cloud storage. In many OT environments, recovering large system images over limited or restricted connectivity simply isn't compatible with the recovery times operations demand. 

Organizations aren't failing because they don't have backups. They're failing because the recovery architecture they're relying on wasn't designed with OT in mind.

The Weakest Link Might Not Be Yours

The CrowdStrike outage proved that resilience isn't just about your own infrastructure. One of the most widely publicized technology disruptions in recent years, it challenged a long-held assumption: that the platforms you depend on will always be there when you need them most.

That thinking extends to recovery. Cloud is another third-party dependency. If restoring critical systems depends on infrastructure outside your control, your recovery strategy is only as resilient as the providers supporting it.

Regulators are catching up. NIS2 and the NIST Cybersecurity Framework both place increasing emphasis on third-party risk - and for manufacturers in scope, that applies directly to how you architect recovery, not just how you protect data. 

Cloud remains an essential part of modern resilience, and manufacturers aren't abandoning it, nor should they. They're becoming more deliberate about the role it plays.

The organizations best positioned to minimize operational disruption will be those that separate backup storage from recovery architecture, using the cloud where it adds value while ensuring their most critical systems can be restored quickly, predictably and with the fewest possible dependencies.

Ultimately, resilience isn't measured by where your backups are stored. It's measured by how quickly you can recover the systems that keep your operations running.

More in Software