Technical Guide

Industrial Power Outage: Avoiding Production Downtime

Prevent production downtime caused by OT network outages with LTE failover, Dual SIM, monitoring, alerts, and industrial redundancy.

Fiber cut, router failure, carrier outage, industrial switch failure, unstable radio link: an industrial network outage can render a SCADA system unavailable, isolate PLCs, interrupt alarm reporting, and slow down on-site diagnostics. To prevent a telecom incident from leading to a production shutdown, it is necessary to design an OT network continuity solution with LTE failover, dual SIM, active monitoring, alerts, and priority rules tailored to critical applications.

The Problem of Industrial Power Outages

In an industrial setting, the network is not merely a convenience. It connects the SCADA system to PLCs, sensors to the monitoring system, operators to control interfaces, maintenance personnel to remote equipment, and logging systems to production data.

When this network goes down, the impact can be immediate.

  • The SCADA system loses contact with the PLCs: no more real-time data, no more remote control, no more centralized visibility.

  • Operators sometimes have to physically access each piece of equipment, which prolongs the process and increases the risk of error.

  • Production logs, measurement histories, and event logs may contain gaps that are difficult to interpret.

  • Critical alarms are no longer being forwarded to the monitoring center or the on-call team.

  • Troubleshooting takes longer because it is necessary to distinguish between a network failure, a PLC failure, an application failure, or an operator error.

  • Remote or multi-site locations are highly dependent on the quality of the primary WAN link.

  • Critical infrastructure—such as water, energy, healthcare, logistics, or certain continuous processes—must demonstrate the ability to maintain continuity and recover.

The main risk is not just the outage itself. It is the lack of an automatic mechanism to detect, switch over, alert, and return to normal operation.

Why a Network Outage Can Halt Production

Not all network outages cause an immediate shutdown. Some systems continue to operate locally thanks to PLCs, control loops, and field safety systems. However, as soon as monitoring, multi-device coordination, or remote intervention becomes necessary, the loss of network connectivity significantly disrupts operations.

Affected componentImmediate effectIndustrial risk
SCADALoss of visibility and controlReduced operational capability or precautionary shutdown
PLCsUnreachable from the supervisory systemSlower diagnostics and recovery
AlarmsNot transmitted or delayedIncident detected too late
HistoryMissing dataWeakened quality analysis and traceability
Remote accessMaintenance impossibleOn-site visit required
IP camerasLoss of visual controlReduced monitoring
Remote meter readingData not reportedDisruptions to billing, reporting, or compliance

An OT network must therefore be designed as an availability infrastructure. Redundancy is not just about “having Internet access,” but about maintaining critical functions during a failure of the primary link.

Common Causes of OT Network Outages

Industrial network outages have various causes. Some are caused by the operator, while others stem from the site itself.

  • Fiber-optic cable breakage due to construction work or a civil engineering incident.

  • Failure of a set-top box, router, firewall, or modem.

  • Power loss in a network rack or field cabinet.

  • Poor quality of an ADSL, SDSL, wireless, or 4G connection.

  • Temporary overload of the main link.

  • Incorrect network configuration after maintenance.

  • Failure of an industrial switch or a fiber converter.

  • Network loop, broadcast storm, or segmentation fault.

  • Unplanned operator maintenance.

  • A cyber incident requiring the isolation of part of the network.

OT network continuity cannot rely on a single scenario. It must monitor the actual availability of the service and switch over when the primary link is no longer functioning properly.

An industrial LTE failover architecture places a gateway between the OT network and the outgoing links. This gateway monitors the primary link, maintains a mobile backup link, and enforces an automatic failover policy.

This architecture makes it possible to maintain a network path when the primary link becomes unavailable. It also provides teams with evidence of what happened: the time of the outage, the link used, the duration of the failover, the return to normal operations, and the observed quality.

Our Approach

The Eziwan gateway continuously monitors the quality of the primary link, whether it is fiber, ADSL, SDSL, carrier Ethernet, or another WAN link. As soon as a performance degradation threshold is detected, it can automatically switch to a backup LTE link.

The switchover can be triggered by several signals.

  • Complete loss of the main link.

  • Abnormal latency.

  • Excessive packet loss.

  • VPN tunnel unavailable.

  • Jitter is too high for sensitive applications.

  • Repeated failures of the availability probes.

  • Long-term degradation beyond a defined threshold.

Two SIM cards from different carriers enhance redundancy. If the primary mobile network is unavailable or experiencing service degradation, the gateway can use a second SIM card. For the most critical sites or those outside mobile coverage, a satellite link—such as Starlink or VSAT, depending on the context—can be added as a third level of backup.

Key Features

Active/Passive Dual SIM

Active/passive dual SIM involves having a primary SIM and a secondary SIM ready to take over. The two SIM cards can be associated with different carriers to reduce dependence on a single mobile network.

This approach is particularly useful when the site normally relies on a high-performance service provider but must remain accessible in the event of a local outage, network congestion, or maintenance.

Transparent Failover for Critical Services

Failover should be as seamless as possible for industrial applications. Depending on the protocols, VPN configuration, application timeouts, and reconnection mechanisms, some sessions may survive the failover or reestablish automatically.

The goal is to maintain critical functions: alarms, monitoring, remote access, status updates, authorized commands, and priority telemetry. Less urgent data streams can be throttled or suspended while operating on LTE.

Network Quality Probes

A sudden outage is easy to detect. A gradual degradation is more dangerous, because the connection still appears to be active even as applications become unstable.

Quality sensors monitor, among other things:

  • Latency.

  • The jig.

  • Packet loss.

  • The status of the VPN tunnel.

  • The availability of reference targets.

  • LTE radio quality.

  • The active mobile carrier.

These measures allow for a switchover before the monitoring system becomes inoperable.

Real-Time Alerts

Alerts must notify the right people at the right time. A useful notification specifies the affected site, the failed link, the failover link used, the time of the switchover, the tunnel status, and the severity level.

Channels may include:

  • Email.

  • Text message.

  • Webhook.

  • Centralized monitoring.

  • Ticket in an ITSM tool.

  • Notification to an on-call team.

The restoration alert is just as important as the outage alert, because it confirms a return to normal conditions and allows for an analysis of the actual duration of the incident.

Availability Report

An availability report provides a factual overview of network incidents. It can be used by OT teams, the IT department, the maintenance manager, the quality manager, or the service provider.

A good report should include:

  • The number of rollovers.

  • The cumulative time spent on the backup link.

  • The links used.

  • Outage and restoration times.

  • The causes identified.

  • The observed radio quality.

  • The periods of deterioration prior to disconnection.

  • Recurring events by site.

This data helps determine whether to switch providers, relocate an antenna, add a satellite link, replace a router, or redesign the network topology.

Compatibility with Existing IP Equipment

A network continuity gateway must integrate without requiring a complete overhaul of industrial equipment. It can work with a wide range of existing IP devices: SCADA systems, Siemens, Schneider, or Rockwell PLCs, IP cameras, HMIs, data logging servers, industrial switches, connected sensors, RTUs, and IoT gateways.

The challenge is to maintain essential data flows while avoiding direct exposure of OT equipment to the Internet.

Configurable Graceful Degradation

When a site operates on LTE, it may be necessary to reduce certain uses to manage bandwidth and data costs. Graceful degradation involves maintaining critical functions and deprioritizing the rest.

Examples of priorities:

  • Maintain critical alarms.

  • Maintain the controls necessary for operations.

  • Keep the VPN tunnel open for on-call duty.

  • Reduce the frequency of non-urgent polling.

  • Pause certain video streams.

  • Postpone large exports.

  • Limit non-critical updates.

This approach transforms the LTE connection into true service continuity, rather than a simple "best-effort" fallback.

Third-Level Satellite Support

For very remote locations, 4G or 5G isn’t always enough. An Ethernet WAN port can allow for the addition of a satellite modem as a third-level backup. This triple-redundant architecture is suitable for sites where a loss of connectivity has a significant impact: water, power, security, critical remote meter reading, monitoring of remote sites, or multi-site operations.

However, the satellite must be designed as a standalone link: line-of-sight, power supply, mounting, latency, monitoring, cost, and priority rules must all be validated.

The Life Cycle of an Industrial Partnership

An effective failover follows a clear cycle: monitoring, detection, validation, failover, monitoring of the failover mode, and controlled return.

The return to the main link must be controlled. If the signal strength recovers for a few seconds and then drops again, a return that is too rapid can cause oscillations. It is therefore necessary to wait for a period of stability before returning to the nominal level.

Which Data Streams to Prioritize During an Outage

Not all OT flows have the same level of criticality. A good continuity plan defines what must be maintained during an outage of the primary link.

FlowPriorityComment
Critical AlarmsVery HighMust be maintained even in degraded mode
Operational CommandsHighMust be limited to authorized uses
SCADA MonitoringHighPolling adjustable based on bandwidth
Remote Maintenance AccessMedium to HighAs per on-call procedures
Detailed LoggingMediumMay be temporarily reduced
IP VideoVariableOften bandwidth-intensive
Batch ExportsLowTo be postponed until service is restored
UpdatesLowTo be blocked during failover

This prioritization prevents a secondary data stream from consuming the LTE connection at the expense of alarms or essential monitoring.

Example of a Failover Policy

A failover policy must be documented and understandable to IT, OT, and maintenance teams. It describes thresholds, failover paths, alerts, and recovery rules.

site:
nom: usine_ligne_conditionnement
criticite: haute

lien_principal:
type: fibre
surveillance:
latence_max: 250ms
perte_paquets_max: 5%
echec_sonde: 3
duree_degradation: 30s

secours:
lte:
mode: dual_sim
sim_1: operateur_a
sim_2: operateur_b
priorite:
- alarmes
- supervision
- acces_maintenance
satellite:
actif: optionnel
usage: dernier_recours

retour_nominal:
stabilite_requise: 10min
retour_automatique: true

alertes:
basculement: sms_email
degradation: email
retour_service: sms_email
rapport_mensuel: pdf

The settings must be tailored to the site. A standalone pumping station, an automated pipeline, and a multi-site monitoring center do not have the same constraints.

Measuring the OT Network MTTR

MTTR, or mean time to recovery, is a key metric. In an industrial network, it is important to measure not only the time it takes to repair the main link, but also the time required to maintain or restore service.

We can distinguish several stages.

IndicatorDefinitionObjective
Detection TimeTime between incident and detectionAs short as possible
Failover TimeTime before backup is activatedApplication-compatible
Notification TimeTime before team is alertedImmediate or near-immediate
Diagnosis timeTime to identify the causeReduced by logs
Repair timeTime to restore the primary linkDepends on the provider
Nominal recovery timeTime before controlled recoveryAvoid oscillations

LTE failover does not repair a broken fiber connection. It minimizes the operational impact while repairs are underway.

SLA, Availability, and Production Continuity

An availability target, such as 99.9% or 99.97%, should not be used as a generic promise without context. It depends on the architecture, available connections, cellular coverage, power supply, maintenance, carrier contracts, and monitoring.

On the other hand, a redundant architecture makes it possible to build a more robust service target than a single link.

ArchitectureExpected ResilienceMain Limitation
Single WAN connectionLowAny outage takes the site offline
WAN + single-SIM LTEMediumDependence on a single mobile carrier
WAN + Dual SIMHighDependence on mobile coverage
WAN + Dual SIM + SatelliteVery highCost, latency, and complexity
Managed multi-sitesHigh if properly operatedRequires clear procedures

Actual availability must be measured over time using actionable reports, not just stated in a contract.

Security: Ensuring Business Continuity Without Exposing the OT

A poorly designed backup connection can pose a cybersecurity risk. For example, a 4G modem connected directly to a PLC, with an exposed administration interface, can create an uncontrolled entry point.

Good security practices are essential.

  • No OT services are directly exposed on the Internet.

  • Encrypted VPN tunnel to a controlled platform.

  • Personal accounts for remote access.

  • Strong authentication when human access is permitted.

  • Segmentation between the IT network, DMZ, and OT.

  • Filtering traffic by destination and usage.

  • Logging of connections and failovers.

  • Quick revocation of service provider access.

  • Monitoring of configuration and events.

This approach is consistent with the defense-in-depth principles used in industrial cybersecurity, particularly in approaches based on IEC 62443.

LTE Failover and IEC 62443

IEC 62443 is not limited to network availability. It addresses the cybersecurity of industrial systems more broadly, covering zones, conduits, security requirements, risk management, hardening, and governance. An LTE failover can contribute to this approach, but it does not, on its own, make an infrastructure compliant.

He can help with several technical issues.

  • Separation of zones and network conduits.

  • Remote access management.

  • Logging of connectivity events.

  • Reduced direct exposure of OT equipment.

  • Maintenance of certain critical functions in the event of a network outage.

  • Documentation of workflows and dependencies.

Compliance must always be assessed at the level of the entire system: architecture, procedures, users, maintenance, suppliers, and operational evidence.

Industrial Use Cases

Plant with Centralized SCADA

A plant that monitors multiple production lines from a centralized SCADA system is heavily dependent on network availability. If the primary link goes down, LTE failover maintains visibility into alarms and critical statuses until the primary link is restored.

Pumping Station or Water Facility

Drinking water and wastewater treatment facilities are often spread out over a large area. Network continuity ensures that alarms, water levels, drive malfunctions, pump statuses, and other information necessary for on-call personnel are consistently transmitted.

Energy Production

A solar power plant, a delivery station, or a storage facility must transmit its production data, inverter faults, alarms, and availability status. A backup link reduces the risk of losing monitoring capabilities during an operator outage.

Logistics Facility or Automated Warehouse

In an automated warehouse, the network connects the WMS, conveyors, PLCs, sensors, monitoring systems, and sometimes cameras. An outage can slow down or halt physical material flows. Network redundancy helps maintain priority functions.

Special Machine at a Customer’s Site

A machine manufacturer can equip its facilities with a gateway that includes a backup connection to maintain secure maintenance access. This reduces the need for on-site visits and speeds up troubleshooting when the customer’s network is unavailable or unstable.

Pre-Deployment Checklist

Before deploying an industrial LTE failover system, it is necessary to validate the key points.

CheckQuestionPriority
MappingHave critical equipment and data flows been identified?High
Mobile CoverageHave two carriers been tested on-site?High
AntennasDoes the location ensure sufficient radio quality?High
PowerDo the gateway and network equipment have backup power?High
ThresholdsAre failover criteria defined?High
PrioritizationAre critical data streams prioritized on LTE?High
SecurityAre OT services not directly exposed?High
LogsAre failover events logged?High
AlertsAre the appropriate teams notified?High
FailbackIs failback to the primary link controlled?Medium
TestsHas an outage drill been conducted?High

The failure test is essential. An untested backup architecture remains merely a hypothetical scenario.

Common Mistakes to Avoid

Add a 4G Modem Without Supervision

A standalone 4G modem can be a temporary solution, but it doesn’t always provide visibility into handoffs, radio quality, data usage, or recurring issues.

If the primary link and the mobile link rely on the same carrier or the same local infrastructure, redundancy may be less effective. True independence should be sought when the site is critical.

Do Not Prioritize Flows

In the event of an LTE failover, certain data streams can saturate the connection: video, backups, exports, updates, or large-scale synchronizations. Without prioritization, alarms and monitoring may be affected.

Forgetting the Power Supply

A network failover is useless if the active gateway, switch, or antenna no longer has power. Business continuity equipment must be connected to an emergency power supply when the criticality of the system requires it.

An unstable primary link can cause the system to switch back and forth between WAN and LTE. A stability period must be defined before the system returns to normal operation.

How Eziwan Helps Reduce the Risk of Shutdown

Eziwan centralizes OT network continuity around a monitored gateway capable of monitoring links, switching to LTE, using two SIM cards, notifying teams, and generating reports.

NeedEziwan’s SolutionBenefit
Fiber or xDSL outageAutomatic failover to LTECritical functions maintained
Mobile carrier outageDual SIM with separate carriersEnhanced redundancy
Difficult diagnosticsLogs and event historyFaster analysis
Reduced monitoringReal-time alertsImmediate response
Remote locationsAdapted antennas and optional satelliteExtended coverage
Data costsConfigurable grace period for degraded servicePrioritization of essential data streams
Reliability auditAvailability reportProof of operation

This continuity can be combined with the Eziwan gateway, monitoring via the Eziwan cloud, and industrial connectivity architectures.

Once the solution has been installed, its operation must be verified through a controlled test.

The test must verify alarm escalation, SCADA behavior, PLC availability, notifications, logs, and the switchback to the primary link. It must be conducted within a controlled environment, with the relevant teams.

Conclusion

An industrial network outage can lead to a production shutdown when SCADA systems, PLCs, alarms, and remote access rely on a single connection. Dual-SIM LTE failover reduces this risk by adding an automatic, monitored, and operational backup path.

The value of a solution like Eziwan goes beyond the 4G modem. It comes from the whole package: network quality monitoring, automatic failover, Dual SIM, real-time alerts, availability reports, traffic prioritization, and integration with a secure OT architecture. For critical industrial sites, this approach transforms network continuity into a controlled process rather than an emergency response.

Further Reading

Frequently Asked Questions

You might also like