Fiber cut, router failure, carrier outage, industrial switch failure, unstable radio link: an industrial network outage can render a SCADA system unavailable, isolate PLCs, interrupt alarm reporting, and slow down on-site diagnostics. To prevent a telecom incident from leading to a production shutdown, it is necessary to design an OT network continuity solution with LTE failover, dual SIM, active monitoring, alerts, and priority rules tailored to critical applications.
The Problem of Industrial Power Outages
In an industrial setting, the network is not merely a convenience. It connects the SCADA system to PLCs, sensors to the monitoring system, operators to control interfaces, maintenance personnel to remote equipment, and logging systems to production data.
When this network goes down, the impact can be immediate.
-
The SCADA system loses contact with the PLCs: no more real-time data, no more remote control, no more centralized visibility.
-
Operators sometimes have to physically access each piece of equipment, which prolongs the process and increases the risk of error.
-
Production logs, measurement histories, and event logs may contain gaps that are difficult to interpret.
-
Critical alarms are no longer being forwarded to the monitoring center or the on-call team.
-
Troubleshooting takes longer because it is necessary to distinguish between a network failure, a PLC failure, an application failure, or an operator error.
-
Remote or multi-site locations are highly dependent on the quality of the primary WAN link.
-
Critical infrastructure—such as water, energy, healthcare, logistics, or certain continuous processes—must demonstrate the ability to maintain continuity and recover.
The main risk is not just the outage itself. It is the lack of an automatic mechanism to detect, switch over, alert, and return to normal operation.
Why a Network Outage Can Halt Production
Not all network outages cause an immediate shutdown. Some systems continue to operate locally thanks to PLCs, control loops, and field safety systems. However, as soon as monitoring, multi-device coordination, or remote intervention becomes necessary, the loss of network connectivity significantly disrupts operations.
| Affected component | Immediate effect | Industrial risk |
|---|---|---|
| SCADA | Loss of visibility and control | Reduced operational capability or precautionary shutdown |
| PLCs | Unreachable from the supervisory system | Slower diagnostics and recovery |
| Alarms | Not transmitted or delayed | Incident detected too late |
| History | Missing data | Weakened quality analysis and traceability |
| Remote access | Maintenance impossible | On-site visit required |
| IP cameras | Loss of visual control | Reduced monitoring |
| Remote meter reading | Data not reported | Disruptions to billing, reporting, or compliance |
An OT network must therefore be designed as an availability infrastructure. Redundancy is not just about “having Internet access,” but about maintaining critical functions during a failure of the primary link.
Common Causes of OT Network Outages
Industrial network outages have various causes. Some are caused by the operator, while others stem from the site itself.
-
Fiber-optic cable breakage due to construction work or a civil engineering incident.
-
Failure of a set-top box, router, firewall, or modem.
-
Power loss in a network rack or field cabinet.
-
Poor quality of an ADSL, SDSL, wireless, or 4G connection.
-
Temporary overload of the main link.
-
Incorrect network configuration after maintenance.
-
Failure of an industrial switch or a fiber converter.
-
Network loop, broadcast storm, or segmentation fault.
-
Unplanned operator maintenance.
-
A cyber incident requiring the isolation of part of the network.
OT network continuity cannot rely on a single scenario. It must monitor the actual availability of the service and switch over when the primary link is no longer functioning properly.
Recommended Architecture with LTE Dual SIM Failover
An industrial LTE failover architecture places a gateway between the OT network and the outgoing links. This gateway monitors the primary link, maintains a mobile backup link, and enforces an automatic failover policy.
This architecture makes it possible to maintain a network path when the primary link becomes unavailable. It also provides teams with evidence of what happened: the time of the outage, the link used, the duration of the failover, the return to normal operations, and the observed quality.
Our Approach
The Eziwan gateway continuously monitors the quality of the primary link, whether it is fiber, ADSL, SDSL, carrier Ethernet, or another WAN link. As soon as a performance degradation threshold is detected, it can automatically switch to a backup LTE link.
The switchover can be triggered by several signals.
-
Complete loss of the main link.
-
Abnormal latency.
-
Excessive packet loss.
-
VPN tunnel unavailable.
-
Jitter is too high for sensitive applications.
-
Repeated failures of the availability probes.
-
Long-term degradation beyond a defined threshold.
Two SIM cards from different carriers enhance redundancy. If the primary mobile network is unavailable or experiencing service degradation, the gateway can use a second SIM card. For the most critical sites or those outside mobile coverage, a satellite link—such as Starlink or VSAT, depending on the context—can be added as a third level of backup.
Key Features
Active/Passive Dual SIM
Active/passive dual SIM involves having a primary SIM and a secondary SIM ready to take over. The two SIM cards can be associated with different carriers to reduce dependence on a single mobile network.
This approach is particularly useful when the site normally relies on a high-performance service provider but must remain accessible in the event of a local outage, network congestion, or maintenance.
Transparent Failover for Critical Services
Failover should be as seamless as possible for industrial applications. Depending on the protocols, VPN configuration, application timeouts, and reconnection mechanisms, some sessions may survive the failover or reestablish automatically.
The goal is to maintain critical functions: alarms, monitoring, remote access, status updates, authorized commands, and priority telemetry. Less urgent data streams can be throttled or suspended while operating on LTE.
Network Quality Probes
A sudden outage is easy to detect. A gradual degradation is more dangerous, because the connection still appears to be active even as applications become unstable.
Quality sensors monitor, among other things:
-
Latency.
-
The jig.
-
Packet loss.
-
The status of the VPN tunnel.
-
The availability of reference targets.
-
LTE radio quality.
-
The active mobile carrier.
These measures allow for a switchover before the monitoring system becomes inoperable.
Real-Time Alerts
Alerts must notify the right people at the right time. A useful notification specifies the affected site, the failed link, the failover link used, the time of the switchover, the tunnel status, and the severity level.
Channels may include:
-
Email.
-
Text message.
-
Webhook.
-
Centralized monitoring.
-
Ticket in an ITSM tool.
-
Notification to an on-call team.
The restoration alert is just as important as the outage alert, because it confirms a return to normal conditions and allows for an analysis of the actual duration of the incident.
Availability Report
An availability report provides a factual overview of network incidents. It can be used by OT teams, the IT department, the maintenance manager, the quality manager, or the service provider.
A good report should include:
-
The number of rollovers.
-
The cumulative time spent on the backup link.
-
The links used.
-
Outage and restoration times.
-
The causes identified.
-
The observed radio quality.
-
The periods of deterioration prior to disconnection.
-
Recurring events by site.
This data helps determine whether to switch providers, relocate an antenna, add a satellite link, replace a router, or redesign the network topology.
Compatibility with Existing IP Equipment
A network continuity gateway must integrate without requiring a complete overhaul of industrial equipment. It can work with a wide range of existing IP devices: SCADA systems, Siemens, Schneider, or Rockwell PLCs, IP cameras, HMIs, data logging servers, industrial switches, connected sensors, RTUs, and IoT gateways.
The challenge is to maintain essential data flows while avoiding direct exposure of OT equipment to the Internet.
Configurable Graceful Degradation
When a site operates on LTE, it may be necessary to reduce certain uses to manage bandwidth and data costs. Graceful degradation involves maintaining critical functions and deprioritizing the rest.
Examples of priorities:
-
Maintain critical alarms.
-
Maintain the controls necessary for operations.
-
Keep the VPN tunnel open for on-call duty.
-
Reduce the frequency of non-urgent polling.
-
Pause certain video streams.
-
Postpone large exports.
-
Limit non-critical updates.
This approach transforms the LTE connection into true service continuity, rather than a simple "best-effort" fallback.
Third-Level Satellite Support
For very remote locations, 4G or 5G isn’t always enough. An Ethernet WAN port can allow for the addition of a satellite modem as a third-level backup. This triple-redundant architecture is suitable for sites where a loss of connectivity has a significant impact: water, power, security, critical remote meter reading, monitoring of remote sites, or multi-site operations.
However, the satellite must be designed as a standalone link: line-of-sight, power supply, mounting, latency, monitoring, cost, and priority rules must all be validated.
The Life Cycle of an Industrial Partnership
An effective failover follows a clear cycle: monitoring, detection, validation, failover, monitoring of the failover mode, and controlled return.
The return to the main link must be controlled. If the signal strength recovers for a few seconds and then drops again, a return that is too rapid can cause oscillations. It is therefore necessary to wait for a period of stability before returning to the nominal level.
Which Data Streams to Prioritize During an Outage
Not all OT flows have the same level of criticality. A good continuity plan defines what must be maintained during an outage of the primary link.
| Flow | Priority | Comment |
|---|---|---|
| Critical Alarms | Very High | Must be maintained even in degraded mode |
| Operational Commands | High | Must be limited to authorized uses |
| SCADA Monitoring | High | Polling adjustable based on bandwidth |
| Remote Maintenance Access | Medium to High | As per on-call procedures |
| Detailed Logging | Medium | May be temporarily reduced |
| IP Video | Variable | Often bandwidth-intensive |
| Batch Exports | Low | To be postponed until service is restored |
| Updates | Low | To be blocked during failover |
This prioritization prevents a secondary data stream from consuming the LTE connection at the expense of alarms or essential monitoring.
Example of a Failover Policy
A failover policy must be documented and understandable to IT, OT, and maintenance teams. It describes thresholds, failover paths, alerts, and recovery rules.
site:
nom: usine_ligne_conditionnement
criticite: haute
lien_principal:
type: fibre
surveillance:
latence_max: 250ms
perte_paquets_max: 5%
echec_sonde: 3
duree_degradation: 30s
secours:
lte:
mode: dual_sim
sim_1: operateur_a
sim_2: operateur_b
priorite:
- alarmes
- supervision
- acces_maintenance
satellite:
actif: optionnel
usage: dernier_recours
retour_nominal:
stabilite_requise: 10min
retour_automatique: true
alertes:
basculement: sms_email
degradation: email
retour_service: sms_email
rapport_mensuel: pdf
The settings must be tailored to the site. A standalone pumping station, an automated pipeline, and a multi-site monitoring center do not have the same constraints.
Measuring the OT Network MTTR
MTTR, or mean time to recovery, is a key metric. In an industrial network, it is important to measure not only the time it takes to repair the main link, but also the time required to maintain or restore service.
We can distinguish several stages.
| Indicator | Definition | Objective |
|---|---|---|
| Detection Time | Time between incident and detection | As short as possible |
| Failover Time | Time before backup is activated | Application-compatible |
| Notification Time | Time before team is alerted | Immediate or near-immediate |
| Diagnosis time | Time to identify the cause | Reduced by logs |
| Repair time | Time to restore the primary link | Depends on the provider |
| Nominal recovery time | Time before controlled recovery | Avoid oscillations |
LTE failover does not repair a broken fiber connection. It minimizes the operational impact while repairs are underway.
SLA, Availability, and Production Continuity
An availability target, such as 99.9% or 99.97%, should not be used as a generic promise without context. It depends on the architecture, available connections, cellular coverage, power supply, maintenance, carrier contracts, and monitoring.
On the other hand, a redundant architecture makes it possible to build a more robust service target than a single link.
| Architecture | Expected Resilience | Main Limitation |
|---|---|---|
| Single WAN connection | Low | Any outage takes the site offline |
| WAN + single-SIM LTE | Medium | Dependence on a single mobile carrier |
| WAN + Dual SIM | High | Dependence on mobile coverage |
| WAN + Dual SIM + Satellite | Very high | Cost, latency, and complexity |
| Managed multi-sites | High if properly operated | Requires clear procedures |
Actual availability must be measured over time using actionable reports, not just stated in a contract.
Security: Ensuring Business Continuity Without Exposing the OT
A poorly designed backup connection can pose a cybersecurity risk. For example, a 4G modem connected directly to a PLC, with an exposed administration interface, can create an uncontrolled entry point.
Good security practices are essential.
-
No OT services are directly exposed on the Internet.
-
Encrypted VPN tunnel to a controlled platform.
-
Personal accounts for remote access.
-
Strong authentication when human access is permitted.
-
Segmentation between the IT network, DMZ, and OT.
-
Filtering traffic by destination and usage.
-
Logging of connections and failovers.
-
Quick revocation of service provider access.
-
Monitoring of configuration and events.
This approach is consistent with the defense-in-depth principles used in industrial cybersecurity, particularly in approaches based on IEC 62443.
LTE Failover and IEC 62443
IEC 62443 is not limited to network availability. It addresses the cybersecurity of industrial systems more broadly, covering zones, conduits, security requirements, risk management, hardening, and governance. An LTE failover can contribute to this approach, but it does not, on its own, make an infrastructure compliant.
He can help with several technical issues.
-
Separation of zones and network conduits.
-
Remote access management.
-
Logging of connectivity events.
-
Reduced direct exposure of OT equipment.
-
Maintenance of certain critical functions in the event of a network outage.
-
Documentation of workflows and dependencies.
Compliance must always be assessed at the level of the entire system: architecture, procedures, users, maintenance, suppliers, and operational evidence.
Industrial Use Cases
Plant with Centralized SCADA
A plant that monitors multiple production lines from a centralized SCADA system is heavily dependent on network availability. If the primary link goes down, LTE failover maintains visibility into alarms and critical statuses until the primary link is restored.
Pumping Station or Water Facility
Drinking water and wastewater treatment facilities are often spread out over a large area. Network continuity ensures that alarms, water levels, drive malfunctions, pump statuses, and other information necessary for on-call personnel are consistently transmitted.
Energy Production
A solar power plant, a delivery station, or a storage facility must transmit its production data, inverter faults, alarms, and availability status. A backup link reduces the risk of losing monitoring capabilities during an operator outage.
Logistics Facility or Automated Warehouse
In an automated warehouse, the network connects the WMS, conveyors, PLCs, sensors, monitoring systems, and sometimes cameras. An outage can slow down or halt physical material flows. Network redundancy helps maintain priority functions.
Special Machine at a Customer’s Site
A machine manufacturer can equip its facilities with a gateway that includes a backup connection to maintain secure maintenance access. This reduces the need for on-site visits and speeds up troubleshooting when the customer’s network is unavailable or unstable.
Pre-Deployment Checklist
Before deploying an industrial LTE failover system, it is necessary to validate the key points.
| Check | Question | Priority |
|---|---|---|
| Mapping | Have critical equipment and data flows been identified? | High |
| Mobile Coverage | Have two carriers been tested on-site? | High |
| Antennas | Does the location ensure sufficient radio quality? | High |
| Power | Do the gateway and network equipment have backup power? | High |
| Thresholds | Are failover criteria defined? | High |
| Prioritization | Are critical data streams prioritized on LTE? | High |
| Security | Are OT services not directly exposed? | High |
| Logs | Are failover events logged? | High |
| Alerts | Are the appropriate teams notified? | High |
| Failback | Is failback to the primary link controlled? | Medium |
| Tests | Has an outage drill been conducted? | High |
The failure test is essential. An untested backup architecture remains merely a hypothetical scenario.
Common Mistakes to Avoid
Add a 4G Modem Without Supervision
A standalone 4G modem can be a temporary solution, but it doesn’t always provide visibility into handoffs, radio quality, data usage, or recurring issues.
Use the Same Operator for All Links
If the primary link and the mobile link rely on the same carrier or the same local infrastructure, redundancy may be less effective. True independence should be sought when the site is critical.
Do Not Prioritize Flows
In the event of an LTE failover, certain data streams can saturate the connection: video, backups, exports, updates, or large-scale synchronizations. Without prioritization, alarms and monitoring may be affected.
Forgetting the Power Supply
A network failover is useless if the active gateway, switch, or antenna no longer has power. Business continuity equipment must be connected to an emergency power supply when the criticality of the system requires it.
Return to the Main Link Too Soon
An unstable primary link can cause the system to switch back and forth between WAN and LTE. A stability period must be defined before the system returns to normal operation.
How Eziwan Helps Reduce the Risk of Shutdown
Eziwan centralizes OT network continuity around a monitored gateway capable of monitoring links, switching to LTE, using two SIM cards, notifying teams, and generating reports.
| Need | Eziwan’s Solution | Benefit |
|---|---|---|
| Fiber or xDSL outage | Automatic failover to LTE | Critical functions maintained |
| Mobile carrier outage | Dual SIM with separate carriers | Enhanced redundancy |
| Difficult diagnostics | Logs and event history | Faster analysis |
| Reduced monitoring | Real-time alerts | Immediate response |
| Remote locations | Adapted antennas and optional satellite | Extended coverage |
| Data costs | Configurable grace period for degraded service | Prioritization of essential data streams |
| Reliability audit | Availability report | Proof of operation |
This continuity can be combined with the Eziwan gateway, monitoring via the Eziwan cloud, and industrial connectivity architectures.
Recommended Test Plan
Once the solution has been installed, its operation must be verified through a controlled test.
The test must verify alarm escalation, SCADA behavior, PLC availability, notifications, logs, and the switchback to the primary link. It must be conducted within a controlled environment, with the relevant teams.
Conclusion
An industrial network outage can lead to a production shutdown when SCADA systems, PLCs, alarms, and remote access rely on a single connection. Dual-SIM LTE failover reduces this risk by adding an automatic, monitored, and operational backup path.
The value of a solution like Eziwan goes beyond the 4G modem. It comes from the whole package: network quality monitoring, automatic failover, Dual SIM, real-time alerts, availability reports, traffic prioritization, and integration with a secure OT architecture. For critical industrial sites, this approach transforms network continuity into a controlled process rather than an emergency response.
Further Reading
- Industrial 4G Backup — automatically activate a 4G backup connection in the event of a fiber outage
- Industrial Internet Redundancy — redundancy architectures to ensure production continuity
- Industrial Connectivity — compare connectivity technologies for your industrial sites
- Industrial 4G Router — industrial routers with automatic failover and built-in Dual SIM
- Industrial LTE Router — high-performance LTE routers for critical industrial environments