5G and LTE Branch Failover WAN Design - 夜莺博客

5G and LTE Branch Failover WAN Design

Branch connectivity is increasingly a wired primary with a cellular backup, chosen because cellular does not share the physical path with the fibre or copper in the ground. The engineering problem is not whether the cellular link works, it is how the network decides to use it, how fast it switches, and whether the site can still be managed when the primary fails.

Three deployment shapes

Shape How it works Trade-off
Overlay adapter Cellular converted to Ethernet, presented as a second WAN to the existing router Least disruptive; the existing router makes the failover decision
Integrated cellular router One device with wired and cellular WANs plus LAN switching Fewer devices; a single point of failure unless paired
Router redundancy with cellular A cellular router stands by as a hardware backup, with VRRP between the two Survives primary router failure as well as WAN failure

The third option addresses the failure most teams forget. WAN failover protects against a carrier outage; it does nothing when the primary router itself stops forwarding. Running VRRP between the primary router and a cellular-capable standby gives simultaneous WAN and hardware failover, with the LAN continuing on the same gateway address.

What should trigger failover

Link state alone is a poor trigger, because a WAN interface can be up while the path beyond it is dead. Before committing to a design, answer these questions:

  • What is probed? A synthetic target beyond the local ISP (not the next hop) is the usual choice. The target must be reachable on both paths, or the test itself causes a failover loop.
  • How many probes must fail? A single dropped probe should never fail the path; use consecutive failures with a short interval, and require a longer stable period before failing back.
  • How are metrics combined? SD-WAN products usually allow a weighted policy combining packet loss, latency and jitter thresholds rather than a simple up/down test. Explicit thresholds are what stop a congested but working link from flapping.
  • What is the failback policy? Automatic failback after a short outage causes repeated disruption. Prefer a hold-down time measured in minutes, or a manual decision for sites where the primary fluctuates.
! Conceptual failover policy
probe: 8.8.8.8 and 1.1.1.1, ICMP, interval 3 s
fail after: 3 consecutive losses on both targets
loss threshold: 10%
latency threshold: 150 ms
jitter threshold: 30 ms
failback hold-down: 10 min

Dual SIM and carrier diversity

Two SIMs from the same carrier on the same cell tower are not redundancy. Dual-SIM designs are only useful when they provide genuinely different carrier coverage, and the failover logic should support switching on coverage loss and on data-plan exhaustion, not just on link failure. Dual-modem devices go further, offering active/active cellular or cellular-to-cellular failover, which matters for sites where the cellular link is the primary and the wired connection is the backup.

Data plans, performance and out-of-band management

Two operational realities constrain every cellular design. The first is the data plan: a failover link that suddenly carries production traffic can exhaust a monthly allowance in hours, so metering, alerting and traffic selection (which applications are allowed to use the backup) belong in the design. The second is throughput: cellular performance varies with signal quality, load and location, so do not design a site's capacity around the backup path - design degradation, not equivalence.

Out-of-band management is the third benefit and often the compelling one. A cellular adapter connected to the primary router's console and management port provides a path that survives the wired outage entirely and removes the need for a POTS line or a truck roll. The design principles for that layer are covered in our Zabbix proxy distributed monitoring guide; the SD-WAN control plane itself is described in Cisco SD-WAN architecture, and monitoring the resulting paths is covered by blackbox_exporter probes.

原文链接:https://cradlepoint.ericsson.com/blog/how-lte-enables-failover-oobm-for-branch-continuity/