FortiGate HA Failover Troubleshooting (FGCP) - 夜莺博客

FortiGate HA Failover Troubleshooting (FGCP)

Fortinet's FGCP clustering is easy to configure and hard to debug, because a broken cluster usually looks healthy from the web UI. Two members can both believe they are primary, configurations can drift apart without an error, and a failover can happen without anyone noticing until traffic drops. The fix is to stop reading dashboards and start reading the cluster's own diagnostics.

What a healthy cluster looks like

One primary and one or more secondaries, with session and configuration synchronisation running over dedicated heartbeat links. Priority influences the election, and with override enabled a higher-priority unit takes the primary role back when it returns. Monitored interfaces decide whether a member is eligible: if a monitored link goes down, the unit stops being a failover candidate - or in some designs, triggers a failover.

The commands that matter

# get system status                 # HA mode, current role, cluster uptime
# diagnose sys ha status            # detailed per-member HA information
# execute ha md5sum                 # SYS and CLI checksums for both members
# diagnose debug application hatalk 7
# diagnose debug enable             # heartbeat/failover messages as they happen
# diagnose sys ha checksum show     # compare configuration sections

execute ha md5sum output lists each member with a SYS and a CLI hash. Identical hashes mean the configurations agree. Different hashes mean one member is running a configuration the other has never received - a silent, security-relevant divergence that a "synchronised" status line will not reveal.

Symptom 1: both members are primary (split-brain)

The cluster has lost heartbeats and each unit independently decided it is responsible. In practice this is caused by the heartbeat path, not by software:

  • Heartbeat traffic sharing a path with data, and that path being saturated or filtered mid-way.
  • Heartbeat interfaces on a switch whose spanning-tree or port security state changed during a maintenance window.
  • A firmware mismatch between members, which prevents synchronisation and complicates the election.
  • Heartbeat links on different VLANs, or one side behind a NAT/filtering device.

Re-establish the heartbeat link, then resynchronise rather than "fixing" roles by hand. Both members claiming primary is also the point at which you should check for duplicate IPs on the network and for MAC address conflicts in the neighbours' ARP tables.

Symptom 2: unexpected failover

# diagnose sys ha status
# diagnose debug application hatalk 7

The heartbeat debug shows the last heartbeat time and the reason the cluster decided a member was gone. Common causes: a monitored interface flapped (link debounce too short or a bad optic), CPU or memory crossed a threshold on the secondary, or the heartbeat link itself suffered micro-outages. Raise the heartbeat loss threshold (hb-lost-threshold) if the link is genuinely noisy - but investigate the link first; a threshold change hides a real fault.

Symptom 3: traffic drops after failover

  • Session pickup disabled: without session synchronisation, existing flows are lost on switchover. Enable session pickup for stateful inspection to survive, and accept that long-lived flows may still need to re-establish.
  • Asymmetric routing: return traffic reaching the other member without session synchronisation is silently dropped. Check the upstream routing, not the cluster.
  • Configuration divergence: verify checksums as above. A member that never received a policy update behaves inconsistently after it becomes primary.
  • ARP and neighbour caching: upstream devices may still have the old member's MAC. Forcing a gratuitous ARP or shortening ARP timers on the neighbours helps; VLAN-level failover is cleaner than IP-level when the topology allows.

Operating an FGCP cluster

  • Monitor the cluster role externally (two alerting sources, one per member) so a split-brain is visible immediately.
  • Alert on configuration checksum mismatch, not just on "HA status".
  • Keep firmware identical across members before and after upgrades, and verify synchronisation at the end of every change window.
  • Document which interfaces are monitored; unmonitored uplinks are the most common single point of failure in a "redundant" cluster.
  • Test failover deliberately - pull a heartbeat cable during a window - so the first real failover is not also the first test.

Related: FortiGate debug flow and packet trace, HSRP vs VRRP vs GLBP first-hop redundancy, and PAN-OS commit failure troubleshooting.

原文链接:https://docs.fortinet.com/document/fortiweb/8.0.0/troubleshooting-guide/182034/ha-trouble-shooting