OPNsense CARP and pfsync: Firewall HA That Fails Over - 夜莺博客

OPNsense CARP and pfsync: Firewall HA That Fails Over

A firewall pair without state synchronisation is not highly available, it is merely redundant: the backup takes over and every established session dies. OPNsense combines three mechanisms to avoid that - CARP virtual IPs for address ownership, pfsync for state table replication and XMLRPC for configuration replication. This guide covers the working configuration order and the checks that tell you failover will actually be seamless.

Design prerequisites

  • Two nodes built from the same OPNsense version - mismatched versions cause state sync problems.
  • At least two networks: one transit/carp network between the firewalls, and the LAN/WAN segments they protect. A dedicated interface for pfsync is strongly recommended, for both security (state injection) and performance.
  • Every protected segment gets a CARP virtual IP, shared across the pair, plus a unique physical IP per node - both sets are required.
  • Both nodes in the same layer 2 domain for CARP; multicast is the default transport, and switching to unicast in the Virtual IP settings is the standard workaround when multicast is filtered. OPNsense VLAN Interfaces and Firewall Rule Processing covers the interface and rule semantics you will rely on afterwards.

1. Virtual IPs and CARP parameters

Interfaces > Virtual IPs > Add
  Type:        CARP
  Interface:   LAN
  Address:     10.20.0.1/24        (shared virtual address)
  VHID Group:  1
  Advertising Frequency: 1 (base 1, skew 0 on master)
  Description: LAN virtual IP

# each node keeps its own physical IP, e.g. master 10.20.0.11, backup 10.20.0.12

Every CARP group needs a unique VHID within the broadcast domain. A duplicate VHID produces two masters advertising over each other - the classic intermittent outage that clears by itself.

2. State synchronisation (pfsync) on the master

System > High Availability > Settings
  Synchronize States:        enabled
  Synchronize Interface:     PFSYNC (the dedicated sync interface)
  Synchronize Peer IP:       10.0.0.2      (backup node on the sync link)
  Synchronize Config to IP:  10.0.0.2      (XMLRPC, same link)
  Remote System Username:    syncuser
  Remote System Password:    <strong secret>
  Services to synchronize:   rules, NAT, DHCPD, Virtual IPs, users/groups

Naming an explicit peer IP switches pfsync to unicast, which is the recommended production setting: multicast state updates get dropped in most switched fabrics. XMLRPC replicates configuration objects, and it only runs while all CARP interfaces are in MASTER state - a deliberate guard so a broken backup cannot push over a healthy master.

3. Backup node

Configure pfsync on the backup with the master's IP as the peer, and enable Disable preempt so the pair behaves as a group and does not flap. Do not configure XMLRPC synchronisation on the backup - it must never be a source of configuration. Leave the backup in a known-good state; that is your escape hatch if a change breaks the master.

4. Switch-side considerations

Both firewalls should sit in the same layer 2 fabric, and the upstream and downstream switches should not filter CARP multicast or block the VRRP protocol if you choose VRRP-style addresses. Where the fabric cannot be trusted, use the unicast option above and pin the sync link to its own VLAN.

5. Prove failover works

# status: both nodes should show the same roles
System > High Availability > Status

# from a client, keep a session alive and watch it survive
ping -t 10.20.0.1        # continuous ping to the virtual IP
ssh user@10.20.0.50      # long-lived TCP session

# trigger failover
Interfaces > Virtual IPs > select the VIP > Temporarily disable CARP (master)
# or unplug the master's WAN/LAN link

The continuous ping should show no more than one or two lost replies, and the SSH session must stay connected. If the session drops, pfsync is not replicating state - check that both nodes use the same state-sync protocol version (the "Sync compatibility" selector from 24.7 onwards) and that the sync interface is up.

Failure modes worth memorising

  • Both nodes master - duplicate VHID, or the CARP advertisements are not reaching the peer. Test by comparing the source address seen at the peer.
  • Config drifts apart - XMLRPC failed silently. Verify under Status, and confirm the sync user exists on both sides with the same password.
  • Single link failover only - CARP follows interface state; if only the WAN monitor fails, use gateway groups rather than CARP to trigger the change.
  • Failover works, traffic still black-holes - the upstream switch ARP cache still points at the old master's MAC. Short ARP timers or gratuitous ARP support on the switch fix it. The principle is identical to VRRP-based designs such as keepalived VRRP HAProxy virtual IP failover.

Related reading on this site: Keepalived VRRP for HAProxy: Virtual IP Failover Setup and pfSense VLAN Trunk and Firewall Rules Configuration.

原文链接:https://docs.opnsense.org/manual/how-tos/carp.html