LVS/IPVS Load Balancing with keepalived: NAT vs DR - 夜莺博客

LVS/IPVS Load Balancing with keepalived: NAT vs DR

Before HAProxy, before Kubernetes services, there was IPVS: a kernel-level layer-4 load balancer that has been in the mainline Linux kernel for two decades and still handles more packets per watt than any userspace proxy. It is also the fastest way to get into trouble, because the choice between NAT and Direct Routing is a topology decision, not a preference — and the ARP behaviour of DR real servers will break a production LAN if you skip the sysctls. This article walks both modes with complete, verifiable configurations.

The Three Forwarding Methods

Method Request path Reply path Constraint
NAT (masquerading, -m) Client → director DNAT → real server Real server → director SNAT → client Real servers need their default route via the director
DR (direct routing, -g) Director rewrites only the L2 destination Real server replies to the client directly Director and real servers on the same L2 for the VIP; VIP on lo with ARP suppressed
TUN (-i) Director encapsulates (IPIP/GRE) Real server replies directly Real server must decapsulate; more moving parts

Start with NAT if you want the fewest network surprises. Choose DR when return bandwidth should leave the director — the classic high-throughput HTTP/TCP farm on a single LAN.

Mode 1 — NAT with keepalived

Director sysctls

cat >/etc/sysctl.d/99-ipvs-nat-director.conf <<'EOF'
net.ipv4.ip_forward = 1
# only needed if netfilter stateful rules must see IPVS flows
# net.ipv4.vs.conntrack = 1
EOF
sysctl --system

keepalived virtual server

global_defs {
    router_id lvs-nat-1
    lvs_flush
    lvs_flush_on_stop
}

vrrp_instance VI_HTTP {
    state BACKUP
    interface eth0
    virtual_router_id 51
    priority 100
    advert_int 1
    authentication { auth_type PASS  auth_pass change-me }
    virtual_ipaddress { 203.0.113.10/24 dev eth0 }
}

virtual_server 203.0.113.10 80 {
    delay_loop 6
    lvs_sched wlc
    lvs_method NAT
    protocol TCP
    alpha                     # assume real servers down until checks pass
    omega
    quorum 1
    inhibit_on_failure        # set weight 0 instead of removing the server
    real_server 10.0.0.11 80 {
        weight 1
        TCP_CHECK { connect_port 80  connect_timeout 3  retry 2 }
    }
    real_server 10.0.0.12 80 {
        weight 1
        HTTP_GET {
            url { path /healthz  status_code 200-299 }
            connect_timeout 3
            retry 2
        }
    }
}
keepalived -t -l                # syntax test where supported
systemctl enable --now keepalived
ipvsadm -Ln

Three flags deserve comment. alpha starts checkers pessimistic so a director restart does not briefly advertise dead backends. inhibit_on_failure sets a failed real server's weight to zero, which IPVS treats as quiescent: no new jobs, existing connections can finish. And health checks should test something the user needs — /healthz, not merely a port that still accepts SYNs on a wedged application.

Mode 2 — Direct Routing

Real server preparation (every backend)

cat >/etc/sysctl.d/99-ipvs-dr-realserver.conf <<'EOF'
net.ipv4.conf.all.arp_ignore = 1
net.ipv4.conf.lo.arp_ignore = 1
net.ipv4.conf.all.arp_announce = 2
net.ipv4.conf.lo.arp_announce = 2
EOF
sysctl --system

ip addr add 203.0.113.10/32 dev lo      # VIP on loopback, never on eth0
ip route add 203.0.113.10/32 dev lo

Why this works: arp_ignore=1 answers ARP only when the target address is configured on the incoming interface, so a VIP living on lo never responds to ARP arriving on eth0. arp_announce=2 stops the VIP from becoming the source address of ARP requests on the wrong interface. This pair is the modern replacement for the old hidden sysctl and, if you skip it, two hosts will fight over the VIP and clients will intermittently hit the wrong server.

Director configuration (DR)

virtual_server 203.0.113.10 80 {
    delay_loop 5
    lvs_sched wrr
    lvs_method DR
    protocol TCP
    alpha
    inhibit_on_failure
    real_server 203.0.113.11 80 { weight 2  HTTP_GET { url { path /healthz  status_code 200-299 } } }
    real_server 203.0.113.12 80 { weight 1  HTTP_GET { url { path /healthz  status_code 200-299 } } }
}

In DR and TUN modes the real server port must equal the virtual service port, and the application must accept connections destined to the VIP (bind 0.0.0.0 or the VIP explicitly) — a service bound only to the RIP will silently refuse traffic.

Manual ipvsadm Equivalents (for debugging)

ipvsadm -A -t 203.0.113.10:80 -s wlc
ipvsadm -a -t 203.0.113.10:80 -r 10.0.0.11:80 -m -w 1     # NAT
ipvsadm -a -t 203.0.113.10:80 -r 203.0.113.12:80 -g -w 1 # DR
ipvsadm -Ln --stats
ipvsadm -Ln --rate
ipvsadm -Lnc | head

Verification and Failure Modes

  • All traffic to one backend. Check the scheduler and weights; a heavy real server with wlc legitimately takes more connections, and non-persistent HTTP/1.0 flows are too short to distribute evenly. Confirm with ipvsadm -Ln --stats.
  • Intermittent timeouts in DR. ARP policy missing on the clients' LAN or on the real servers. Verify with ip neigh and a capture on the backend.
  • Director restart floods dead backends. Use alpha and non-zero quorum.
  • Director pair loses established connections. VRRP moves the VIP but not L4 state; add lvs_sync_daemon <iface> inst VI_HTTP id 51 to sync IPVS connection tables between the two directors.
  • Rollback: systemctl stop keepalived; ipvsadm -C and remove the lo VIP on real servers.

Related posts: keepalived, VRRP and HAProxy failover, nginx stream TCP/UDP load balancing and kube-proxy: iptables vs IPVS vs nftables.

原文链接:https://dev.to/lyraalishaikh/stop-sending-all-traffic-to-one-backend-practical-ipvs-load-balancing-with-keepalived-on-linux-5cdm