Junos BFD Timers: Liveness Detection That Scales - 夜莺博客

Junos BFD Timers: Liveness Detection That Scales

BFD is the only way to make routing protocols notice a dead neighbour in milliseconds instead of seconds, but on Junos it is configured per protocol rather than globally, and the timers are negotiated rather than forced. That combination trips up engineers migrating from IOS habits: a minimum-interval that looks fast on paper can be silently slowed down during negotiation, and a value that is too aggressive will flap the session and take BGP or OSPF down with it. This guide walks through the Junos BFD configuration model, the timer hierarchy, the recommended values for Routing Engine-based versus distributed sessions, and the show commands that prove the session is actually running at the interval you asked for.

How Junos BFD differs from IOS

On IOS, BFD timers usually live on the interface (or globally) and protocols merely reference the interface. On Junos there is no global BFD block: bfd-liveness-detection is attached under the protocol configuration, and the timers are negotiated with the peer, so both ends must agree. Junos also derives the multihop mode automatically when the BGP session has a multihop TTL configured, so the same syntax covers single-hop and multihop peers.

Item Cisco IOS Junos
Where timers are configured Interface or global BFD block Under each protocol (OSPF, IS-IS, BGP, static)
Enabling per protocol ip ospf bfd on the interface bfd-liveness-detection inside the protocol stanza
Multihop Explicit bfd multihop Automatic with a multihop TTL
Verification show bfd neighbors show bfd session, show bgp neighbor | match bfd

Timers: what each knob really controls

minimum-interval sets both the local transmit interval and the expected receive interval. multiplier is the number of missed packets that declares the session down, so detection time equals interval times multiplier. detection-time threshold overrides that arithmetic with one absolute value, and no-adaptation stops Junos from negotiating to slower timers than you configured.

set protocols ospf area 0.0.0.0 interface ge-0/0/0.0 bfd-liveness-detection minimum-interval 300
set protocols ospf area 0.0.0.0 interface ge-0/0/0.0 bfd-liveness-detection multiplier 3
set protocols bgp group INTERNAL bfd-liveness-detection minimum-interval 200
set protocols bgp group INTERNAL bfd-liveness-detection multiplier 3
set protocols bgp group EBGP-MH bfd-liveness-detection minimum-interval 500
set protocols bgp group EBGP-MH bfd-liveness-detection multiplier 4
set protocols bgp group EBGP-MH bfd-liveness-detection no-adaptation

Two consequences are easy to miss. First, a global OSPF-level bfd-liveness-detection under protocols ospf applies to every interface, and per-interface overrides then refine it. Second, if one side asks for 300 ms and the other is capped at 1000 ms, the negotiated value follows the slower peer — BFD always picks the larger interval.

Choosing intervals that do not flap

Juniper's own guidance is blunt: intervals below 100 ms on Routing Engine-based sessions, or below 10 ms on distributed sessions, invite flapping. During a Routing Engine switchover, processes such as RPD, MIBD and SNMPD consume CPU for longer than a fast timer allows, so a tight BFD session drops exactly when the control plane is busiest.

  • General Routing Engine switchover: use 5000 ms.
  • Dual chassis cluster with a failed control link: use 6000 ms so the secondary does not flap LACP through BFD.
  • Large deployments: 300 ms for Routing Engine-based sessions, 100 ms for distributed ones.
  • With nonstop active routing (NSR) enabled: at least 2500 ms.

Those numbers look conservative next to the 50 ms promises quoted in marketing decks, but a BFD session that resets a full BGP table is far more expensive than a four-second detection time. For link-level failure detection where the routing protocol will recompute anyway, consider pairing BFD with the faster IGP hello timers instead of chasing single-digit milliseconds. See BGP hold and keepalive timers for how protocol timers interact with BFD-driven teardown.

Authentication and larger topologies

Where the session crosses an untrusted segment, key chains keep a rogue host from forging BFD packets and tearing the adjacency down. The loose-check option lets the session stay up if the peer is not authenticating, which is useful while you roll the change out one side at a time.

set security authentication-key-chains key-chain BFD-KEY key 1 secret "$9$example"
set protocols ospf area 0.0.0.0 interface ge-0/0/0.0 bfd-liveness-detection authentication key-chain BFD-KEY
set protocols ospf area 0.0.0.0 interface ge-0/0/0.0 bfd-liveness-detection authentication algorithm sha-1
set protocols ospf area 0.0.0.0 interface ge-0/0/0.0 bfd-liveness-detection authentication loose-check

For static routes the stanza lives under routing-options and needs an explicit neighbour so Junos knows which next hop to track — without it, the static route never inherits the session.

set routing-options static route 192.168.100.0/24 next-hop 10.0.12.2
set routing-options bfd-liveness-detection minimum-interval 300
set routing-options bfd-liveness-detection multiplier 3
set routing-options bfd-liveness-detection neighbor 10.0.12.2

Verification that actually proves the timer

show bfd session
show bfd session extensive | match "Detect Time|Transmit Interval"
show ospf neighbor detail | match BFD
show bgp neighbor | match bfd
monitor start bfd

show bfd session extensive is the one that matters: it prints transmit interval, receive interval and detection time as negotiated, not as configured. If the detection time is larger than minimum-interval times multiplier, the peer is slowing you down. When a session refuses to come up at all, check for mismatched interface MTU, an ACL blocking UDP 3784 (single hop) or 4784 (multihop), and whether the far end has BFD under a different protocol instance. Adjacency failures with BFD enabled often track back to ordinary BGP problems — see Junos BGP establishment troubleshooting before blaming BFD.

Finally, put the session on the monitoring system. A BFD flap is a leading indicator of a physical problem, and the counters are exposed through standard MIBs described in SNMPv3 configuration on Cisco IOS and Junos.

原文链接:https://juniper.net/documentation/en_US/junos/topics/example/bgp-bfd-ibgp.html