Arista EOS BFD: Timers, Protocol Integration and Tuning - 夜莺博客

Arista EOS BFD: Timers, Protocol Integration and Tuning

A BGP hold timer of 180 seconds is fine for stability and catastrophic for availability. When a link fails, everything above it needs to know quickly - and the routing protocol's own timers are deliberately conservative because they were designed for links that flap. BFD exists to decouple those two concerns: it detects failure in tens or hundreds of milliseconds and tells the routing protocol, which then converges without waiting for a timer that was never meant for this purpose.

This article covers EOS BFD configuration, the arithmetic behind timer choices, integration with BGP and static routes, and the failure modes that stop sessions from coming up.

The EOS BFD Model

EOS uses BFD v1 in asynchronous mode. Parameters can be set globally and overridden per interface:

! global defaults
switch(config)# bfd interval 300 min-rx 300 multiplier 3

! per-interface override - takes precedence over the global value
switch(config)# interface ethernet 3/20
switch(config-if-Et3/20)# bfd interval 200 min-rx 200 multiplier 3

! loopback, management and VLAN interfaces are also valid
switch(config)# interface vlan 100
switch(config-if-Vl100)# bfd interval 150 min-rx 150 multiplier 3

! port-channels too
switch(config)# interface port-channel 10
switch(config-if-Po10)# bfd interval 100 min-rx 100 multiplier 3
Parameter Meaning Range Default
transmit_rate How often this device sends BFD control packets 50-60000 ms 300 ms
min-rx The slowest rate this device is willing to accept 50-60000 ms 300 ms
multiplier Missed packets before the session is declared down 3-50 3

The Arithmetic You Must Do Before Choosing Numbers

A BFD session negotiates the slower of the two ends' preferences for both transmit and receive intervals. Detection time is then:

detection time = multiplier x max(local min-rx, remote transmit_rate)

example: both ends configured interval 200 min-rx 200 multiplier 3
  -> negotiated transmit 200 ms, receive expectation 200 ms
  -> detection = 3 x 200 ms = 600 ms

example: local 200 ms, remote 1000 ms multiplier 3
  -> detection = 3 x 1000 ms = 3 s   (the remote's slow rate governs)

Two consequences follow. First, a remote end with default settings can silently undo your tuning - verify the negotiated values, not just your own configuration. Second, detection is a multiple of the interval, so halving the interval halves detection time and doubles the control packet rate. At very short intervals, the packet rate itself becomes a consideration on CPU-scheduled sessions.

Practical starting points:

Scenario Interval Multiplier Detection
General LAN/WAN routing 300 ms 3 900 ms
Latency-sensitive routed core 150 ms 3 450 ms
Hardware-offloaded, aggressive 50 ms 3 150 ms
Flaky link that flaps 1000 ms 5 5 s

If a link flaps physically, aggressive BFD makes it worse: the session tears down, the protocol reconverges, the link comes back, and the cycle repeats. On links with known physical instability, loosen the timers deliberately or fix the link first.

Integrating With BGP

switch(config)# router bgp 64500
switch(config-router-bgp)# neighbor 10.1.1.2 remote-as 64501
switch(config-router-bgp)# neighbor 10.1.1.2 bfd
switch(config-router-bgp)# neighbor 10.1.1.2 bfd interval 200 min-rx 200 multiplier 3
switch(config-router-bgp)# neighbor 10.1.1.2 fallback
! 'fallback' lets the session fall back to the BGP hold timer if BFD cannot establish
switch(config-router-bgp)# neighbor 10.1.1.2 timers 3 9

! verify the session state and the negotiated parameters
switch# show bfd peers
switch# show bfd peers detail
switch# show bfd peers summary
switch# show bgp neighbors 10.1.1.2 | include BFD

The combination that gives meaningful convergence is BFD plus short BGP keepalive and hold timers, so that the fallback path is also fast. With BFD at 600 ms detection and BGP keepalive 3 / hold 9, a failing session is noticed by BFD and the BGP session drops in the same second rather than after three minutes.

Note the fallback statement. Without it, a BFD session that cannot establish - because the peer does not support it, or an ACL drops the control packets - can prevent the BGP session from coming up at all, which turns a monitoring enhancement into an outage. Use it, and if you deliberately want BFD to be mandatory, know that you have made it so.

Integrating With Static Routes

BFD makes static routes genuinely resilient without a routing protocol:

switch(config)# ip route 203.0.113.0/24 10.1.1.2 bfd interval 200 min-rx 200 multiplier 3

! verify
switch# show ip route static | include bfd
switch# show bfd peers 10.1.1.2

When the session drops, the route is withdrawn, and a backup route takes over. This is the single most useful BFD deployment in a network that has not adopted a dynamic routing protocol for its edge paths.

Where BFD Pairs Well - and Where It Does Not

  • Pairs with LACP, but be careful. Using BFD over a port channel tells you the channel is broken; LACP already tells you that. Running both is fine and sometimes useful, but they should not be layered such that a single physical member failure produces two independent teardown triggers. The channel configuration itself is covered in this EOS port channel and LACP guide.
  • Pairs with VRRP and other first-hop protocols to cut failover time, with the same caveat about not creating a feedback loop where a BFD teardown causes a role change that causes a BFD teardown.
  • Does not replace path monitoring. BFD tells you the peer is reachable. It does not tell you the path is good. For measuring latency, jitter and loss against an SLA you want a probe-based tool, and the Junos equivalent described in this RPM probe guide shows what that measurement model looks like.

Why a Session Will Not Come Up

Work through these in order - they cover most real cases:

  1. Check both ends agree. show bfd peers detail on each side; a session showing Init or Down on one end and nothing on the other means the control packets are not arriving.
  2. ACLs and control-plane policing. BFD control packets are UDP, destination port 3784 (and 4784 for multihop), sourced from 3784. An ACL that permits OSPF and BGP but not UDP 3784 is the classic cause. Check the control-plane policy.
  3. Asymmetric configuration on a port channel. BFD over a port channel needs a consistent view on both ends; a member link that is up but not forwarding, or a mismatched hashing configuration, can cause intermittent flapping.
  4. Timing too aggressive for the path. If the path's real round-trip time exceeds the negotiated interval, sessions flap or never establish. On long-haul or satellite paths, start slow and tune down.
  5. Session scale. BFD sessions consume resources. Hundreds of sessions at 50 ms intervals is a different proposition from a handful. Check the platform's supported session count and whether it is hardware-offloaded or CPU-scheduled - CPU-scheduled sessions at very short intervals can be delayed by other control-plane work, producing apparent peer failure that is really local scheduling jitter.
  6. Sub-interfaces and unusual encapsulations. Not every interface type supports BFD, and support varies by platform and release. If it never establishes and the control packets are arriving, check platform support before assuming a configuration error.

Monitoring

switch# show bfd peers summary
switch# show bfd peers detail | include (Peer|State|Interval|Multiplier|Uptime)
switch# show bfd counters
switch# show bfd peers counters
! session telemetry to CloudVision, if deployed
switch(config)# bfd session-telemetry
switch(config)# bfd session-telemetry interval 15

Alert on the session count as well as on individual sessions. A slow leak of session state, or a count that creeps up after a change, is easier to spot in aggregate than by watching individual peers. And treat a BFD flap that does not correlate with a physical event as a signal to check the control plane, not to lengthen the timers - the timers are usually not the problem, they are the messenger. Verifying that the authentication and management plane around all this is coherent is worth doing at the same time, and the AAA and logging configuration for the same platform is covered in this EOS AAA TACACS and RADIUS guide; the equivalent BFD configuration and verification commands on another widely deployed platform are shown in this Nexus 9000 BFD guide.

原文链接:https://www.arista.com/en/um-eos/eos-bidirectional-forwarding-detection