Juniper BGP Session and Route Flaps: Prevention Guide - 夜莺博客

Juniper BGP Session and Route Flaps: Prevention Guide

BGP session flapping is one of the most disruptive events on a Junos router: every flap withdraws and re-advertises routes, triggering route churn across the entire network. This guide, based on the official Juniper Networks Junos documentation, explains why BGP session flaps cascade when VPN address families are configured on route reflectors and AS boundary routers, and shows the exact configuration statements that stop unnecessary flaps before they impact production traffic. You will learn practical prevention techniques including passive EBGP sessions, BGP-static routes, and address-family-level damping.

Why BGP Sessions Flap on Junos OS

When a router or switch is configured as a route reflector (RR) or an AS boundary router (an external BGP peer), and a VPN family such as family inet-vpn unicast is configured, a flap of either the RR IBGP session or the EBGP session causes flaps of all other BGP sessions configured with the same family. The root cause is that Junos re-advertises the full VPN table on session re-establishment, so one unstable session drags down the whole group. Understanding this behavior is the first step of BGP session flapping troubleshooting on Juniper equipment.

Why one flap becomes many: the amplification effect

The amplification is worth spelling out, because it explains why the symptom looks far bigger than the cause. When a VPN family is present, the amount of state Junos must rebuild on session re-establishment is not one prefix — it is the entire VPN routing table for that family. Concretely:

  • A session in a peer group with family inet-vpn unicast goes down. Junos tears down the adjacency and starts re-advertising.
  • The re-advertisement is batched through the shared update groups that the peer group members have in common. Sessions in the same update group share advertisement state, so a change that invalidates the group invalidates every member’s view at once.
  • Because the VPN NLRI count is orders of magnitude larger than plain IPv4 unicast, the resulting update burst is large enough to affect the control plane and, on smaller platforms, the forwarding plane’s ability to program new paths.
  • Any peer whose TCP session cannot absorb the burst — small receive window, congested path, low keepalive headroom — times out and flaps in turn. That is the cascade.

The trigger is usually external: an interface flap, a fibre fault, a keepalive timeout caused by a saturated control plane, or an MTU/fragmentation problem on the BGP path. The fix in this article does not attempt to make the underlying fault disappear; it removes the mechanism that turns a single fault into a network-wide event.

Verifying That Unnecessary Session Flaps Are Occurring

Use the following procedure to confirm that VPN-family configuration is amplifying session flaps:

  1. Run show bgp summary to verify that the sessions have been established.
  2. Deactivate the EBGP session with deactivate protocols bgp group ebgp.
  3. Run show bgp summary again and observe the flap counters on the remaining sessions.
show bgp summary
show bgp summary | match "Flaps"
show bgp neighbor 192.0.2.1 | match "Last flap"
show bgp neighbor 192.0.2.1 | match "Flaps:"
show log messages | match "BGP_TRACE"

An hour or two of flap counters is not enough evidence. Junos records the total flap count and the time of the last flap per neighbour; a neighbour with several flaps in a single maintenance window, correlated with an interface event in show log messages, is the signature you are hunting for. Correlate the flap timestamps against interface and routing events in the log rather than assuming BGP itself is at fault.

Baseline evidence: traceoptions without the noise

Before changing configuration, capture what is actually happening. BGP traceoptions at the right flags produce a manageable log; at the wrong flags they will fill the disk and make the problem worse.

set protocols bgp traceoptions file bgp-trace.log size 10m files 5
set protocols bgp traceoptions file world-readable
set protocols bgp traceoptions flag state
set protocols bgp traceoptions flag update
set protocols bgp traceoptions flag keepalive
set protocols bgp traceoptions flag normal

Keep flag packet off unless you are chasing a specific handshake failure — it logs every message and will dominate the file. Rotate with size and files so the trace cannot exhaust the disk, and remove the configuration once the investigation is finished. Our Junos traceoptions for BGP guide covers the flag set in detail.

Preventing Flaps with a Passive EBGP Session

The recommended fix is to add a passive EBGP session with a neighbor address that does not exist in the peer autonomous system. Because the session is passive, Junos never attempts to actively connect to it, yet the VPN address family still has a valid configuration target, which stops the cascading flap behavior:

set protocols bgp group ebgp passive
set protocols bgp group ebgp neighbor 192.0.2.254

After committing, run show bgp summary to verify that the real sessions are established and the passive session stays idle. Three cautions apply:

  • Use the neighbour address at group level or as an explicit stub neighbour. The intent is to keep the group’s family configuration anchored to a stable member, so that deactivating or flapping a real neighbour does not invalidate the group’s VPN advertisement state.
  • Pick an address that genuinely does not exist. Document it as reserved. If somebody later assigns it to a live device, you get an unexpected session.
  • Do not confuse passive with a substitute for fixing the underlying fault. A flapping physical link still needs fixing; the passive session only stops the amplification.

Also review hold-time while you are in the group. The default of 90/30 seconds is tuned for stability, not convergence; extending hold time on a link you know to be flaky buys you a real reduction in session teardown, at the cost of slower failure detection. For fast detection with stability, use BFD instead of shortening BGP timers.

Configuring BGP-Static Routes to Prevent Route Flaps

Beginning with Junos OS Release 14.2, you can configure and advertise BGP-static routes in a BGP network. A BGP-static route can be advertised even when it is not the active route for the prefix, and it does not flap unless deleted manually. This is ideal for stable prefix origination:

set routing-options static route 203.0.113.0/24 discard
set policy-options policy-statement export-bgp term 10 from protocol static
set policy-options policy-statement export-bgp term 10 then accept
set protocols bgp group ibgp export export-bgp

The distinction that matters is between a route that exists because an IGP found a path to it and a route that is declared. An IGP-learned prefix disappears the moment the IGP loses reachability, and BGP dutifully withdraws it. A static (or BGP-static) prefix stays in the BGP table regardless of what the underlay is doing, so downstream peers see no churn. Use it for the prefixes you actually want to be stable: customer summaries, service addresses, anycast endpoints, and the loopbacks that anchor sessions and next hops.

Two practical notes:

  • Pair the static route with a discard or reject next hop so that traffic for a prefix with no real path is dropped cleanly rather than following a stale route.
  • Do not blanket-static prefixes whose reachability is genuinely dynamic. Hiding a real failure behind a static route is worse than the flap you were trying to avoid.

Route Flap Damping at the Address-Family Level

Since Junos OS Release 12.2, flap damping can be applied at the address-family level, for example on family inet-mvpn signaling damping. Damping suppresses unstable prefixes instead of letting every flap propagate to the rest of the network, which is a valuable complement to session-level prevention.

set protocols bgp group ibgp family inet-vpn unicast damping
set protocols bgp group ibgp family inet-vpn unicast damping half-life 15
set protocols bgp group ibgp family inet-vpn unicast damping reuse 750
set protocols bgp group ibgp family inet-vpn unicast damping suppress 3000
set protocols bgp group ibgp family inet-vpn unicast damping max-suppress 60
set protocols bgp group ibgp family inet-vpn unicast damping disable

Damping works by assigning each prefix a penalty that decays over time. Every withdraw or attribute change adds penalty; if the total crosses suppress, the prefix is hidden from peers until decay brings it back under reuse. The classic failure of damping is being too aggressive: in the early 2000s, default parameters suppressed stable prefixes during a large-scale event and caused more harm than the flapping itself. That is why the widely repeated advice is to be conservative — long half-life, high thresholds — or to leave damping off and rely on session-level stability instead. Read our breakdown of route flap damping penalty, half-life and reuse thresholds before turning it on for a production address family.

Session-level protection beyond the passive session

  • BFD. Sub-second failure detection with a stable BGP hold time is the right combination for links where you want fast reconvergence without the false teardowns that short BGP timers cause.
  • Graceful restart and LLGR. Graceful restart lets a peer keep using the routes it learned from you while your control plane restarts. On VPN families, long-lived graceful restart (LLGR) extends the same idea to prefixes that would otherwise be stale-route-deleted.
  • TCP authentication (MD5/ao). A TCP MD5 mismatch produces a session that never establishes — and a well-meaning engineer troubleshooting by toggling the peer will produce exactly the flap pattern you are trying to eliminate. Turn on tcp-auth-keys deliberately and fully, not partially.
  • Keepalive and hold timer discipline. The defaults exist because they are stable. Shortening them to 10/30 across a large peer group multiplies control-plane load and makes a transient oversubscription far more likely to become a teardown.
  • Keep the advertised set small. Tight export policy and disciplined aggregation reduce the volume that must be re-sent every time a session re-establishes, which directly reduces the cost of a flap.

Underlay and platform causes of flap

The majority of genuine BGP flaps on Junos originate below BGP. Before tuning BGP, rule these out:

  • Interface errors. CRC, framing and input errors on the underlying link cause intermittent reachability loss to the peer address. Check show interfaces extensive | match error.
  • MTU and fragmentation. Large update bursts fragmented and dropped on the path cause the session to stall and then time out.
  • Control-plane saturation. A routing engine that is busy rebuilding a table cannot service keepalives promptly. Watch RE CPU and the BGP task utilisation during the flap.
  • Aggressive uRPF or firewall filters. An inbound filter that drops TCP 179 with the wrong source address, or that drops fragments, gives a session that flaps under load but looks idle at rest.
  • Asymmetric multi-hop reachability. A multihop EBGP session whose return path differs from the forward path will establish and then fail unpredictably — usually after a routing event.

Monitoring and alerting

Prevention is only half the job; you also want to know when the mechanism is leaking. Useful checks, in the order of value:

show bgp summary | match "Flaps"
show bgp neighbor | match "Flaps|Last flap"
show bgp neighbor | match "Notifications"
show bgp neighbor | match "Received|Sent"
show task memory
show system processes extensive | match rpd

Alert on any neighbour whose flap count increments outside a change window, and on any drop in the number of established VPN-family sessions. Also watch rpd memory and CPU: a route reflector carrying a full VPN table reacts to a flap with a measurable spike, and that spike is a leading indicator that the next cascade will hurt.

Prevention checklist

  • Add a passive EBGP session with a documented, unused neighbour address so a flapping peer cannot invalidate the group’s VPN advertisement state.
  • Advertise service prefixes as BGP-static routes so they do not flap with the underlay.
  • Keep damping off, or configure it conservatively at the address-family level.
  • Use BFD for fast failure detection rather than shortening BGP keepalive and hold timers.
  • Correct interface, MTU and control-plane problems before tuning BGP at all.
  • Monitor flap counters, session counts and rpd CPU, with an alert on unexplained increments.

Related Reading

原文链接:https://www.juniper.net/documentation/us/en/software/junos/bgp/topics/topic-map/bgp-session-flaps.html