Junos BGP Establishment Troubleshooting: Commands - 夜莺博客

Junos BGP Establishment Troubleshooting: Commands

A BGP session that refuses to come up leaves the network silently broken, and the Junos CLI gives you every tool needed to find out why - if you know where to look. This article, based on the Network Curiosity Junos BGP troubleshooting series, walks through the main troubleshooting tools (show bgp summary, show bgp neighbor, log files, traceoptions, and monitor traffic interface) and the common scenarios that prevent BGP from establishing, including filtered BGP traffic, shutdown peers, prefix limits and session state analysis.

The BGP Finite State Machine, Briefly

Every establishment failure is somewhere in the state machine, so knowing what each state means narrows the search before you run a single command.

  • Idle — the session is administratively down, or the router is waiting for a retry after a failure. A peer that never leaves Idle is usually shut down locally, or the peer address is not accepted by the local configuration.
  • Connect — a TCP connection is in progress. The TryConnect flag in show bgp neighbor means the local router is trying and receiving no answer.
  • Active — the router is retrying the TCP connection on the ConnectRetry timer, typically every two minutes. This is where filtered port 179, missing routes to the peer, and wrong local addresses all show up, and it is also the state name that appears in show bgp summary for any session that is not established.
  • OpenSent — the TCP session is up and an OPEN has been sent; the router is waiting for the peer's OPEN. Stuck here usually means the peer is answering TCP but not running BGP, or the OPEN is being dropped by a filter.
  • OpenConfirm — OPENs have been exchanged and the router is waiting for the first KEEPALIVE. Stuck here points at an authentication mismatch, a hold-time negotiation problem, or a one-way packet path.
  • Established — the session is up. Failures from here are not establishment problems and belong to flap, prefix-limit or policy analysis.

Timers matter as much as states: the ConnectRetry interval governs how often a failed attempt repeats, while the negotiated hold time (90 seconds by default, keepalive one third of it) governs how quickly a silent peer is declared dead. When you are watching a trace, a message that repeats exactly every 120 seconds is a ConnectRetry loop, not a protocol error.

Main BGP Troubleshooting Tools on Junos

  • Show commands - show bgp summary and show bgp neighbor give an overview of configured sessions and their current state.
  • Log files - such as the messages log file, capture BGP events around the failure.
  • Traceoptions - specific debug configuration to trace BGP events at protocol level.
  • Monitor traffic interface - acts like a packet capture, showing traffic destined to the router on a given interface to verify TCP 179 connectivity.
  • ping and traceroute - confirm the peer address is reachable at all, and where the path breaks if it is not.

Reading show bgp summary

show bgp summary
show bgp summary | match "10.0.0.2"
show bgp neighbor 10.0.0.2

The summary table carries more information than it appears to. Alongside each peer you get the peer AS, the session state, the number of prefixes received (Active/Received/Accepted/Damped), the flap count, and the time since the session last went up or down. Two readings repay attention. First, the state column: Junos prints Establ for a working session and Active for one that has failed, which is easy to misread as "the session is doing something". Second, the prefix counters: Active/Received/Accepted tells you whether routes arrived at all, whether they were accepted by import policy, and how many were rejected. A session that is established with received-but-not-accepted counters is a policy problem, not an establishment problem.

Reading the Session State

A session stuck in the Connect state - with the BGP neighbor output showing the TryConnect flag - indicates the router is attempting to establish the TCP connection but the peer is not responding. Check whether the peer is reachable, whether TCP port 179 is filtered between the devices, and whether the local address used for peering is correct.

Decoding show bgp neighbor

show bgp neighbor 10.0.0.2
show bgp neighbor 10.0.0.2 | match "Last|Error|State|flags"
show bgp neighbor 10.0.0.2 | match "Local|Remote|Hold"

This command is the single most valuable output on a failing session. It shows the configured and negotiated hold time, the local and remote addresses used for the session, the type (internal or external) and the peer AS, the current flags, and — most importantly — Last State, Last Event and Last Error. Those three fields record why the session last dropped, in the protocol's own words. A Last Error containing a NOTIFICATION code and subcode is the diagnosis; everything else is corroboration. If Last State is blank, the session has never been up, which immediately tells you to look at transport and configuration rather than at what changed.

BGP Notification Codes Worth Memorising

Code Meaning Typical cause
1 Message header error Malformed or truncated TCP segment; often an MTU problem.
2, sub 2 Open message error — bad peer AS The remote AS number does not match what the local peer expects.
2, sub 3 Open message error — bad BGP identifier Two peers have the same router ID, or an invalid one.
2, sub 5 Open message error — authentication failure An MD5/TCP-MD5 key mismatch between the two ends.
2, sub 6 Unacceptable hold time Hold time outside the supported range or a configured minimum that is not met.
3 Update message error Malformed UPDATE, commonly from a software defect or an illegal attribute.
4 Hold timer expired No KEEPALIVE received within the negotiated hold time; a silent path or a heavily loaded peer.
5 Finite state machine error Unexpected message for the current state.
6, sub 1 Cease — max prefixes exceeded The peer sent more routes than the configured prefix limit allows; the session is torn down deliberately.

Codes 1 and 2 are configuration problems, codes 4 and 6 are capacity or path problems, and code 6 sub 1 is the one you will meet most often on a production edge that has grown past its configured limit.

Common Causes of BGP Establishment Failure

BGP Traffic Being Filtered

Firewall or ACL rules in the path can silently drop BGP traffic. This is less likely with eBGP sessions, which typically run over a direct interface, but more likely with iBGP peering to loopback addresses where intermediate devices sit in the path. Verify with monitor traffic interface that TCP packets to port 179 arrive at the router.

BGP Peer Is Shutdown

Check show bgp summary for the peer state: a peer showing Active for a long time while the neighbor's session never moves to Established often means the remote side has the session administratively disabled or is unreachable.

Missing Route to the Peer Address

With loopback-based peering — the normal iBGP design — the peer address is reachable only if an IGP carries it. When the IGP breaks, every iBGP session that rides on those loopbacks goes down at once, and the symptom is a set of sessions stuck in Active with the peer address unreachable. Confirm with show route 10.0.0.2 before touching BGP configuration, and fix the IGP or the static route that carries the loopback.

Wrong Local Address

A BGP group configured with local-address must use an address that is actually configured and up. A stale local-address left over from a renumbered link produces a TCP connection attempt sourced from an address the peer has no route to, which looks identical to a firewall problem from the local end. Compare show bgp neighbor output with show interfaces terse | match inet.

Authentication Mismatch

TCP-MD5 authentication, configured with authentication-key, must match exactly on both ends. A mismatch produces a NOTIFICATION code 2 subcode 5, and — depending on platform and version — a session that fails immediately after the OPEN exchange. It is the first thing to check whenever a session reaches OpenConfirm and then drops.

Hold Time and Timer Mismatch

Hold times are negotiated, so a straightforward mismatch is not fatal; what is fatal is a configured minimum that the peer's proposal cannot satisfy, which produces code 2 subcode 6. Short hold times on a congested link are a separate problem: a peer that cannot deliver keepalives inside the window will be declared dead repeatedly, producing code 4 and a session that flaps rather than fails.

Peer AS or Type Mismatch

Configuring type internal for an external peer, or setting the wrong peer-as, produces a code 2 subcode 2 notification at OPEN. On a new turn-up this is usually a copy-paste error in the group template.

MTU and MSS Problems

A session that establishes and then drops as soon as the first full UPDATE is sent may be hitting an MTU problem. A large UPDATE is a large TCP segment and must survive the path; where a tunnel or a carrier network is in the middle, the router's own traffic can be fragmented or dropped. Clamp the TCP MSS on the peering interface, or reduce the interface MTU, and watch for code 1 header errors in the meantime.

Prefix Limits and Route Flooding

A router sending too many routes may see the BGP session reset or fail to establish, depending on how the prefix limit is configured on the receiving side. Review the configured prefix limits and the received route counts on both peers.

eBGP Multihop and TTL

An eBGP session between addresses that are not directly connected needs multihop with a TTL large enough to cross the path, and it needs a route to the peer to exist in the first place. Without it, the session sits in Connect while the SYN leaves with a TTL of 1.

Step-by-Step Diagnosis Flow

  1. Run show bgp summary and note the session state (Idle, Connect, Active, OpenConfirm, Established).
  2. Run show bgp neighbor for the failing peer and check flags and last error codes.
  3. Check the messages log and BGP traceoptions output for the last events.
  4. Use monitor traffic interface to confirm whether TCP SYN packets reach the router and whether replies come back.
  5. Verify routing to the peer address and the local-address configured for the BGP group.
  6. Compare configuration on both ends: AS numbers, group type, local and peer addresses, authentication keys and hold times.
  7. If the session is up but routes are missing, switch to policy and prefix-limit analysis rather than session troubleshooting.

Proving TCP 179 Connectivity

monitor traffic interface ge-0/0/0.0 matching "tcp port 179" no-resolve
monitor traffic interface ge-0/0/0.0 matching "tcp port 179" extensive no-resolve
ping 10.0.0.2 count 5
traceroute 10.0.0.2

The packet capture answers the question that no show command can: is the SYN arriving at all, and if so, is a SYN-ACK going back? Outbound SYNs with no replies point at a filter or a routing problem on the far side; inbound SYNs with no reply from the local router point at a local filter on the loopback interface, which is a classic iBGP failure when the lo0 filter is written for a different address family or a different protocol. Remember that a filter on lo0 is the one that matters for loopback-based peering, and that Junos filters lo0 traffic only if you explicitly apply a filter there.

Traceoptions Tailored to Establishment Problems

set protocols bgp group my-internal-group neighbor 10.0.0.2 traceoptions file bgp-int size 5m files 4
set protocols bgp group my-internal-group neighbor 10.0.0.2 traceoptions flag state
set protocols bgp group my-internal-group neighbor 10.0.0.2 traceoptions flag error
set protocols bgp group my-internal-group neighbor 10.0.0.2 traceoptions flag packet
commit
show log bgp-int | match "10.0.0.2" | last 30

Three flags are enough for establishment work, and the trace must be bounded and removed afterwards. The tail of the trace separates the three failure modes cleanly: connect failures repeated on the retry timer mean transport; an OPEN followed by a NOTIFICATION means configuration or authentication; TCP up with no BGP messages at all means the remote side is not running BGP on that address. The full production-safe workflow, including file sizing and cleanup, is in our guide to Junos traceoptions for BGP troubleshooting.

What to Do Once the Session Is Up

  • Watch the flap counters for an hour after the fix; a session that comes up is not automatically a session that stays up.
  • Check the received and accepted prefix counters against the peer's advertised view.
  • Confirm the route selection outcome — a working session with the wrong best path is still an outage.
  • Remove the traceoptions configuration and record the evidence in the change ticket.
  • If the cause was a missing route, a filter or an MTU problem, fix the underlying design rather than the session: the same fault will hit the next peer on the same design. Interface-level instability behind these symptoms is covered in our notes on Junos interface flapping, hold time and damping, and the damping side of repeated drops in Junos BGP flap damping parameters.

Related Reading

原文链接:https://www.networkcuriosity.com/junos-bgp-establishment-troubleshooting