Junos Troubleshooting Commands Reference: Route to Flows - 夜莺博客

Junos Troubleshooting Commands Reference: Route to Flows

Having a reliable list of Junos operational commands at hand turns chaotic troubleshooting sessions into structured investigations. This reference, compiled by a network engineer working daily with Junos, groups the most useful commands by functional area — generic diagnostics, BGP, OSPF, forwarding table, security flows and IPsec — so you can jump straight to the right show command for the problem you are facing. It also explains how to enable traceoptions when standard outputs are not enough.

How to Approach a Junos Troubleshooting Session

Junos rewards a methodical approach more than most network operating systems, for one structural reason: the control plane and the forwarding plane are separate, and the commands that read them are separate too. show route reads the routing table the Routing Engine built. show route forwarding-table reads what the Packet Forwarding Engine actually programmed. When those two disagree, you have a control-plane problem. When they agree but traffic still fails, you have a data-plane, policy or physical problem — and no amount of staring at the routing table will help.

That single split drives a useful order of operations for almost any Junos incident. First establish reachability and interface state. Then confirm the route exists and the next hop resolves. Then confirm the forwarding table has the same answer. Only then move into protocol-specific debugging, and only then enable traceoptions, because traceoptions consume CPU and fill the disk and are far more useful once you know which protocol is actually misbehaving.

The reference below follows that order. Each section lists commands that are safe to run in production; the few commands that carry a cost — traceoptions, extensive outputs on a busy box — are called out explicitly.

Step 1: Baseline Health and Reachability

Start with the things that tell you whether the device is fundamentally healthy, and finish with a reachability test that uses the routing instance you actually care about. Many "routing problems" are power events, fan failures or a flapping uplink that nobody noticed.

show chassis alarms
show system alarms
show chassis environment
show system uptime
show interfaces terse
show interfaces <if> extensive

show interfaces terse gives you the fastest possible view of every interface's up/down state and address family. Once you have narrowed to one interface, show interfaces extensive gives you the error taxonomy that matters: input errors, CRC, framing, runts, giants and the flap counter. If you are unsure which of those fields actually indicates a physical fault, the field-by-field explanation in Junos show interfaces extensive error fields breaks them down, and Junos interface flapping layer 1 troubleshooting walks the physical checklist when the flap counter is climbing.

Reachability tests should always be source-aware. Testing from the default routing instance when the traffic in question arrives in a VRF tells you nothing useful.

ping 10.0.0.1 source 10.0.0.2 rapid count 100
ping 10.0.0.1 routing-instance CUSTOMER-A source 10.0.0.2 count 20
traceroute 10.0.0.1 routing-instance CUSTOMER-A
traceroute 10.0.0.1 no-resolve source lo0.0

Enabling traceoptions for Protocol Debugging

set <protocol> traceoptions flag all
set <protocol> traceoptions file <name> size 100m

Traceoptions can be enabled in the hierarchy of the specific protocol (bgp, ospf, isis, rsvp, and so on) and help with immediate identification of the issue type; press ? after flag to list available options.

Two habits make traceoptions safe in production. First, always set a file size limit and a file count so the daemon cannot fill the disk — an unconstrained traceoptions file is a classic cause of a self-inflicted outage. Second, narrow the flags before you enable them. flag all is fine for a lab or for a session that will not establish; on a full BGP table it produces an enormous volume and slows the protocol daemon.

set protocols bgp traceoptions file bgp-log size 50m files 5
set protocols bgp traceoptions flag state
set protocols bgp traceoptions flag packets
commit
run show log bgp-log | last 100

Remember to deactivate the traceoptions once the investigation ends. A forgotten debug that runs for six months is a slow-motion incident. For a deeper workflow, including how to filter traceoptions output down to the interesting neighbours, see Junos traceoptions: debug BGP without guessing.

Generic Junos Troubleshooting Commands

show route
show route table <table-name>
show route protocol <protocol> table <table-name>
show route hidden table inet.0
show route summary
ping source <address> rapid count <n>
traceroute routing-instance <name>
show interfaces detail
show interfaces extensive
monitor interface
show log messages
show chassis alarms
show chassis hardware detail
show chassis fpc

Two of these deserve emphasis. show route hidden table inet.0 lists routes that the device learned but refused to install — the reason is always shown, and it is one of the fastest ways to explain a "missing route" that is actually present in the protocol database. monitor interface is a rolling, per-second view rather than a cumulative counter, which makes it the right tool for confirming that an interface is passing traffic at all before you go looking for something subtler.

BGP and OSPF Diagnostics

show bgp summary
show bgp neighbor
show bgp neighbor <address> | match "Last|State|Error"
show route receive protocol bgp extensive
show route advertising protocol bgp hidden
show ospf route
show ospf database detail
show ospf neighbor extensive
show ospf interface detail
show ospf statistics

The neighbour-focused commands are the ones that resolve most BGP incidents. show bgp neighbor <address> prints the last state transition, the last error and the notification data — that triple is usually enough to name the cause outright, whether it is a hold-timer mismatch, an authentication failure, a capability disagreement or an administrative shutdown on one side. If the session is stuck rather than flapping, the state-by-state diagnosis in BGP session stuck in Idle or Active and Junos BGP establishment troubleshooting covers both the Idle and Active cases in detail.

A practical tip for both protocols: always check both directions. A prefix that is present in the routing table but never advertised usually means an export policy term is rejecting it, and show route advertising protocol bgp hidden surfaces exactly that, including the policy that suppressed it.

For forwarding plane verification use show route forwarding-table destination <prefix> to see the actual next hop programmed in the PFE.

Security Flows and IPsec on SRX

show security flow session
show security flow status
show security flow statistics
show security policy
show security ike security-association
show security ike security-association index <#> detail
show security ipsec security-association
show security ipsec statistics
show security ipsec next-hop-tunnels
monitor interface st0.x
show interfaces extensive st0.x
show security pki local-cert detail

For route-based VPNs, combine show security flow session tunnel with monitor interface st0.x to confirm encrypted traffic is flowing over the tunnel interface.

Work IPsec faults outside-in. Confirm that IKE phase 1 is up before you debug phase 2 — if there is no IKE security association, the IPsec output will be empty and misleading. Then confirm that a flow session exists for the traffic in question and that its policy name is the one you expected, because a traffic-selector mismatch or a policy that does not match produces a session that looks perfectly healthy while carrying nothing. See IPsec IKEv2 troubleshooting: SA_INIT and IKE_AUTH debugs for the phase-1 walkthrough and Juniper SRX route-based site-to-site IPsec VPN for the configuration side.

Step 5: Logs, Change History and Why They Matter

Junos keeps an unusually complete record of what changed and who changed it, and reading that record early often shortcuts an entire investigation. show system commit lists commit history with the responsible user; comparing the active configuration against an older rollback with show | compare rollback N shows the exact delta. If an incident began at a specific time, correlate that time against the commit list before you spend an hour on counters.

show system commit
show configuration | compare rollback 1
show log messages | last 50
show log messages | match "error|Error|ERROR"
show log interactive-commands

The interactive-commands log records operator activity, which is the fastest way to answer "did someone touch this box?" — a question that, in practice, resolves a surprising share of unexplained incidents.

Common Pitfalls That Waste the Most Time

Five traps account for a disproportionate share of long Junos troubleshooting sessions, and all five are avoidable once you know they exist.

Debugging the control plane when the fault is in the forwarding plane. If show route forwarding-table matches the routing table and both look correct, stop reading routing output. The next thing to check is interface counters, firewall filters and class-of-service, in that order.

Forgetting that a filter can discard silently. A firewall filter with a term that counts and discards produces exactly the symptom of a missing route, and the routing table will look perfect. Always check for filters attached to the interface or to the routing instance before concluding that the packet path is clean — the terms and counters are described in Junos firewall filters: terms, match conditions and actions.

Reading cumulative counters as if they were rates. Every counter on this page is cumulative unless stated otherwise. Sample twice with a known interval and look at the delta.

Assuming both directions behave the same. An interface that transmits fine and receives nothing is a different fault from one that is down in both directions, and most protocols surface these differently.

Leaving traceoptions enabled. It is the single most common source of a second, self-inflicted incident in the same maintenance window.

Quick Reference Table

Symptom First command Second command
No connectivity at all show interfaces terse show chassis alarms
Route missing show route <prefix> show route hidden
Route present, no traffic show route forwarding-table destination <prefix> monitor interface
BGP session down show bgp summary show bgp neighbor <addr>
OSPF adjacency stuck show ospf neighbor extensive show ospf interface detail
VPN up, no data show security ipsec statistics show security flow session tunnel
Unexplained change show system commit show configuration | compare rollback 1

FAQ

Should I pipe everything through | match? Not at first. Match filters are excellent once you know what you are looking for, but they hide context — the line above the match is often the most informative. Use | last, | find and | except alongside | match as appropriate.

Is show configuration | display set better than the stanza view? For troubleshooting, usually yes: it is greppable and it shows inherited values less ambiguously. The stanza view is better for understanding intent and structure, because it keeps apply-groups and interfaces in their hierarchy.

How do I compare this vendor's commands to Cisco? Most Junos concepts have a close Cisco equivalent with a different syntax and a different CLI model — commit-based versus immediate. The mapping in Cisco vs Juniper troubleshooting commands cheat sheet and the broader multi-vendor CLI cheat sheet cover the common pairs.

See also inter-VLAN troubleshooting on EX switches and SRX device upgrade steps.

原文链接:https://ard92.github.io/2022/04/25/junos-troubleshooting.html