Junos Troubleshooting: 10 Essential Operational Commands - 夜莺博客

Junos Troubleshooting: 10 Essential Operational Commands

When a Juniper device starts misbehaving, the fastest path to the root cause is usually a handful of well-chosen operational commands rather than a deep dive into logs. These ten Junos tips, collected from a technical account director who supports enterprise and campus networks daily, cover the scenarios that account for most unexpected issues: power events, virtual chassis problems, cooling and hardware alarms, and interface or LACP faults. All commands work across the EX, QFX and MX families.

Why Ten Commands Cover Most Incidents

Enterprise switching incidents follow a predictable distribution. A large share are physical or environmental — someone bumped a power cable, a fan failed, a rack got hot. A slightly smaller share are topology or redundancy events, where a virtual chassis member dropped out or a LAG lost a member without losing the bundle. A third group are genuine configuration or software faults, and those are the ones where you need logs and traceoptions. The commands below are ordered to match that distribution: you check the cheap, physical things before you spend time on the expensive, logical ones.

That ordering is not arbitrary. Environmental and redundancy faults produce loud, unambiguous signals if you ask for them, and they frequently masquerade as something else — a switch that keeps electing a new master looks like a spanning-tree problem, and a switch whose fan tray has failed looks like random performance degradation. Asking the environment and chassis questions first removes those false trails before they cost you an hour.

Every command below is read-only and safe to run in production. None of them requires a commit, and none of them changes device state. Where a command produces a lot of output on a large chassis, pipe it through | except or | last rather than running it raw.

1. Show System Uptime

show system uptime reveals whether the device recently power-cycled — in stable enterprise networks, unexpected issues are often caused by someone knocking a power cable while working in the rack. The output shows current time, time since boot, time since the last configuration commit, and the list of logged-in users.

show system uptime
show system boot-messages | last 20
show chassis alarms

Read all three lines of the uptime output separately. A device that has been up for 400 days but committed a configuration two minutes ago has not rebooted — someone changed something, and that change is your prime suspect. A device with an uptime of eleven minutes has rebooted regardless of what the ticket says, and show system boot-messages will show whether it came up cleanly.

2. Show Virtual-Chassis Status

show virtual-chassis status verifies that every switch in an EX/QFX Virtual Chassis is present and operating. Any member shown as Not Present needs immediate investigation.

show virtual-chassis status
show virtual-chassis vc-port
show virtual-chassis protocol adjacency

The status output is a table of members with their role, state and neighbour. What matters is the split between control-plane and data-plane state: a member can be present in the status table while its VCP links are down and traffic is no longer traversing the fabric. That is why the second and third commands exist. If show virtual-chassis vc-port shows fewer active VCPs than you designed for, you have lost redundancy even though the member still appears in the list.

A member that flaps in and out of the chassis is worse than one that stays down, because each flap forces an election and a topology re-convergence, and the symptom usually surfaces as intermittent application timeouts rather than as an obvious link alarm. For the configuration side of this, including how members are pre-provisioned so they can rejoin predictably, see Juniper EX Virtual Chassis configuration step by step and Juniper Virtual Chassis explained: roles, members and benefits.

3. Show Log Messages | Last 10

show log messages | last 10 surfaces the most recent syslog entries; pipe through match to filter keywords such as link down or chassis.

show log messages | last 10
show log messages | match "link down|LINK_DOWN"
show log messages | match chassis
show log messages | last 200 | match "error|Error|ERROR"

Two practical refinements. First, the last ten lines tell you what happened most recently, not what caused the incident — if the incident was twenty minutes ago, use | last 200 or drill in by time rather than relying on the default. Second, on a busy device the messages file fills fast, so show log messages.0.gz and the rotated files may hold the relevant window. Junos rotates logs under /var/log and the archived copies are often overlooked.

Match filters on the raw log are much more useful than they look, because Junos log messages are consistently tagged — RT_FLOW for security flow events, UI_COMMIT for configuration commits, CHASSISD for chassis daemon activity. Filtering on the tag rather than on free text gives far cleaner output.

4. Show Chassis Environment

show chassis environment reports temperatures, power and fan state — fans spinning at high speed usually mean insufficient cooling. Read the temperature rows against the thresholds shown in the output rather than against your intuition; a 55°C intake is fine, a 55°C exhaust is not, and the threshold column tells you which is which.

show chassis environment
show chassis environment power
show chassis environment fan
show chassis temperature-thresholds

Power supply state is worth checking independently of the aggregate view: a chassis with two supplies where one has failed will usually keep running, and the only symptom is a minor alarm and an increased thermal load on the surviving supply. The dedicated show chassis environment power view makes that visible immediately.

5. Show Chassis Alarms

show chassis alarms covers hardware alarms (power supply, fans), while show system alarms covers software issues such as an expired UTM license or a missing rescue configuration.

show chassis alarms
show system alarms
show chassis hardware detail

The distinction between the two alarm sources is important and frequently forgotten. Chassis alarms come from the hardware subsystem and generally mean a physical component needs attention. System alarms come from software and generally mean something needs configuring, licensing or committing. Treating a system alarm as hardware chases a fault that does not exist; ignoring a chassis alarm because the device still forwards traffic stores up an outage for the next power event.

6. Show Chassis Hardware Detail

When an alarm does indicate hardware, show chassis hardware detail gives you the serial numbers and part numbers you will need for an RMA, along with the revision levels. Collect it before you open a support case — it is the first thing that will be requested, and on a device that is about to be rebooted it may be the last chance to capture it.

show chassis hardware detail
show chassis hardware detail | match "FPC|PIC"
show chassis fpc

7. Show Interfaces Descriptions | Match Down

show interfaces descriptions | match down
show interfaces terse | match down
show interfaces extensive | match "flap|CRC|error"

If you maintain descriptions on important uplinks, the first command highlights any important interface that is down. This is a case where preventive hygiene pays off: a switch whose interfaces have no descriptions gives you a list of ge-0/0/3 style names and no clue which one matters. Every operational team that has adopted this habit reports the same result — descriptions turn a five-minute triage into a five-second one.

The third command is the follow-up when an interface is up but performing badly. A rising flap counter, CRC errors or input errors on a single member of a bundle points at the physical layer on that member, not at the bundle. The physical checklist in Junos interface flapping: layer 1 troubleshooting covers optics, patch leads and speed negotiation in the order that finds the fault fastest.

8. Show LACP Interfaces

show lacp interfaces
show lacp statistics interfaces ae0
show interfaces ae0 extensive | match "flap"
show lacp interfaces ae0 | match "Current state"

show lacp interfaces identifies redundant links in a LAG bundle that are down when they should be up. The output shows, per member, the actor and partner state columns — what the local switch thinks and what its neighbour thinks. A member showing Current state: Attached on one side and Detached or Collecting/Distributing on the other is a half-formed bundle, and traffic hash distribution across a half-formed bundle is a classic source of intermittent packet loss that looks like an application problem.

Check the LACP statistics counters as well, because a clean state with rising receive errors means the negotiation succeeded but the link underneath is not healthy. On MC-LAG designs the equivalent failure has additional behaviours — see MC-LAG ICCP failure scenarios and LACP system ID and MC-LAG ICCP failure: ARP, PIM and ICL behaviour for what changes during an ICCP outage.

9. Show Ethernet-Switching Table

show ethernet-switching table
show ethernet-switching table interface ge-0/0/1
show ethernet-switching table vlan VLAN10
show ethernet-switching table | match <mac>

In L2 environments, the switching table shows whether the switch has learned MAC addresses on a port and in which VLAN. This answers two very different questions with the same output. First: is the switch seeing this host at all, and on the port you expect? Second: is a single MAC appearing on multiple ports in quick succession, which indicates a loop or a duplicate address?

A MAC that moves between ports repeatedly is a strong loop indicator, and if you see that pattern, check storm control configuration before the broadcast traffic from the loop consumes the uplink — the settings in Junos storm control configuration for BUM traffic are the containment mechanism, not the fix, but they buy you time to find the loop.

10. Show Interfaces in Detail

show interfaces ge-0/0/1
show interfaces ge-0/0/1 extensive | match "Physical|Speed|MTU|flap"
show interfaces ge-0/0/1 | match "rate|packets"
monitor interface ge-0/0/1

Finally, show interfaces returns last flapped timestamps and input/output PPS rates — a healthy output rate with zero input on a WAN link points at the provider. That asymmetry is one of the most useful diagnostics in the whole list: it distinguishes between a link that is broken in both directions and a link where the far side has stopped sending while your side keeps transmitting into the void.

For a single-interface problem, monitor interface is often better than the static counters, because it refreshes every second and shows the current rate rather than a cumulative tally. Watch it during a controlled test — generate traffic from a known source and confirm the counter moves. If it does not, the question becomes where the traffic is being dropped before it reaches this port, and the switching table and VLAN checks above answer that.

Putting the Ten Together

The value of this list is the order, not the individual commands. A sensible sequence for any unexpected incident is: uptime and boot messages (did it reboot?), virtual chassis status (did redundancy change?), recent logs (what did it say?), chassis environment and alarms (is anything physically wrong?), hardware detail if so, then descriptions and LACP (which link changed?), then the switching table and interface detail (what does the data plane actually see?).

That sequence answers the majority of enterprise switching incidents without ever opening a traceoptions file, and when it does not, it narrows the problem enough that traceoptions become targeted rather than a fishing expedition. For a broader command set covering routing protocols and SRX security flows, see the Junos troubleshooting commands reference.

FAQ

Do these commands work on MX routers as well as EX and QFX switches? Yes, with two caveats: command 2 (virtual chassis) and command 9 (ethernet-switching table) are specific to switching platforms and have no meaningful equivalent on an MX, where you would look at show bridge domain and show ethernet-switching alternatives instead.

Which of the ten should I automate? Uptime, chassis alarms and LACP state are the three that benefit most from polled monitoring, because they all produce a clean signal that a monitoring system can alert on without false positives.

What should I add to the list for a data centre fabric? Replace the virtual-chassis checks with EVPN-VXLAN underlay and overlay state, and add the storm-control and MAC-move equivalents for the fabric — the physical-first ordering still applies.

Related Junos articles: troubleshooting IRB VLAN interfaces on EX, inter-VLAN communication on EX switches, and SRX device upgrade steps.

原文链接:https://www.nomios.com/news-blog/ten-tips-junos-troubleshooting/