Junos RPM Probes: SLA Monitoring with Ping, HTTP and UDP - 夜莺博客

Junos RPM Probes: SLA Monitoring with Ping, HTTP and UDP

Junos has its own answer to IP SLA and NQA, and it is called Real-Time Performance Monitoring - RPM. It sends synthetic probes to a target and computes round-trip time, jitter and packet loss over a defined test window, with thresholds that can raise SNMP traps. The reason to use it rather than a management-station ping is that RPM measures from the router's own forwarding path, out of the interface you choose, at a defined interval - which is what you need when you are verifying an SLA on a specific WAN link or proving that one member of a LAG is worse than the others.

This article covers the configuration model, a working ICMP and HTTP probe, thresholding, and the details that make results trustworthy.

The Configuration Model

RPM is organised as an owner with one or more tests under it. The owner name plus the test name is the unique instance, so choose names that describe the path, not the interface index - you will be reading them in trap messages and graphs for years.

# Junos - owner "WAN-SLA", test "to-branch-01"
set services rpm probe WAN-SLA test to-branch-01 probe-type icmp-ping
set services rpm probe WAN-SLA test to-branch-01 target address 203.0.113.45
set services rpm probe WAN-SLA test to-branch-01 source-address 198.51.100.1
set services rpm probe WAN-SLA test to-branch-01 probe-count 10
set services rpm probe WAN-SLA test to-branch-01 probe-interval 1
set services rpm probe WAN-SLA test to-branch-01 test-interval 60
set services rpm probe WAN-SLA test to-branch-01 history-size 50
set services rpm probe WAN-SLA test to-branch-01 data-size 64
set services rpm probe WAN-SLA test to-branch-01 dscp-code-point 001110
set services rpm probe WAN-SLA test to-branch-01 traps rtt-jitter

The semantics that matter:

  • probe-count × probe-interval is the measurement window. Ten probes one second apart gives you a ten-sample view of a ten-second period.
  • test-interval is how often the whole test repeats. Sixty seconds is a reasonable default; shorter intervals cost CPU and add traffic.
  • history-size keeps the most recent results for trend analysis and graphing.
  • data-size changes the packet size, which is how you test for MTU-sensitive paths - a probe that succeeds at 64 bytes and fails at 1400 is telling you something specific.
  • dscp-code-point lets you mark the probes so they traverse the same queue as the traffic they represent. A probe that is not marked like production traffic measures the best-case path, not the production path.

Probe Types

Probe type Measures Notes
icmp-ping RTT, jitter, loss The default; works against anything that answers ICMP
icmp-ping-timestamp One-way delay estimates Requires a cooperating responder; enables ingress/egress breakdown
udp-ping RTT to a UDP port Use when ICMP is filtered but you can pick a port
udp-ping-timestamp One-way delay and jitter Needs a hardware-timestamp-capable responder; the most accurate option
tcp-connect Time to establish a TCP session The right choice for application-path verification
http-get Time to fetch a URL Not available for BGP RPM; useful for proxy and WAN optimisation paths

A TCP connect test against the real service port is often more honest than ICMP: it measures reachability of the thing you actually care about, on the port you actually use.

set services rpm probe APP-SLA test to-erp-https probe-type tcp-connect
set services rpm probe APP-SLA test to-erp-https target url https://erp.example.com
set services rpm probe APP-SLA test to-erp-https probe-count 5
set services rpm probe APP-SLA test to-erp-https probe-interval 2
set services rpm probe APP-SLA test to-erp-https test-interval 120
set services rpm probe APP-SLA test to-erp-https history-size 30
set services rpm probe APP-SLA test to-erp-https destination-interface ge-0/0/0.0
set services rpm probe APP-SLA test to-erp-https dscp-code-point af21

Thresholds and Traps

RPM is only useful as monitoring if it can tell you when something changed. Set thresholds on the metrics you care about and let the failure trigger an alarm rather than a dashboard nobody watches:

set services rpm probe WAN-SLA test to-branch-01 thresholds rtt 40000
set services rpm probe WAN-SLA test to-branch-01 thresholds jitter-rtt 5000
set services rpm probe WAN-SLA test to-branch-01 thresholds std-dev-rtt 10000
set services rpm probe WAN-SLA test to-branch-01 thresholds successive-loss 3
set services rpm probe WAN-SLA test to-branch-01 thresholds total-loss 5
set services rpm probe WAN-SLA test to-branch-01 traps rtt-jitter
set services rpm probe WAN-SLA test to-branch-01 traps probe-failure

Values are in microseconds for the time metrics. Four thresholds worth thinking about:

  • rtt - the maximum round-trip time permitted; crossing it means the path is outside its SLA.
  • jitter-rtt - the maximum jitter, which is the number that actually breaks voice and video.
  • std-dev-rtt - variability across the test window; a low average with high deviation is a worse user experience than a slightly higher average with low deviation.
  • successive-loss / total-loss - consecutive and total probe loss. A single lost probe is noise; three consecutive losses is a path problem.

Verification

show services rpm probe-results
show services rpm probe-results owner WAN-SLA
show services rpm history-results
show services rpm probe-results | match "Probe Type|Target|Round Trip|Jitter|Loss"

# example output shape
Owner: WAN-SLA, Test: to-branch-01
    Probe Type: icmp-ping, Target address: 203.0.113.45
    Probe results:
      Response received, Fri Oct  1 09:15:02 2026
        Round trip time: 18432 usec
        Jitter:           2211 usec
        Std deviation:    3980 usec
        Probe loss:       0
    Results over current test:
      Probes sent: 10, Probes received: 10, Loss: 0 percent
      Round trip time: avg 19021, min 17880, max 21440 usec
      Jitter:          avg  2604, min  1100, max  4410 usec

Read the results as a set, not a single number. Average RTT that looks fine combined with a rising maximum and rising jitter is the signature of a link approaching saturation, and it appears long before loss does.

Instrumenting a Path Properly

Two techniques make RPM results meaningful:

  1. Measure from the interface under test. Use source-address or destination-interface so the probe leaves through the link you are evaluating. A probe that follows the default route measures whatever path the router would choose, which may be the backup.
  2. Match the forwarding treatment. Set the DSCP so the probe shares the queue with the traffic it represents. To go further, route the probe through a specific queue with a firewall filter and an interface configured for hardware timestamping - this removes the router's own queuing delay from the measurement and gives you a much cleaner view of the network.
set firewall filter RPM-OUT term probes from dscp af21
set firewall filter RPM-OUT term probes then forwarding-class assured-forwarding
set interfaces ge-0/0/0 unit 0 family inet filter output RPM-OUT

Where RPM Fits (and Where It Does Not)

  • SLA verification and reporting. RPM is the right tool: defined interval, defined probes, thresholds, history, traps.
  • Fast failure detection for routing. Use BFD instead. RPM's interval granularity is for measurement, not sub-second reconvergence - the timer design and interaction with protocol hold-downs are covered in this Junos BFD timers guide.
  • Continuous full-mesh mesh monitoring. Do not build an N-squared probe mesh on the routers. Use a management platform, or a streaming-telemetry pipeline, and keep RPM for the paths with an actual SLA.

Pitfalls

  • Time synchronisation. One-way delay and jitter metrics require the requester and responder to share an accurate clock. If they do not, the numbers are meaningless rather than merely imprecise. Ensure both ends are synchronised before trusting one-way figures, and be aware of the accuracy difference between the two approaches described in this PTP vs NTP comparison.
  • Probes competing with production. A probe every second with a large data size on a congested low-speed link adds load. Size the test to the link.
  • Jitter needs timestamp support. If the responder does not support hardware timestamps, RPM can report round-trip measurements but cannot compute round-trip jitter. Check before promising a jitter SLA.
  • Platform differences. Probe types and threshold statements vary between the EX, MX and SRX families, and some features are release-dependent. Verify the specific statement against the documentation for your platform and release rather than assuming a configuration that works on one will load on another.
  • Treating a passing test as a healthy service. RPM proves the network path. It does not prove the application answered correctly. A TCP connect test that succeeds while the application returns errors still passes.

Used for what it is good at - measuring a defined path on a defined schedule and raising an alarm when the SLA is missed - RPM replaces a surprising amount of ad-hoc pinging and gives you the history to argue about a carrier's performance with evidence. When the symptom is application-level rather than path-level, the session-oriented debugging workflow in this SRX flow session debugging guide is the next place to look.

原文链接:https://juniper.net/documentation/en_US/junos/topics/usage-guidelines/services-configuring-real-time-performance-monitoring.html