Huawei NQA: Test Instances for Path and SLA Monitoring - 夜莺博客

Huawei NQA: Test Instances for Path and SLA Monitoring

NQA - Network Quality Analysis - is Huawei's probe framework on VRP and on CloudEngine switches. It is the equivalent of IP SLA on Cisco and RPM on Junos: it generates synthetic traffic to a target on a schedule, measures round-trip time, jitter and packet loss, and can react to a threshold crossing by sending a trap, writing a log, or changing the state of a tracked route. Where BFD gives you fast failure detection, NQA gives you measurement over time - and the two solve genuinely different problems.

This article covers the test-instance model, working configurations for the common test types, thresholding and reaction, and the details that decide whether the results are trustworthy.

The Test Instance Model

An NQA test is identified by an administrator name and a test (operation) name, and it has a type. The naming is free-form, so use something that describes the path - you will be reading it in display output and in trap messages for years.

system-view
nqa test-instance ADMIN WAN-TO-BRANCH01
 nqa type icmp
 destination-address ipv4 203.0.113.45
 source-address ipv4 198.51.100.1
 probe-count 10
 interval seconds 1
 frequency 60
 timeout 3
 records history 50
 description SLA monitor for branch 01 WAN link
 commit
quit
Statement Meaning Guidance
probe-count Probes per test cycle 10 gives a useful sample without flooding
interval seconds Spacing between probes 1 s default; shorter costs CPU and adds traffic
frequency Seconds between test cycles 60 s is a sound default
timeout Seconds before a probe is treated as lost Must exceed the expected RTT; too short inflates loss
records history Results retained per test Enough to see a trend, not so many that memory matters

Note the commit. On CloudEngine and newer VRP versions configuration changes are staged until committed, and an uncommitted NQA test simply does not run. If a test appears to do nothing, check that first.

Test Types

Type Measures Statement
icmp RTT, jitter, loss destination-address ipv4 ...
tcp Time to establish a TCP session destination-port 443
udp RTT to a UDP port destination-port 53
http Time to fetch a URL destination-address plus a filename/URL statement
trace Hop-by-hop path Path discovery and change detection
jitter UDP jitter, loss, one-way metrics Needs the peer as an NQA server

TCP connect test against a real service

system-view
nqa test-instance ADMIN ERP-HTTPS
 nqa type tcp
 destination-address ipv4 10.60.20.15
 destination-port 443
 source-address ipv4 10.60.20.1
 probe-count 5
 interval seconds 2
 frequency 120
 timeout 5
 records history 30
 commit
quit

Testing the real service port is usually more informative than ICMP. It measures whether the application is reachable, not just whether the host answers pings - and a device that permits ICMP but blocks 443 is a common configuration.

UDP jitter test between two devices

! on the responder
system-view
nqa-server udpecho 10.60.20.1 9020
quit

! on the initiator
system-view
nqa test-instance ADMIN WAN-JITTER
 nqa type jitter
 destination-address ipv4 10.70.20.1
 destination-port 9020
 source-address ipv4 10.60.20.1
 probe-count 20
 interval milliseconds 20
 frequency 60
 commit
quit

The jitter test is the one that matters for voice and video - it produces packet loss, minimum and maximum delay, and jitter for the path, which is far more useful than an average RTT when the complaint is "calls sound bad".

! the other side must be configured as an NQA server for the relevant port and address
nqa-server udpecho 10.70.20.1 9020

Thresholds and Reaction

A measurement nobody acts on is a wasted probe. Configure a threshold and a reaction:

nqa test-instance ADMIN WAN-TO-BRANCH01
 threshold rtd 100
 threshold rtd 100 action trap
 threshold rtd 100 action log
 threshold successive-loss 3 action trap
 threshold rtd 300 action trap
 commit

Thresholds are in milliseconds for RTT-related metrics, and reactions can include sending traps, writing logs, or triggering a state change that a tracked route or VRRP group observes. The reaction types worth setting on day one:

  • RTT threshold at the SLA boundary - the number in your contract.
  • RTT threshold at a worse value that indicates degradation rather than the edge.
  • Successive loss - consecutive lost probes, which distinguishes a real path failure from random loss.
  • Log reaction in addition to trap, so there is local evidence when the SNMP path is also affected.

Verification

display nqa results
display nqa results test-instance ADMIN WAN-TO-BRANCH01
display nqa history test-instance ADMIN WAN-TO-BRANCH01
display nqa statistics test-instance ADMIN WAN-TO-BRANCH01
display nqa server status

! a typical result block
NQA entry(admin, WAN-TO-BRANCH01) test result:
  Destination IP address: 203.0.113.45
  Send operation times: 10        Receive response times: 10
  Completion: success             RTD OverThresholds number: 0
  Min/Max/Average Completion Time: 17/24/19 ms
  Sum/Square-Sum Completion Time: 190/3672
  Lost packet ratio: 0 %

Read min, max and average together. An average of 19 ms with a maximum of 24 ms is a stable path. The same average with a maximum of 300 ms is a path with periodic stalling that users will notice, and it is invisible if you only chart the average.

Using NQA to Detect Path Failure

The classic use is to couple a probe to a route or a first-hop protocol, so a monitored path is withdrawn when it degrades:

nqa test-instance ADMIN WAN-PRIMARY
 nqa type icmp
 destination-address ipv4 203.0.113.45
 frequency 10
 probe-count 3
 interval milliseconds 500
 commit
quit

! track the test and bind it to the static route
ip route-static 0.0.0.0 0 198.51.100.2 track nqa ADMIN WAN-PRIMARY
! or to a VRRP group's priority
interface Vlanif100
 vrrp vrid 10 track nqa ADMIN WAN-PRIMARY reduced 30

This is where NQA earns its place: it turns a measured path quality into a routing decision, without waiting for a physical link to fail. The static-route side of that design is covered in this Huawei static route guide, and the first-hop protocol side in this Huawei VRRP master/backup guide.

Pitfalls

  • Source address matters. Without a source address, probes can leave through the wrong interface and measure a path you did not intend to test. Set it explicitly.
  • Uncommitted configuration. On CloudEngine and newer VRP, an uncommitted NQA test does nothing. It is the first thing to check when a test never produces results.
  • Timeout too short. A timeout shorter than the real path RTT reports 100% loss on a perfectly healthy link. Set it from measurement, not from hope.
  • Aggressive frequencies across many targets. A 10-second frequency against fifty destinations is a lot of control-plane work. Scale the frequency to the number of tests.
  • Reaction loops. A reaction that withdraws a route which causes the probe to fail which withdraws another route is a design error. Track one thing at a time and dampen the reaction.
  • Clock accuracy for one-way metrics. Jitter tests that report one-way delay need both ends synchronised. Verify time synchronisation before believing one-way numbers.
  • Overlap with BFD. If you already run BFD on the same path for fast detection, do not duplicate the failure reaction in NQA. BFD detects, NQA measures - keep the roles separate.

When the results are collected centrally, the same trap and result data feeds an NMS - and the SNMP side of that integration, including the common reasons traps do not arrive, is covered in this Huawei SNMP and NMS troubleshooting guide. If the monitored path is a LAG, remember that a probe measures the path the hash selects, not every member - the aggregation configuration itself is in this Huawei Eth-Trunk and LACP guide.

原文链接:https://support.huawei.com/enterprise/en/doc/EDOC1100034074/2e80aedf/nqa-configuration