iPerf3 Throughput Testing: A Repeatable Methodology - 夜莺博客

iPerf3 Throughput Testing: A Repeatable Methodology

Throughput numbers that cannot be reproduced are worse than no numbers, because they end up in a change ticket as evidence. iPerf3 is a small tool with a handful of options that decide whether the result measures the network or the test host, and the difference between "we got 4 Gb/s on a 10G link" and a genuine finding usually comes down to three flags.

Start the server and size the window

# server
iperf3 -s -D
iperf3 -s -p 5203

# client: 30 seconds, 1 second reports
iperf3 -c 10.10.10.20 -i 1 -t 30

TCP throughput is bounded by the window in flight divided by round-trip time. On any link with meaningful latency, the default socket buffer will cap the result long before the link is saturated, so the window has to be set explicitly, and it is negotiated for both sides:

iperf3 -c 10.10.10.20 -w 32M -P 4 -t 30 -O 2

Four things are happening in that command. -w 32M sets a 32 MB window. -P 4 runs four parallel streams, which matters twice over: a single flow hashes to a single member of a port channel, and even on a single link a single TCP connection may not fill it. -O 2 omits the first two seconds so the measurement excludes TCP slow start. -i 1 gives per-second reports so you can see whether the rate is stable or collapsing.

Direction matters

iperf3 -c 10.10.10.20 -R -t 30 -i 1

Reverse mode has the server send and the client receive. Run both directions every time: asymmetric results are extremely common at 10G and above and point at a specific host's PCIe or interrupt handling rather than at the network. A link that does 9.4 Gb/s forward and 2 Gb/s reverse is not a "10G link" in any useful sense.

UDP tests: loss and jitter, not throughput

iperf3 -c 10.10.10.20 -u -b 200M -t 30
iperf3 -c 10.10.10.20 -u -b 0 -t 10     # no rate limit

UDP results report datagrams sent, received, lost, and jitter. Use them to characterise a path for real-time traffic and to find the point at which a policer or shaper starts dropping, but do not report UDP loss as a link fault without checking whether the receiver was simply overloaded.

Host-side effects to eliminate

  • CPU affinity. -A 2,3 pins the sender and receiver to specific cores. On a busy host, letting the kernel schedule the test across cores produces variance that is not a network property.
  • Zero copy. -Z uses sendfile(), reducing PCIe and CPU overhead. If throughput jumps when -Z is added, the previous result described the host, not the link.
  • Multiple address pairs. One flow hashes to one path. To measure the capacity of a bundle or a fabric, run concurrent tests using different source and destination addresses and ports, then add the results.
  • Interrupt and ring settings. A petabyte of tuning lore lives here; on Linux the practical starting point is checking drop counters and ring buffers as described in ethtool ring buffer and coalescing tuning.
  • JSON output for the record. -J produces a complete result set that can be filed with a change ticket, which is what makes the test repeatable by the next engineer.
# a defensible standard test - file the JSON with the ticket
iperf3 -c 10.10.10.20 -w 8M -P 4 -t 30 -O 2 -i 1 -J > test-$(date +%F).json

Interpreting results

A clean 10G result is typically north of 9.4 Gb/s with zero retransmits. A result between 5 and 9 Gb/s usually indicates the window or the host. Retransmits in the iperf3 output alongside full bandwidth suggest a marginal physical link, a duplex mismatch or a congested path, and those cases belong to latency and loss monitoring and interface error counters rather than to more iperf runs. For continuously available synthetic probing, blackbox_exporter probes cover the ongoing case that one-off tests cannot.

原文链接:https://docs.nvidia.com/networking-ethernet-software/knowledge-base/Configuration-and-Usage/Monitoring/Throughput-Testing-and-Troubleshooting