Linux IRQ Affinity, RSS, RPS and XPS Tuning - 夜莺博客

Linux IRQ Affinity, RSS, RPS and XPS Tuning

At 10G and above, a single saturated CPU handling interrupts is the usual ceiling on throughput. The tools that fix it are layered: RSS spreads flows across hardware queues in the NIC, IRQ affinity decides which CPUs handle those queues, RPS pushes protocol processing to another CPU entirely, and XPS does the same for transmit. Each layer solves a different problem, and applying them in the wrong order is how people make throughput worse.

RSS: hardware flow distribution

Multi-queue NICs hash each packet into one of N receive queues, and each queue has its own interrupt. The mapping is visible in /proc/interrupts and the hashing can usually be inspected and adjusted:

ethtool -l eth0                # max / current combined queue counts
sudo ethtool -L eth0 combined 8

ethtool -x eth0                # current RSS indirection table
ethtool -n eth0                # n-tuple / flow director filters

grep eth0 /proc/interrupts
! 57:  120931  0  0  0  0  0  0  0  PCI-MSI  eth0-TxRx-0
! 58:  0  118422 0  0  0  0  0  0  PCI-MSI  eth0-TxRx-1

For low latency, use as many queues as CPUs (or the NIC maximum). For maximum throughput, the efficient configuration is usually the smallest number of queues where no queue saturates a CPU, because with interrupt coalescing each additional queue also adds work.

IRQ SMP affinity

cat /proc/irq/57/smp_affinity_list        # e.g. "0"
echo 4 > /proc/irq/57/smp_affinity        # hex bitmask -> CPU 2
echo 2 > /proc/irq/57/smp_affinity_list   # CPU numbering is 1-based here

systemctl status irqbalance               # may override your settings on restart

Two traps. First, smp_affinity is a hexadecimal bitmask while smp_affinity_list is a CPU list — a careless write of a decimal value into the bitmask file pins the queue to the wrong CPU. Second, irqbalance will happily undo manual placement; either stop it or configure its banned-CPU lists, otherwise your tuning disappears at the next daemon restart.

RPS, RFS and XPS

# RPS: which CPUs do protocol processing for a receive queue
echo 0f > /sys/class/net/eth0/queues/rx-0/rps_cpus

# RFS: keep flows on the CPU running the application thread
echo 32768 > /proc/sys/net/core/rps_sock_flow_entries
echo 4096  > /sys/class/net/eth0/queues/rx-0/rps_flow_cnt

# XPS: which queues a CPU may transmit on
echo 4 > /sys/class/net/eth0/queues/tx-0/xps_cpus

RPS is logically a software implementation of RSS: it is invoked later in the datapath and hands the packet to another CPU's backlog queue. On a system where every hardware queue already maps to its own CPU, RPS is redundant. Where there are fewer hardware queues than CPUs — common in virtual machines — RPS is often the difference between one CPU on fire and the work spread across four.

RFS (Receive Flow Steering) improves on RPS by steering packets to the CPU where the consuming application thread actually runs, which raises cache locality. It requires rps_sock_flow_entries to be set globally and a per-queue rps_flow_cnt; the classic failure is setting only one of the two and seeing no effect.

XPS does for transmit what RPS does for receive, and the most efficient configuration is a 1:1 CPU-to-queue mapping when queues equal CPUs, so no two CPUs contend for one transmit queue.

NUMA and locality

cat /sys/class/net/eth0/device/numa_node
numactl --hardware
taskset -c 6,7 ./server            # pin the application
echo 40 > /proc/irq/60/smp_affinity   # pin IRQs to CPUs on the same node

Cross-socket memory access costs more than any of the micro-optimisations above. Keep the NIC's IRQs, the application threads and the memory they use on the same NUMA node, and verify with numastat -p <pid> that you achieved it.

Verifying the result

mpstat -P ALL 2            # watch %soft on each CPU
cat /proc/softirqs | grep -E "CPU|NET_RX"
perf top -C 2              # what the hot CPU is spending time in
nstat -az | grep -iE "softnet|backlog"

Healthy behaviour looks like %soft spread roughly evenly across the CPUs you pinned, rather than one CPU at 100 % softirq and the rest idle. If evenly pinned CPUs are still saturated, the bottleneck has moved up the stack — check the application and socket buffers before adding more queues.

Related: softnet_stat and netdev_max_backlog tuning for the backlog that RPS writes into, and DPDK hugepages and NUMA tuning if you are considering skipping the kernel stack entirely.

原文链接:https://docs.kernel.org/networking/scaling.html