Linux softnet_stat, RPS and netdev_max_backlog Tuning - 夜莺博客

Linux softnet_stat, RPS and netdev_max_backlog Tuning

Interface counters can be perfectly clean while a Linux host silently drops packets: no CRC errors, no discards visible in ifconfig, and yet netstat -s shows receive buffer errors climbing. The reason is that the drop happens above the driver, in the softnet backlog — and the counters for it live in a file most engineers have never opened. This article explains the receive path, how to read /proc/net/softnet_stat, and which tunables actually fix per-CPU drops as opposed to merely making the numbers look better.

The Receive Path, Briefly

  1. The NIC receives a frame and, depending on RSS, selects a hardware queue and an IRQ (thus a CPU).
  2. The driver's NAPI poll moves the packet into the kernel and calls netif_receive_skb().
  3. If RPS is enabled, the packet is pushed onto the backlog queue of a chosen CPU; otherwise it is processed on the receiving CPU immediately.
  4. Protocol processing happens, then the socket is woken.

Step 3 is where backlog drops appear: if the target CPU's input_pkt_queue exceeds net.core.netdev_max_backlog, the packet is dropped and the per-CPU dropped counter increments. No NIC counter moves, no interface error appears — which is exactly why this class of loss is misdiagnosed as an application problem.

Reading /proc/net/softnet_stat

# Columns (hex): processed, dropped, time_squeeze, cpu_collision,
#                 received_rps, flow_limit_count
awk '{printf "cpu%d processed=%u dropped=%u time_squeeze=%u rps=%u flow_limit=%u
",
      NR-1, strtonum("0x"$1), strtonum("0x"$2), strtonum("0x"$3),
      strtonum("0x"$5), strtonum("0x"$6)}' /proc/net/softnet_stat

The interpretations that matter:

  • dropped climbing steadily — the backlog queue is too small for the packet rate. Raise netdev_max_backlog (start at 3000, 10000 for high-rate hosts) and re-measure.
  • time_squeeze climbing — the NAPI poll ran out of its budget while the CPU was busy elsewhere. This is a CPU capacity/interrupt-distribution problem, not a queue-size problem; fix with RPS, more queues, or fewer competing workloads.
  • flow_limit_count climbing — large flows are being preempted in favour of small ones. That is the feature working (see below), not a fault.
  • Everything zero but drops are still reported — the loss is above the backlog (socket receive buffer, or the application not reading fast enough). Check nstat -az | grep -i drop and ss -tm for per-socket queue sizes.

Enable RPS/RFS Where RSS Is Not Enough

# RPS: distribute protocol processing for queue 0 of ens1f0
# CPU mask: CPUs 0-3 => 0x0f
echo f > /sys/class/net/ens1f0/queues/rx-0/rps_cpus
echo f > /sys/class/net/ens1f0/queues/rx-1/rps_cpus

# RFS: steer packets to the CPU running the consuming application
sysctl -w net.core.rps_sock_flow_entries=65536     # 1M+ on very busy servers
echo 32768 > /sys/class/net/ens1f0/queues/rx-0/rps_flow_cnt

# Verify
cat /sys/class/net/ens1f0/queues/rx-0/rps_cpus
grep -i rps /proc/interrupts | head

RPS is a software implementation of RSS and runs later in the path: RSS chooses the hardware queue and IRQ CPU, RPS chooses the CPU that does protocol processing above the interrupt handler. On a virtual machine with a single virtio queue, or on a NIC whose queue count is smaller than the CPU count, RPS is the difference between one saturated core and a well-distributed host. RFS goes one step further and uses the socket flow table to send packets to the CPU where the application thread is actually running, which improves data-cache hit rates — relevant for busy database and proxy servers.

Flow Limiting: Protecting Small Flows From Big Ones

# Enable flow limiting on rx-0 (same bitmap format as rps_cpus)
echo f > /sys/class/net/ens1f0/queues/rx-0/rps_flow_cnt    # RFS entries first
cat /sys/class/net/ens1f0/queues/rx-0/rps_flow_cnt

# Flow limit bitmap is per-device in older kernels; check what your kernel exposes
find /sys/class/net/ens1f0 -name 'flow_limit*' 2>/dev/null
ls /sys/class/net/ens1f0/queues/rx-0/

Once a CPU's input queue exceeds half of netdev_max_backlog, the kernel starts counting packets per flow over the last 256 packets; a flow that exceeds the configured ratio has its next packet dropped early. Small flows continue to be accepted. This is a deliberate, bounded form of unfairness that stops one bulk transfer from starving a thousand interactive sessions — worth enabling on web proxies and API gateways, unnecessary on pure throughput hosts.

Sysctls Worth Setting Together

cat >/etc/sysctl.d/99-netrx.conf <<'EOF'
net.core.netdev_max_backlog = 10000
net.core.somaxconn = 4096
net.core.rmem_max = 33554432
net.core.netdev_budget = 600
net.core.netdev_budget_usecs = 8000
EOF
sysctl --system

netdev_budget and netdev_budget_usecs cap how much work one poll cycle performs before yielding; raising them reduces time_squeeze at the cost of latency for other work on that CPU. Increase them deliberately and measure both throughput and tail latency.

Diagnostic Workflow

  1. nstat -az | egrep 'TcpExtListenDrops|TcpExtTCPBacklogDrop|UdpInErrors' to classify the loss.
  2. awk the softnet_stat columns (above) to see which CPU and which counter is climbing.
  3. If dropped: raise netdev_max_backlog; if time_squeeze: add RPS/RFS or increase queue count.
  4. If neither moves: the loss is in the socket buffer — check ss -tm and raise rmem_max plus the application's read buffer.
  5. Re-measure under the same load; a tuning change without a before/after measurement is a guess.

Related reading: ethtool ring buffers and interrupt coalescing, nf_conntrack table full and drops and DPDK hugepages and NUMA tuning.

原文链接:https://docs.kernel.org/networking/scaling.html