DPDK Hugepages and NUMA Memory Tuning - 夜莺博客

DPDK Hugepages and NUMA Memory Tuning

Most DPDK performance problems are not driver problems. They are memory-placement problems: a NIC on socket 0 DMA-ing into packet buffers allocated on socket 1 will lose a large fraction of its throughput to cross-socket traffic, and the symptom looks exactly like a NIC firmware issue. This article explains how DPDK organises memory, why hugepages matter far more than "using less TLB", and how to configure the kernel and EAL so that queues, buffers and the polling threads all live on the same NUMA node as the PCIe device.

Why DPDK Needs Its Own Memory Model

DPDK bypasses the kernel network stack, so it also bypasses the kernel's allocator. Instead of malloc(), DPDK reserves hugepage-backed memory at startup and manages it with its own allocator (mempools, heaps and memory segments). Three properties of that design matter operationally:

  • Pre-allocation. All packet buffers are reserved before traffic starts. If the hugepage pool is too small, the application fails at startup rather than degrading later — which is what you want.
  • Pinning. Hugepage memory is locked, never swapped, never migrated. That removes page faults and TLB shootdowns from the fast path.
  • DMA-safe addressing. Physical addresses are known and stable, which is a requirement for IOMMU mappings and for devices using the legacy UIO/VFIO paths.

1 GB hugepages give an additional structural benefit: a single 1 GB page can cover a whole ring of packet buffers without any TLB entry change, so PMD loops do not thrash the TLB under load.

Reserving Hugepages Correctly

Runtime (test and lab, does not survive reboot)

# 8 x 1G pages on NUMA node 0 only
echo 8 > /sys/devices/system/node/node0/hugepages/hugepages-1048576kB/nr_hugepages
# 4096 x 2M pages on node 0
echo 4096 > /sys/devices/system/node/node0/hugepages/hugepages-2048000kB/nr_hugepages

grep -H . /sys/devices/system/node/node*/hugepages/*/nr_hugepages
grep -H . /sys/devices/system/node/node*/hugepages/*/free_hugepages

Persistent (production)

# /etc/default/grub
GRUB_CMDLINE_LINUX="default_hugepagesz=1G hugepagesz=1G hugepages=16   hugepagesz=2M hugepages=8192 transparent_hugepage=never intel_iommu=on iommu=pt"
# then: update-grub && reboot

# Reclaim at runtime instead of rebooting (only works while memory is free)
echo 8 > /sys/devices/system/node/node0/hugepages/hugepages-1048576kB/nr_hugepages

Two rules avoid the classic failures. Reserve hugepages per NUMA node, not globally — a global count can be satisfied entirely on node 1 and then be useless for a NIC on node 0. And reserve them before the workload that fragments memory starts, because 1 GB pages require contiguous physical memory and cannot be allocated late on a busy host.

NUMA Awareness in the EAL

# Show the topology you are actually working with
lscpu | egrep 'NUMA|Socket|Core'
cat /sys/class/net/ens1f0/device/numa_node
cat /sys/class/net/ens1f0/device/local_cpulist

# Bind to socket 0 devices and CPUs only
dpdk-testpmd -l 4-15 --socket-mem 8192,0   -a 0000:3b:00.0 --huge-unlink --   -i --nb-cores=8 --rxq=4 --txq=4

Everything in that line matters:

  • --socket-mem 8192,0 — 8 GB on socket 0, zero on socket 1. The trailing zero is not cosmetic: without it DPDK may place buffers locally to the thread that allocates them.
  • -l 4-15 — CPU cores on socket 0, matching the NIC. Use --lcores with explicit (socket:core) mapping when the host has more than two sockets or you need hyperthread siblings isolated.
  • --huge-unlink — removes the hugepage files after mapping; do not use it if your monitoring relies on /dev/hugepages contents or if a secondary process must attach later.

Keeping Threads and Memory Together

Three layers must agree on the NUMA node: the NIC's PCIe root complex, the PMD threads, and the mempool backing the queues.

# Pin threads and verify where the memory actually landed
taskset -c 4-11 dpdk-l2fwd ...
numastat -p $(pgrep -f dpdk-l2fwd)
cat /proc/$(pgrep -f dpdk-l2fwd)/numa_maps | grep -c huge

numastat -p is the fastest sanity check: a healthy run shows the vast majority of the process's memory allocated on a single node, matching the NIC's numa_node file. If half the memory sits on the other socket, fix the mempool or the thread affinity — no amount of driver tuning will recover that throughput.

Mempool Sizing and Common Failure Modes

# Rule of thumb for a port with 4 RX queues and RX/TX rings of 1024:
#   mbufs = (ring_size * queues * 2) + burst * threads + headroom
# Example: (1024 * 4 * 2) + 32 * 8 + 2048 = 10,496 → round up to 16,384
  • "Not enough memory to allocate mbuf pool". Hugepage count too small, or reserved on the wrong node. Re-check free_hugepages per node before blaming the driver.
  • Throughput collapses at 60–70% of line rate. Usually cross-socket DMA. Verify NIC node, thread affinity and mempool socket id together.
  • Random drops with no interface errors. Mempool exhaustion from too-small a cache or rings; increase --rxd/--txd and the mbuf count, then watch show port stats for rx_mbuf_alloc_failed.
  • Works in the lab, fails in production. Transparent hugepages or a competing VM/container took the contiguous memory first. Reserve large pages at boot.

Verification Checklist

  1. Hugepages reserved per node, matching the NIC's node.
  2. --socket-mem limits allocation to that node.
  3. PMD threads bound to CPUs of the same node, siblings accounted for.
  4. numastat -p shows the process resident on one node.
  5. No rx_mbuf_alloc_failed or rx_nombuf counters under a 30-minute soak test.

Related reading on this site: SR-IOV, Multus and DPDK in Kubernetes, FD.io VPP on Linux and ethtool ring buffer and coalescing tuning.

原文链接:https://www.dpdk.org/memory-in-dpdk-part-1-general-concepts/