FD.io VPP Tutorial: A High-Speed Data Plane on Linux - 夜莺博客

FD.io VPP Tutorial: A High-Speed Data Plane on Linux

VPP is a userspace, layer 2-4 network stack that runs on Linux and processes packets in vectors - batches of up to 256 frames - instead of one at a time. The trick is cache locality: only the first packet in a vector needs to pull the forwarding instructions into the CPU cache, so the remaining 255 execute against instructions that are already hot. The result is a packet forwarding engine that scales far beyond what an in-kernel path can do on the same hardware, and it is the data plane inside a lot of appliances you already own.

This tutorial takes you from a bare Ubuntu host to a working VPP instance with DPDK-bound interfaces, a bridge domain, packet tracing and NAT44.

Install and Prepare the Host

# add the FD.io repository (Ubuntu 22.04 shown)
curl -s https://packagecloud.io/install/repositories/fdio/release/script.deb.sh | sudo bash
sudo apt-get update
sudo apt-get install -y vpp vpp-plugin-core vpp-plugin-dpdk

# hugepages: VPP allocates from here, not from the normal page pool
sudo mkdir -p /dev/hugepages
sudo mount -t hugetlbfs none /dev/hugepages
echo 1024 | sudo tee /proc/sys/vm/nr_hugepages
# persist
echo 'vm.nr_hugepages=1024' | sudo tee /etc/sysctl.d/80-vpp.conf
sudo sysctl --system

Hugepages are the first thing to check when VPP refuses to start or reports allocation failures. A host that has already fragmented its memory may not be able to satisfy the request even when the count looks correct - reboot is often the fastest fix.

Minimal startup.conf

unix {
  nodaemon
  full-coredump
  cli-listen /run/vpp/cli.sock
  log /var/log/vpp/vpp.log
}
api-segment { prefix vpp1 }
cpu {
  main-core 1
  corelist-workers 2-5
}
dpdk {
  dev default {
    num-rx-queues 2
    num-tx-queues 2
  }
}
plugins {
  plugin dpdk_plugin.so { enable }
}

Only one VPP instance on a host can own the DPDK plugin, because it takes a lock on the device. If you want several instances for a lab, disable DPDK in each and connect them with memif or vhost interfaces - the multi-instance pattern is exactly what the official progressive tutorial uses to build a two-node lab from a single VM.

Bind a NIC to DPDK

VPP will not touch an interface the kernel is still managing. Unbind it first.

# find the PCI address
lspci -nn | grep -i ethernet
ip -br link
ethtool -i enp3s0 | head -3

# take it away from the kernel driver
sudo ip link set enp3s0 down
sudo modprobe vfio-pci
sudo dpdk-devbind.py --bind=vfio-pci 0000:03:00.0
sudo dpdk-devbind.py --status

With vfio-pci, IOMMU must be enabled in the BIOS and the kernel command line (typically intel_iommu=on iommu=pt). If the bind fails, that is nearly always the reason.

Bring Up an Interface and Forward

sudo systemctl restart vpp
sudo vppctl

vpp# show version
vpp# show interface
vpp# set interface state GigabitEthernet0/8/0 up
vpp# set interface ip address GigabitEthernet0/8/0 10.10.1.1/24
vpp# set interface state GigabitEthernet0/9/0 up
vpp# set interface ip address GigabitEthernet0/9/0 10.10.2.1/24
vpp# show interface address
vpp# ip route add 10.10.3.0/24 via 10.10.2.2 GigabitEthernet0/9/0

None of that survives a restart unless it is in a startup configuration file. Put interface, address and route statements in /etc/vpp/startup.conf or load them from the CLI socket at boot; ad-hoc CLI typing is fine for a lab and a liability in production.

Bridge Two Interfaces at Layer 2

vpp# create bridge-domain 100
vpp# set interface l2 bridge GigabitEthernet0/8/0 100
vpp# set interface l2 bridge GigabitEthernet0/9/0 100
vpp# show bridge-domain 100 detail
vpp# show l2fib verbose

The L2 FIB is the table that matters once forwarding starts. If traffic disappears, show l2fib tells you whether the source MAC was learned and on which interface - the same question you would ask of a hardware switch.

Trace Packets

This is VPP's best debugging feature and the fastest way to answer "where did the packet go".

vpp# trace add dpdk-input 20
# generate traffic from a peer host, then:
vpp# show trace
------------------- Start of thread 0 vpp_main -------------------
Packet 1
00:00:00:000000: dpdk-input
  GigabitEthernet0/8/0 rx queue 0
  buffer 0x4d2a: current data 0, length 98, ...
00:00:00:000000: ethernet-input
  IP4: 00:50:56:aa:bb:cc -> 00:50:56:dd:ee:ff
00:00:00:000000: ip4-input
  ICMP: 10.10.1.10 -> 10.10.2.10
00:00:00:000000: ip4-lookup
  fib 0 dpo-idx 5 flow hash: 0x00000000
00:00:00:000000: ip4-rewrite
00:00:00:000000: GigabitEthernet0/9/0-output
vpp# clear trace

The node path is the answer: dpdk-input -> ethernet-input -> ip4-input -> ip4-lookup -> ip4-rewrite -> output. If your trace stops after ethernet-input, the frame was not IP and there is no route for it; if it stops at lookup, the FIB has no entry.

NAT44

vpp# set interface nat44 in GigabitEthernet0/8/0 out GigabitEthernet0/9/0
vpp# nat44 add interface address GigabitEthernet0/9/0
vpp# show nat44 interfaces
vpp# show nat44 sessions
vpp# show nat44 summary

Session table growth is the thing to watch here, and it is the same capacity problem that in-kernel NAT has - just at a much higher packet rate. Sizing sessions and timeouts is a design decision, not a default.

Connecting VPP to Containers and the Host

DPDK interfaces cannot be reached from the kernel, so you need an interface that both sides can see:

# host-side tap
vpp# create tap id 0 host-if-name vtap0
vpp# set interface state tap0 up
vpp# set interface ip address tap0 192.168.99.1/24
sudo ip addr add 192.168.99.2/24 dev vtap0 && sudo ip link set vtap0 up

# memif for container-to-VPP, used by CNFs
vpp# create memif id 1 socket /run/vpp/memif.sock
vpp# set interface state memif0/1 up

For Kubernetes deployments where VPP is the CNF data plane, the plumbing around SR-IOV, Multus and DPDK is the part that decides whether it works at all - that setup is covered in this SR-IOV, Multus and DPDK guide. For a comparison with the kernel-resident alternative, this Open vSwitch bridge and VLAN guide shows the equivalent operations in OVS, including its DPDK path.

Tuning Notes

  • Pin workers, isolate cores. Use isolcpus and make sure no other process schedules on the worker cores.
  • Match RX queues to cores. A single queue feeding eight workers creates a cache-coherency bottleneck.
  • Disable unused plugins. Every loaded node that never runs still costs memory and build time.
  • Watch the vector rate, not just throughput. show runtime gives per-node clocks and vectors processed; a node with a rising clock-per-vector is where the regression is.
  • Bigger vectors are better. If average vector size is small, the input path is starving and the cache advantage disappears - look at the NIC's RX queue count and descriptor settings.

原文链接:https://docs.fd.io/vpp/25.06/gettingstarted/progressivevpp/index.html