Open vSwitch DPDK vHost-User and Multiqueue Guide - 夜莺博客

Open vSwitch DPDK vHost-User and Multiqueue Guide

Running Open vSwitch on the DPDK datapath moves packet processing out of the kernel and into user space, which is how you get millions of packets per second to virtual machines without a kernel bridge bottleneck. The trade-off is configuration complexity: hugepages must be reserved, PMD threads pinned, and the vhost-user port type chosen correctly or the VM simply will not have connectivity. This guide covers the working sequence from hugepage reservation through multiqueue validation, using the dpdkvhostuserclient port type that is the only sane choice for production.

Reserve Hugepages and Bind the NIC

DPDK needs 1G or 2G hugepages and a NIC driven by vfio-pci rather than the kernel driver. Reserve the pages at boot and verify them before touching OVS.

# /etc/default/grub
GRUB_CMDLINE_LINUX="default_hugepagesz=1G hugepagesz=1G hugepages=16 iommu=pt intel_iommu=on"
grep HugePages /proc/meminfo
dpdk-devbind.py --status
dpdk-devbind.py --bind=vfio-pci 0000:01:00.0

The number of hugepages should be at least the socket memory OVS is configured to use. Reserving too few produces the classic "Not enough memory to create mempool" failure at startup.

Create the DPDK Datapath Bridge

Bridges that use DPDK must be created with datapath_type=netdev. This is the switch that governs everything else - a missing datapath type creates a kernel bridge and the DPDK ports attach to nothing.

ovs-vsctl add-br br0 -- set bridge br0 datapath_type=netdev
ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-init=true
ovs-vsctl --no-wait set Open_vSwitch . other_config:dpdk-socket-mem="4096,2048"
ovs-vsctl --no-wait set Open_vSwitch . other_config:pmd-cpu-mask=0xc
systemctl restart openvswitch-switch
ovs-vsctl get Open_vSwitch . other_config

pmd-cpu-mask pins poll-mode driver threads to specific cores. Cores listed here must not be used by the kernel; isolate them with isolcpus on the kernel command line for best results.

Add Physical Ports and vHost-User Ports

Two port types exist for virtio: dpdkvhostuser, where OVS is the server and QEMU the client, and dpdkvhostuserclient, where OVS is the client and QEMU owns the socket. The client variant lets you restart OVS without restarting every VM, so it is the recommended type.

ovs-vsctl add-port br0 dpdk0 -- set Interface dpdk0 type=dpdk   options:dpdk-devargs=0000:01:00.0
ovs-vsctl add-port br0 vhu1 -- set Interface vhu1 type=dpdkvhostuserclient   options:vhost-server-path=/var/run/openvswitch/vhu1.sock
ovs-vsctl show

Attach the socket to the guest with QEMU's netdev vhost-user syntax, and remember that the socket path must exist and be writable by the user running QEMU:

-chardev socket,id=char1,path=/var/run/openvswitch/vhu1.sock -netdev type=vhost-user,id=net1,chardev=char1,vhostforce,queues=4 -device virtio-net-pci,netdev=net1,mq=yes,vectors=10

Multiqueue: Sizing Queues Correctly

For multiqueue to actually spread traffic, at least two PMD threads must be configured. With a single PMD all vhost queues are served by the same thread and throughput does not scale. The number of RX queues on the physical DPDK port should also be at least two so different PMDs handle different ingress queues.

ovs-vsctl set interface dpdk0 options:n_rxq=4
ovs-vsctl set Open_vSwitch . other_config:n-dpdk-rxqs=4
ovs-appctl dpif-netdev/pmd-rxq-show
ovs-appctl dpif-netdev/pmd-stats-show

rxq-show is the definitive check: every queue should be assigned to a different PMD core. If they all land on PMD 0, the PMD mask covers one core only.

Verifying Traffic and Troubleshooting

ovs-vsctl list interface dpdk0 | grep -E "link_state|mtu|n_rxq"
ovs-appctl dpctl/show
ovs-appctl dpif/dump-flows -m br0
ovs-appctl dpif-netdev/dpctl-show
ovs-appctl dpif-netdev/pmd-perf-show

Zero packets on a vhost port usually means the socket never connected - check that the guest is running and that the path matches exactly. For the higher-level bridge and VLAN configuration on a kernel datapath, see Open vSwitch ovs-vsctl bridge and VLAN configuration; for the memory and NUMA side, DPDK hugepages and NUMA memory tuning goes deeper.

原文链接:https://docs.openvswitch.org/en/stable/topics/dpdk/vhost-user/