SR-IOV VF Passthrough with VFIO on KVM and Proxmox - 夜莺博客

SR-IOV VF Passthrough with VFIO on KVM and Proxmox

Virtual machines usually pay for their networking twice: once in the hypervisor's virtual switch and once in the guest's virtio driver. SR-IOV removes that cost by letting a single physical adapter present itself as many lightweight PCIe devices, each of which can be handed straight to a guest through VFIO. The catch is that SR-IOV is a stack of dependent layers — firmware, kernel, IOMMU, modules and driver — and a miss in any one of them produces a failure that looks like a different problem. This article walks the whole stack with verifiable checks at every step.

The Architecture in One Paragraph

The physical adapter is the Physical Function (PF): it owns the link, the MAC/PHY and the firmware. Each Virtual Function (VF) is a cut-down PCIe function with its own configuration space, MAC address, TX/RX queues and BARs, sharing the PF's physical resources. Because a VF appears as an independent PCI device, the standard VFIO passthrough mechanism can give a guest direct access: the VF's DMA and interrupts are remapped by the IOMMU (Intel VT-d or AMD-Vi) into the guest's address space, and neither the host CPU nor QEMU touches the payload.

Step 1 — Prove the hardware supports It

lspci -nn | grep -i ethernet
lspci -vvv -s 0000:01:00.0 | grep -A 3 "Single Root I/O Virtualization"
# look for: Total VFs, Initial VFs, VF Offset

If the capability block is absent the card cannot do SR-IOV, whatever the driver documentation says. Note Total VFs: that is the firmware ceiling and it varies by card and even by firmware revision (commonly 4–64 on Intel X710/E810 or Broadcom adapters).

Step 2 — Enable IOMMU in firmware and kernel

# BIOS/UEFI: enable VT-d (Intel) or SVM/IOMMU (AMD), plus "SR-IOV Support" / "PCIe ARI"
# where those exist as separate options - some boards also gate it per PCIe root port.

# GRUB hosts:
GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt"
update-grub

# proxmox-boot-tool hosts (ZFS root):
proxmox-boot-tool status
echo 'root=ZFS=rpool/ROOT/pve-1 boot=zfs intel_iommu=on iommu=pt' > /etc/kernel/cmdline
proxmox-boot-tool refresh

# AMD: replace intel_iommu=on with amd_iommu=on
dmesg | grep -i -e DMAR -e IOMMU | head

iommu=pt (pass-through mode) is worth using: devices that are not being passed through run with native DMA performance, and translation is enabled only for devices actually bound to VFIO.

Step 3 — Load the VFIO modules early

printf 'vfio
vfio_iommu_type1
vfio_pci
' >> /etc/modules
update-initramfs -u -k all
reboot
lsmod | grep vfio

Step 4 — Create the virtual functions

cat /sys/bus/pci/devices/0000:01:00.0/sriov_totalvfs      # firmware ceiling
echo 4 > /sys/bus/pci/devices/0000:01:00.0/sriov_numvfs   # create 4 VFs
lspci -nn | grep -i "Virtual Function"

# Verify isolation: each VF should normally sit alone in its IOMMU group
for d in $(lspci -D | grep "Virtual Function" | awk '{print $1}'); do
  echo "$d group=$(basename $(readlink /sys/bus/pci/devices/$d/iommu_group))"
done

To change the VF count later, first write 0, then the new number — the sysfs attribute rejects going directly from one non-zero value to another.

Persist the VF count across reboots

# Proxmox / Debian: hook the PF interface stanza in /etc/network/interfaces
auto ens1f0
iface ens1f0 inet manual
    pre-up echo 4 > /sys/class/net/$IFACE/device/sriov_numvfs

Step 5 — Bind VFs to VFIO and attach them to a guest

modprobe vfio-pci
echo 0000:01:10.0 > /sys/bus/pci/devices/0000:01:10.0/driver/unbind 2>/dev/null
echo 0000:01:10.0 > /sys/bus/pci/drivers/vfio-pci/bind

# Proxmox CLI
qm set 100 -hostpci0 0000:01:10.0,pcie=1

# libvirt XML (equivalent)
# <interface type='hostdev' managed='yes'>
#   <driver name='vfio'/>
#   <source><address type='pci' domain='0x0000' bus='0x01' slot='0x10' function='0x0'/></source>
#   <mac address='52:54:00:11:22:33'/> <vlan><tag id='100'/></vlan>
# </interface>

Inside the guest, install the appropriate VF driver (iavf for Intel, mlx5_core for Mellanox) if the kernel does not bind it automatically, then confirm the adapter appears as a native NIC.

VF Attributes Worth Setting from the Host

ip link set ens1f0 vf 0 mac 52:54:00:11:22:33
ip link set ens1f0 vf 0 vlan 100
ip link set ens1f0 vf 0 trust on          # allows the guest to set its own MAC/VLAN
ip link set ens1f0 vf 0 spoofchk off      # needed by some hypervisor/overlay setups
ip link set ens1f0 vf 0 max_tx_rate 1000  # Mbps rate limit
ip link show ens1f0                        # confirm the VF table

Failure Modes Seen in Production

  • VF creation silently fails. Usually SR-IOV disabled in firmware, or the PF driver does not support the requested count. Check dmesg immediately after the echo.
  • VF not in its own IOMMU group. Passthrough may be rejected or unsafe; move the card to a slot with ACS support, or use kernel ACS override patches only as a lab measure.
  • Handshake failure on the guest NIC. Often trust is off while the guest sets a custom MAC, or spoofchk rejects the guest's frames.
  • No live migration. VF passthrough breaks migration because address translation state cannot be transferred. A common workaround is bonding a virtio/macvtap port with the SR-IOV port inside the guest, migrating on the virtio path, then bringing the VF up.
  • Throughput below expectations. Check that the VF used is on the same NUMA node as the guest's vCPUs — see our DPDK and NUMA memory tuning notes.

Related posts: SR-IOV and Multus in Kubernetes, Proxmox VLAN-aware bridges and LACP bonds and network namespaces and veth pairs.

原文链接:https://kernel-internals.org/virtualization/vfio