Proxmox VE SDN: Zones, VNets and VXLAN Overlays - 夜莺博客

Proxmox VE SDN: Zones, VNets and VXLAN Overlays

Managing dozens of bridges by hand across a Proxmox cluster is where configuration drift starts: one node gets the VLAN tag, another does not, and a guest loses connectivity after a migration nobody planned. Proxmox VE's built-in SDN feature replaces per-node bridge edits with cluster-wide objects — zones, VNets and subnets — that are compiled into local network configuration on every node. This guide covers the zone types, the CLI workflow, and the verification steps that confirm the overlay is actually up.

Zones, VNets and Subnets

  • Zone — a virtually separated network area with a technology: vlan, qinq, vxlan, simple (isolated/NAT bridge) or evpn.
  • VNet — a virtual network that belongs to a zone; deployed locally on each node and appearing to guests as a normal bridge.
  • Subnet — an IP network on a VNet, used for the built-in DHCP/IPAM in L3-capable zones.

All SDN state lives in /etc/pve/sdn (pmxcfs, replicated to every node) and is compiled into /etc/network/interfaces.d/sdn when you apply it. Zones can be restricted to specific nodes and assigned permissions, which is how you give a tenant access to one VNet without exposing the rest of the fabric.

Zone Types and When to Use Them

Type Layer Cross-node? Typical use
simple L2 isolated bridge No Test VMs, no physical uplink
vlan L2 tagged on existing bridge Via the physical fabric Standard segmentation on a VLAN-enabled uplink
qinq L2 stacked 802.1ad Via the fabric Tenant separation with an S-TAG
vxlan L2 overlay Yes, over L3 Stretched L2 where adding VLANs to the switch is impractical
evpn L2/L3 overlay with BGP control plane Yes Datacentre fabrics, inter-VNet routing, exit nodes

Prerequisites

pveversion                       # any supported 7.x/8.x/9.x
# ifupdown2 must manage networking
dpkg -l | grep ifupdown2
# the SDN include must exist in /etc/network/interfaces
tail -3 /etc/network/interfaces      # expect: source /etc/network/interfaces.d/*

Building a VXLAN Overlay with pvesh

# 1. Create the zone, listing every node IP that participates
pvesh create /cluster/sdn/zones --zone vxzone1 --type vxlan   --peers 10.10.10.11,10.10.10.12,10.10.10.13 --mtu 1450

# 2. Create a VNet inside the zone with its own VXLAN tag
pvesh create /cluster/sdn/vnets --vnet prod-net --zone vxzone1 --tag 100100

# 3. Optional: subnets (needed for SDN DHCP/IPAM in L3 zones)
pvesh create /cluster/sdn/vnets/prod-net/subnets   --subnet 10.100.0.0/24 --gateway 10.100.0.1 --type subnet   --dhcp-range start-address=10.100.0.100,end-address=10.100.0.200

# 4. Commit - this compiles and reloads networking on every node
pvesh set /cluster/sdn

MTU is the number-one cause of VXLAN problems: the overlay adds roughly 50 bytes of header, so a 1500-byte underlay requires a 1450-byte zone MTU. If the physical fabric carries jumbo frames, set the zone MTU accordingly rather than leaving the default and wondering why large transfers stall.

EVPN Zones: Adding a Control Plane

# Controller (BGP/FRR) first
pvesh create /cluster/sdn/controllers --controller evpnctl --type evpn   --asn 65001 --peers 10.10.10.11,10.10.10.12,10.10.10.13

# Then the zone, referencing the controller and the exit nodes
pvesh create /cluster/sdn/zones --zone evpnzone --type evpn   --controller evpnctl --vrf-vxlan 4000 --exitnodes pve1,pve2

With a BGP or IS-IS controller that has a loopback interface, VTEP IPs are auto-derived from the loopback — which makes VTEP addressing stable across reboots and is the behaviour you want in production. Exit nodes provide the L3 gateway out of the overlay.

Verification

# Applied configuration on a node
cat /etc/network/interfaces.d/sdn
ip -br link show type vxlan
bridge link show

# VNet runtime state
pvesh get /cluster/sdn/vnets/prod-net
pvesh get /nodes/pve1/sdn/zones
ip -d link show dev prod-net        # confirm vxlan id, local/remote, port 4789

# Underlay reachability between VTEPs (peers must be mutually reachable)
ping -c2 10.10.10.12
# FDB entries learned over the overlay
bridge fdb show dev vxlan_prod-net | grep -v permanent

For EVPN zones also check the FRR daemon on each node (vtysh -c "show bgp l2vpn evpn summary") — an overlay with no BGP session will still show interfaces up while no MACs are ever learned.

Operational Notes and Pitfalls

  • Nothing is live until you apply. Creating zones and VNets only writes objects; guests see the bridge after pvesh set /cluster/sdn (or the Apply button).
  • VXLAN is not encrypted. Joining sites over the internet requires an additional site-to-site VPN — plan it before, not after.
  • Isolate ports on a VNet prevents guest-to-guest traffic on the same bridge, which is exactly what you want on multi-tenant front-end networks. Remember the isolation applies to the bridge ports, not to the VNet interface itself.
  • Rollback is a config operation. Keep a copy of /etc/network/interfaces on each node before the first apply; a bad VNet can leave a node unreachable except through the console.

Related reading: Proxmox VLAN-aware bridges and LACP bonds, Linux bridge VLAN configuration and VXLAN MTU, fragmentation and jumbo frames.

原文链接:https://pve.proxmox.com/pve-docs/chapter-pvesdn.html