VMware NSX Segments, Transport Zones and TEPs - 夜莺博客

VMware NSX Segments, Transport Zones and TEPs

VMware NSX replaced the old virtual switch sprawl with a distributed overlay controlled by a central manager, and most of the confusion around it comes from vocabulary. A "port group" becomes a segment, a "vDS uplink" becomes a transport node uplink profile, and the tunnels between hosts are GENEVE TEPs owned by the transport zone. Once those three concepts click, NSX operations become predictable. This guide explains transport zones, TEP addressing and uplink profiles, then shows how to build and verify an overlay segment for a workload that must talk to the physical network.

Three building blocks: transport zone, transport node, segment

  • Transport zone — the scope of the overlay. Hosts in the same transport zone can reach each other over the TEP network and share the same segment namespace. Different clusters or different networks should generally be separate transport zones.
  • Transport node — an ESXi host (or a bare-metal Edge node) prepared by NSX. It runs the NSX virtual switch and holds the TEP interfaces where GENEVE traffic is encapsulated.
  • Segment — the replacement for a VLAN port group. A segment is a logical L2 domain distributed across every transport node in its transport zone, independent of physical rack boundaries.
# check transport node state from an ESXi host
nsxdp-cli nst get          # NSX datapath state
esxcli network ip interface list
esxcli network ip route ipv4 list
net-vdl2 -l -s              # NSX logical device state

Design the TEP network before you touch the UI

The TEP network is the underlay: it carries encapsulated traffic between hosts and must have its own addressing and its own routing. Rules that save a lot of pain:

  1. Use a dedicated VLAN and subnet (or a routed /26 per rack) for TEPs. Never share the management network.
  2. Size it for the encapsulation overhead: reserve at least 1600-byte MTU, and 9000 if the physical fabric supports jumbo frames, because GENEVE plus the inner frame easily exceeds 1500.
  3. Make sure every TEP can reach every other TEP in the same transport zone with symmetric routing. Asymmetric paths break the flow cache and cause intermittent, hard-to-diagnose drops.
  4. Plan for N+1 links. A single uplink per host is a design fault; NSX will happily use it and equally happily lose half your east-west traffic.
# verify TEP reachability between two hosts (from each host's shell)
vmkping -I vmk10 -s 1572 -d 10.10.0.22
# and confirm the MTU on the TEP interface itself
esxcli network ip interface ipv4 get | grep -i tep

Uplink profiles and NIOC

An uplink profile maps each host's physical NICs to the traffic types NSX recognises — overlay (TEP), VLAN-backed uplinks, and Edge traffic. Two profiles are normally enough: one for hypervisor hosts and one for Edge nodes. Pair them with Network I/O Control shares so that overlay traffic has guaranteed bandwidth when a single uplink fails and everything funnels through the survivor.

Host uplink profile (example values)
- Uplinks: uplink-1 (vmnic0), uplink-2 (vmnic1)
- Transport VLAN: 203
- MTU: 9000
- Teaming: failover with explicit active/standby, or LACP if the ToR supports it
- NIOC: overlay 50 shares, vMotion 30 shares, management 20 shares

If the top-of-rack switch runs LACP, configure the port channel first and then tell NSX to use it. Mixing a static port channel on the switch with independent NIC failover on the host is a classic cause of silent half-bandwidth operation.

Building a segment and connecting it to the physical world

# create a segment via the manager API
curl -k -u 'admin:' -X POST \
  https://nsx-mgr.example.com/policy/api/v1/infra/segments/web-tier \
  -H "Content-Type: application/json" -d '{
    "display_name": "web-tier",
    "transport_zone_path": "/infra/sites/default/enforcement-points/default/transport-zones/tz-overlay",
    "subnets": [{"gateway_address": "10.20.30.1/24"}]
  }'

A segment with a gateway address gets distributed routing in every transport node, so a VM on host A reaches a VM on host B without hairpinning to an Edge. Traffic leaving the overlay for the physical network is handled by a Tier-0 gateway with an uplink, which is where you define BGP or static routing toward your fabric.

# list segments and their realised state
curl -k -u 'admin:' https://nsx-mgr.example.com/policy/api/v1/infra/segments | jq '.results[].display_name'
# Tier-0 gateway routing state
curl -k -u 'admin:' -X POST https://nsx-mgr.example.com/policy/api/v1/infra/tier-0s/t0-gw/locale-services/default/bgp/neighbors | jq

Verification checklist

# from the manager: is the segment realised everywhere?
GET /policy/api/v1/infra/segments/web-tier/state

# from a host: is the logical port up and the flow cache populating?
netstat -an | grep 6081
vmknic-stats -l TCPIP 2>&1 | tail

# from a VM: does it see the gateway and reach another rack?
ip route
ping -M do -s 1472 
traceroute 

Expect GENEVE traffic on UDP 6081 between TEP addresses, a populated flow cache (zero indicates the overlay is not being used at all), and successful full-size pings. If pings succeed but a bulk copy stalls, revisit MTU: the encapsulation overhead demands an underlay that supports at least 1600 bytes.

Common operational problems

  • Segment up, no connectivity between racks: the TEP network is not routed symmetrically. Check the underlay routing table and any firewall that inspects UDP 6081.
  • Intermittent packet loss under load: a single uplink is the active path for overlay traffic. Fix the uplink profile, not the overlay.
  • A VM cannot reach the physical gateway: the Tier-0 uplink or its routing adjacency is down. Verify the BGP session in the Tier-0 locale service.
  • Post-upgrade, half the hosts lose TEP reachability: a NIC driver or firmware mismatch on the TEP vmkernel interface. Compare driver versions across the cluster before rolling back.

Best practices

  • Keep the TEP underlay boring: separate VLAN, static routes or a single IGP, and no policy in the middle.
  • Document every segment with its gateway and owning team; segments multiply faster than VLANs ever did.
  • Use LACP where the ToR supports it and failover where it does not — do not mix the two on the same host.
  • Back up the manager configuration, and test that you can restore it; the ESXi configuration backup procedure covers the host half of that job.
  • Size the overlay MTU once, in the design phase — the VXLAN/GENEVE VTEP notes explain the same arithmetic for Linux endpoints.

原文链接:https://techdocs.broadcom.com/us/en/vmware-cis/nsx/vmware-nsx/4-2.html