Spine-Leaf Architecture: Modern Data Center Topology - 夜莺博客

Spine-Leaf Architecture: Modern Data Center Topology

Modern data center networks are built on spine-leaf, a flat two-tier fabric that replaced the classic three-tier core/distribution/access model. The reason is traffic patterns: microservices, virtualization and big-data workloads generate huge east-west traffic, and three-tier oversubscription makes that traffic slow and unpredictable. This article explains the spine-leaf layout, why every leaf-to-leaf path is exactly two hops, how ECMP spreads load across spines, and where VXLAN and EVPN fit on top of the fabric — plus the number-one cabling mistake that breaks the whole design.

The Spine-Leaf Layout

Every leaf switch (top-of-rack) connects directly to every spine switch (backbone). Rules: no leaf-to-leaf uplinks, no spine-to-spine uplinks, every leaf connects to every spine, and spines have uniform port counts so all leaves have equal reach. With 4 spines and 24 leaves that is 96 spine-leaf uplinks. Two servers on different racks always talk over the same two-hop path: server to leaf1, leaf1 to spine, spine to leaf2, server.

The shape is a Clos network, and the important property is not the picture — it is the arithmetic. Any leaf can reach any other leaf through exactly N spines, where N is the number of spines in the fabric. Bandwidth between any two leaves is a single link's worth of capacity multiplied by N, because the fabric accepts simultaneous flows on all N paths through ECMP. That is why the topology is described as a fabric rather than a hierarchy: there is no "center" that bottlenecks.

Compare that to the three-tier model, where a server in one access switch reaching a server in another access switch may traverse access, distribution, core, then back down distribution and access. The hop count grows, the core concentrates traffic, and spanning tree disables half the links to stay loop-free. Spine-leaf removes all three problems at once with a single design decision: route at Layer 3 from every leaf.

Why It Beats Three-Tier in the Data Center

  • Latency: 2 hops always, deterministic — versus up to 5 hops in three-tier
  • Bandwidth scaling: add another spine (parallel, non-disruptive) instead of upgrading core links
  • Port scaling: add another leaf
  • Redundancy: any single spine or leaf can fail; all others keep forwarding
  • No STP bottleneck: layer 3 to the edge kills classic spanning-tree problems
  • Uniform performance: every server is the same distance from every other server, so capacity planning does not depend on which rack you land in

Oversubscription: The Number That Actually Decides the Design

Oversubscription is the ratio between (a) the bandwidth available toward the leaf's uplinks and (b) the bandwidth offered by the servers hanging off that leaf. A leaf with 48 server ports at 25G offers 1200G of potential ingress; if it has 8 uplinks at 100G, it has 800G of spine-facing capacity. That is a 1.5:1 oversubscription ratio. A 1:1 design is expensive but non-blocking; 3:1 to 4:1 is common for general-purpose racks; AI and storage fabric racks often demand 1:1 or better.

The formula to keep in mind: oversubscription = (downlink capacity) / (uplink capacity). Design the leaf uplink count from that ratio, then choose the spine count so that a single spine failure still leaves acceptable capacity. If losing one spine drops the fabric from 8 to 7 spines, capacity falls by 12.5 percent — usually fine. If losing a spine halves leaf-to-leaf bandwidth, the fabric is under-built.

How Many Spines and Leaves Do You Need?

The sizing questions are sequential and each answer constrains the next.

  • How many leaves? Total rack-facing ports required, divided by ports per leaf. Add margin for growth — putting a new switch in service is easy, but uplink rebundling is disruptive.
  • How many uplinks per leaf? Derived from the oversubscription target. Round to a number the ASIC can hash across evenly.
  • How many spines? The uplink port count of a spine divided by the uplinks each leaf consumes. All leaves must connect to all spines, so spine port count is the hard limit.
  • Radix check. If leaves multiply so far that a spine cannot host them all, the fabric has exceeded its radix and needs a second tier (spine-super-spine) or larger fixed-form-factor switches.

Work an example. 40 racks, two leaves per rack, each leaf needing 8×100G uplinks. That is 80 leaves and 640 uplink ports. A 32-port 400G spine configured as 64×100G gives 64 uplink ports per spine, so 10 spines would be required — well beyond the common 4-to-8 spine sweet spot, and each link would be under-utilised. A better answer is fewer, fatter uplinks per leaf (for example 4×400G) which cuts the port count to 320 and lets 8 spines of 32×400G carry the fabric. This trade-off — more parallel links or fewer faster links — is where most design arguments actually happen.

What Runs on Spine-Leaf in Practice

Layer 3 terminates at every leaf, so no VLANs stretch across leaves. ECMP routing (OSPF, eBGP-in-the-DC, or IS-IS) spreads traffic across all spines. VXLAN overlays extend the L2 segments tenants need over the L3 fabric, letting VMs move between racks without renumbering, and BGP EVPN advertises MAC/IP reachability across the overlay.

In a modern leaf-spine fabric the design is usually described in three layers that are explicitly separate:

  • Underlay — point-to-point L3 links between each leaf and each spine, running eBGP (typically an AS per leaf, a common AS for all spines) or OSPF/IS-IS. Its only job is to give every leaf a route to every other leaf's loopback with maximum-paths set high enough that all spines are in the forwarding table.
  • Overlay — VXLAN tunnels with an EVPN control plane. EVPN type-2 routes carry MAC and IP reachability, type-3 routes handle BUM replication via ingress replication, and type-5 routes carry prefix routes for routing between subnets or toward external networks.
  • Services — border leaves, firewalls, load balancers and WAN handoff. These attach to one or two leaves as "service leaves" so that policy lives at a defined edge rather than being spread across the fabric.

Underlay Routing Options Compared

  • eBGP — the dominant choice. Each leaf is its own AS, all spines share an AS. It is simple, has no flooding, gives clean per-link visibility, and maximum-paths plus bestpath as-path multipath-relax produces full ECMP. Convergence is per-session and predictable.
  • OSPF — familiar and easy, but link-state flooding grows with fabric size and area design becomes necessary in larger builds. Fine for small fabrics, awkward at scale.
  • IS-IS — scales cleanly, no dependence on IP for the routing protocol itself, and widely used in large service-provider-style fabrics and in hyperscale designs.

ECMP Hashing and Flow Distribution

ECMP does not balance packets; it balances flows. The switch computes a hash over packet header fields and picks one of the equal-cost next hops. The default hash is usually a five-tuple (source IP, destination IP, source port, destination port, protocol). The consequence is important: a single large TCP flow traverses one spine only, so one elephant flow can never exceed a single link's capacity no matter how many spines exist. Many small flows distribute beautifully; a few huge flows do not.

Standard mitigations, in increasing order of complexity:

  • Ensure the hash includes L4 ports and, where supported, is configured for symmetric hashing so both directions of a flow take the same path (important for stateful devices in the path).
  • Increase the number of parallel paths so statistical distribution smooths out. More spines reduces the elephant-flow penalty as a fraction of total capacity.
  • Use flowlet or dynamic load balancing, which re-hashes when a flow's packet gap crosses a threshold. This can rebalance long-lived flows without reordering packets within a burst.
  • For AI/GPU traffic, consider the fabric-level answer instead: RoCE with PFC and ECN, or an InfiniBand/non-Ethernet fabric where the switch performs per-packet adaptive routing.

A Minimal eBGP Underlay Configuration

The following shows the shape of a leaf's underlay on a NX-OS-style CLI. Spines use the same pattern with the neighbour list reversed.

feature bgp
feature interface-vlan
feature nv overlay

interface Ethernet1/1
  description to-spine1
  no switchport
  ip address 10.1.0.1/31
  no shutdown

interface loopback0
  ip address 10.255.0.1/32

router bgp 65001
  router-id 10.255.0.1
  maximum-paths 8
  bestpath as-path multipath-relax
  address-family ipv4 unicast
    network 10.255.0.1/32
  neighbor 10.1.0.0
    remote-as 65000
    description spine1
    address-family ipv4 unicast

Repeat the interface and neighbour blocks for every spine, then verify that the loopback of every other leaf appears with all spines as next hops.

Verifying the Fabric

Verification follows the same three-layer logic as the design.

show ip route 10.255.0.2/32
show ip bgp summary
show ip bgp 10.255.0.2/32
show forwarding route 10.255.0.2/32
show nve peers
show nve vni
show bgp l2vpn evpn summary
  • Underlay check — the remote loopback route should list every spine IP as a next hop. If only one appears, maximum-paths is missing or the AS-path is not being relaxed.
  • Hardware check — show forwarding route proves the ASIC programmed the ECMP group, not just that the routing table looks right. Software/hardware divergence is a classic cause of "the route exists but traffic does not flow".
  • Overlay check — show nve peers should show all remote VTEPs up, and the EVPN table should contain the expected type-2 and type-3 routes.
  • Tunnel path check — the VXLAN tunnel source must be the loopback, and the underlay must reach that loopback with ECMP. If tunnels pin to a physical interface instead, a single uplink failure takes down the overlay.

Troubleshooting and Common Failure Modes

  • One leaf cannot reach another. Check the loopback advertisement first, then confirm both leaves peer with all spines. A missing session on one spine usually means a physical or address-family error, not a routing problem.
  • Traffic asymmetric or single-pathed. Confirm ECMP is in the forwarding table with the expected width, and check the hash configuration for symmetric hashing.
  • VXLAN MTU black holes. The tunnel adds 50 bytes of overhead. If the underlay MTU is 1500, jumbo frames will be dropped inside the tunnel; the fabric links need MTU 9216 or the servers need MSS clamping. Symptoms are confusing: ping works, bulk transfer stalls.
  • Duplicate VTEP or IP conflict. A leaf with the wrong loopback address produces intermittent EVPN withdrawals and flapping MACs.
  • Spine-to-spine or leaf-to-leaf links found in the cabling. These create exactly the loops the design exists to avoid. See the mistake below.

Cabling and Optics in Practice

Leaf-to-server is usually copper or short-reach optics; leaf-to-spine is usually multimode or single-mode optics over structured cabling or direct-attach. Two rules matter more than the rest. First, keep every leaf-to-spine link the same speed and type wherever possible: mixed speeds mean unequal cost paths, and ECMP will only use the ones the routing protocol considers equal. Second, document the physical mapping between leaf uplink port and spine, because a swapped pair of fibres is the single most common commissioning defect and produces exactly the symptoms of a routing problem.

When Three-Tier Still Makes Sense

Campus networks with heavy north-south traffic, very small deployments where a collapsed core suffices, and environments with legacy L2 requirements that cannot tolerate L3 boundaries at every rack.

The #1 Mistake

Cabling a spine-leaf like a three-tier network — adding leaf-to-leaf or spine-to-spine uplinks. This creates L3 forwarding loops or convergence issues and destroys the deterministic two-hop property. Spines never talk to each other; leaves never talk to each other directly.

The second most common mistake is subtler: building a correct topology but sizing the spines for today's leaf count only. Since every leaf must reach every spine, adding a leaf after the spines are full is impossible without a forklift. Buy spine headroom at build time.

FAQ

Can I have more than two tiers? Yes — a super-spine tier appears when leaf count exceeds what one spine tier can hold. The principle is unchanged; only the radix arithmetic is larger.

Do I need VXLAN? No. A pure L3 fabric with routed subnets per rack is simpler and faster. VXLAN/EVPN is needed when workloads must keep addresses across racks, or when tenant segmentation must extend over the L3 boundary.

Is MLAG still used? For servers dual-homed to a pair of leaves, yes — MLAG or EVPN multihoming serves that purpose. It is orthogonal to the spine-leaf fabric itself.

How many spines is "right"? Four is the usual minimum for meaningful failure resilience, eight is a common production sweet spot, and larger fabrics scale by radix rather than by preference.

Related: EVPN MAC-VRF validation on ACX7000, Arista MLAG run book, and the SONiC troubleshooting guide. For the comparison view, read Leaf-Spine vs Three-Tier: 2026 Data Center Design Guide and Spine-Leaf vs Three-Tier: STP Blocking and Bandwidth. For the overlay, see VXLAN EVPN Multi-Site: Border Gateway Design Guide and Cisco Nexus NX-OS OSPF Configuration.

原文链接:https://packetmentor.com/topics/spine-leaf-architecture