Microbursts and Switch Buffer Sizing in the Data Center - 夜莺博客

Microbursts and Switch Buffer Sizing in the Data Center

A 100G link at 10 percent average utilisation can still drop packets, because the congestion that matters happens in bursts of a few milliseconds during which several senders transmit simultaneously to the same egress port. Average-rate monitoring cannot see this, and neither can most of the tuning advice written for 1G networks.

Why buffers overflow at low average utilisation

A switch buffer is shared across all ports and queues on an ASIC. When several egress queues are active at the same instant, each gets a smaller share, so a burst that would have been absorbed on an idle switch is dropped on a busy one. Large-scale measurement work from Meta quantified the effect: contention varies over milliseconds and produces per-queue buffer reductions in the range of 34-70 percent over short periods, and nearly all bursts encounter some contention - but higher contention does not automatically mean more loss, because burst shape and workload matter as much as buffer share.

  • Burst length. Very short bursts are absorbed by whatever buffer exists. Long bursts that last multiple round-trip times can be handled by congestion control. The dangerous middle - a few milliseconds - is where losses concentrate.
  • Connection count. A burst composed of many parallel connections behaves differently from one composed of a few large ones.
  • Workload placement. Contention characteristics are persistent per rack and are influenced by which services are co-located, so placement decisions have network consequences that do not appear in any capacity plan.

Instrumenting for bursts

The first requirement is measurement at the right time scale. Interface counters sampled every minute report a maximum burst rate that may be several times the average and still invisible.

! What to collect, at what granularity
- Per-queue tx/rx instantaneous occupancy (watermark), not just max
- Per-queue drop counters, polled every 5-15 s
- PFC pause frame counts per priority queue where lossless is configured
- ECN marking counters
- ULL/telemetry streaming where the platform supports it

Platform features differ: some switches expose per-queue watermark and buffer utilisation counters, some expose them only through streaming telemetry, and some expose only aggregate drops. Determine which of these your hardware provides before designing a response - the answer often rules out half the options.

Tuning levers, in order of effectiveness

  1. Reduce the burst at the source. Host-side changes - pacing, MTU, fewer parallel flows per destination - are usually more effective than switch tuning, and they are free of side effects on other traffic.
  2. ECN rather than drop. Marking packets as congestion experienced lets the transport react before the buffer overflows. It requires both endpoints to implement it and to react correctly, which is the usual reason it is enabled on paper and ineffective in practice.
  3. Buffer sharing policy. Dynamic threshold settings determine how much of the shared pool a queue may claim. A larger share protects bursts but can starve other queues; the setting trades one class of loss for another and must be evaluated with the actual application mix.
  4. Queue scheduling. Strict priority for latency-sensitive traffic with weighted scheduling for the rest prevents bulk flows from delaying control traffic, but does nothing for a burst targeting its own queue.
  5. PFC where lossless is a hard requirement. Pause frames remove loss but introduce head-of-line blocking and pause storms; they need watchdog and storm-control protection, and their interaction with ECN must be tuned deliberately rather than left at defaults.

Diagnosing a reported microburst problem

Work from the host outward. Capture on the sender to see the actual burst shape and inter-packet gaps; on the switch, look for per-queue drops correlated in time with the application's timing; and on the receiver, separate loss from reordering. A pattern of drops all on one egress queue with simultaneous activity on several ingress ports points to buffer contention. Drops spread across queues at high ingress rates point to an oversubscribed fabric. And drops with no corresponding burst on the wire usually mean the problem is a policing or shaping configuration rather than the buffer at all.

Related reading: WRED versus tail drop, Flexible NetFlow configuration for flow-level visibility, and Nexus 9300 versus 9500 platform selection when buffer depth is the deciding purchase criterion.

原文链接:https://engineering.purdue.edu/~isl/papers/imc2022.pdf