RoCEv2 Lossless Tuning: PFC, ECN and DCQCN - 夜莺博客

RoCEv2 Lossless Tuning: PFC, ECN and DCQCN

RDMA over Converged Ethernet carries InfiniBand's transport over Ethernet, and its retransmission logic is far cruder than TCP: one dropped packet forces a go-back-N retransmission of everything behind it, which can collapse throughput on an otherwise healthy fabric. The industry answer is a stack of three mechanisms - ECN marking, the DCQCN rate algorithm on the NIC, and PFC pause frames as a last resort. This article explains the intended firing order, the single tuning rule that decides whether the design works, and the counters that prove it on a real fabric.

Why RDMA Punishes Loss So Hard

TCP was designed for networks that lose packets. Every flow carries a sliding window, a congestion window, selective acknowledgements and a fast-retransmit path, and losing one segment costs roughly one round trip. RDMA's transport is deliberately lean: it was built for lossless InfiniBand fabrics where the link layer guarantees delivery, so its recovery path is crude by comparison. When a RoCEv2 packet is dropped, the receiving queue pair notices a sequence gap, and the sender has to go back to the last acknowledged position and replay everything behind it - a go-back-N retransmission.

Three consequences follow, and all three are operational, not theoretical:

  • Loss amplifies. A single dropped packet can force retransmission of hundreds of packets, because the lost packet was almost certainly at the head of a large burst.
  • Congestion collapses throughput rather than merely degrading it. A queue that overflows briefly during an incast event can take a collective-training job from line rate to a fraction of it.
  • Tail latency is shared. On a multi-tenant AI cluster, one fat flow hitting a congested link raises tail latency for every other job on the fabric, not only the victim.

The design target is therefore not "never drop a packet" as an absolute, but to make drops statistically negligible by signalling congestion early and often instead of waiting for a queue to fill and overflow. That is precisely what ECN plus DCQCN exist to do, with PFC as the backstop when the signalling does not keep up.

The Three Layers and Their Job

Mechanism Layer Role
ECN IP (L3) Marks CE at an early queue threshold; per-flow, no drops
DCQCN NIC Turns CE marks into CNPs and cuts the sender rate, then ramps back up
PFC Ethernet (L2) PAUSE frame at a late threshold; stops the whole priority class

The intended order is ECN first (a gentle signal), DCQCN second (rate adaptation), PFC last (a safety net). If ECN and DCQCN are tuned correctly, PFC should almost never fire; a steady PFC pause count is a warning that the fabric is closer to dropping than the design assumed.

It helps to see the three mechanisms as a gradient of bluntness. ECN touches one flow and asks the sender to slow down, without any loss. DCQCN is the sender obeying that request, adjusting its rate multiplicatively on congestion and additively when idle. PFC is the blunt instrument: it stops an entire priority class for a period measured in microseconds, and while it is asserted every flow sharing that class is frozen, including flows that were behaving perfectly.

The Rule That Decides Everything

Keep the ECN marking threshold comfortably below the PFC XOFF threshold. In practice that means an ECN band spanning hundreds of kilobytes of queue for a 400G port (commonly quoted starting points are around 100 KB minimum and 1.5 MB maximum, with a marking probability in the region of 10%) and a PFC watchdog timer near 100 ms to break a stuck pause before it deadlocks the fabric.

The reason the ordering is non-negotiable is direction of causality. ECN is a request; PFC is a command with a promise attached - PFC guarantees that if the peer obeys the pause, the buffer will not overflow. If the ECN threshold sits at or above the PFC XOFF threshold, congestion is answered by freezing the link before the sender ever gets the chance to react, and the fabric spends its life in pause/resume cycles instead of in smooth rate control. The classic symptom is a fabric that passes every functional test and yet delivers a fraction of expected collective bandwidth.

Headroom, Buffers and the Bandwidth-Delay Product

PFC works by pausing a peer, but a pause frame takes time to arrive and the peer may already have data in flight. The buffer a switch must reserve to absorb that in-flight data before the pause takes effect is the headroom. A useful approximation is one bandwidth-delay product of the link: for a 400G port and roughly 10 microseconds of round trip, that is already several megabytes of reserved buffer per port, which is why deep-buffer switching silicon exists for storage and AI fabrics and why lossless configuration cannot be copied blindly from an access-layer leaf.

Headroom has two practical knobs:

  • Cable and routing length. A longer path means a larger BDP and therefore more mandatory headroom. A configuration that runs clean on a 3-metre DAC inside a rack can start pausing constantly when the same leaf talks across a 2 km campus link.
  • Number of no-drop ports per pool. Headroom is shared across the ports configured for the same no-drop priority. Oversubscribing the pool - many ports configured lossless that were never meant to be - steals buffer from the ports that need it.

Buffer counters matter here: if you never look at queue depth, you cannot tell whether the fabric is running at 20% of its ECN band or permanently pinned against the PFC threshold.

Switch-side Configuration Shape

! Arista EOS example
interface Ethernet1-32
   priority-flow-control on
   priority-flow-control priority 3 no-drop
   priority-flow-control mode auto
   priority-flow-control watchdog action drop timer 100
!
qos profile lossless
   queue 3 ecn min 102400 max 1536000 probability 0.1
!
interface Ethernet1-32
   service-policy type qos output lossless

Every vendor expresses the same three things - no-drop on the chosen priority, a WRED/ECN band on that queue, and a watchdog - but the priority and DSCP marking must match the host. The conventional mapping is DSCP 26 to priority 3, and it only works end to end if the switch is configured to trust DSCP rather than re-marking at ingress.

The same intent on NX-OS and Junos looks like this:

! Cisco NX-OS
interface Ethernet1/1
   priority-flow-control mode on
   priority-flow-control watchdog interval 100
policy-map type network-qos lossless
   class type network-qos c-nq-3
      pause pfc-cos 3
      set cos 3
      congestion-control ecn minimum 102400 maximum 1536000 probability 10
!
system qos
   service-policy type network-qos lossless

! Junos QFX
set class-of-service congestion-notification-profile roce
set class-of-service congestion-notification-profile roce input ieee-802.1 code-point 011 pfc
set class-of-service classifiers dscp roce-classifier forwarding-class f-no-loss
set class-of-service drop-profiles ecn-profile interpolate fill-level 15 drop-probability 0
set class-of-service drop-profiles ecn-profile interpolate fill-level 90 drop-probability 100

Note the watchdog line in both. A watchdog is what stops a PFC storm from turning a healthy fabric into a deadlocked one: if a port stays paused for longer than the watchdog interval, the switch declares a storm and either drops the priority or shuts the port, depending on the action configured. Always enable it - a fabric without a PFC watchdog is one misbehaving NIC away from a site-wide outage.

Host-side Checks

rdma link show
ibdev2netdev
cat /sys/class/net/enp1s0f0/ecn/roce_np/enable/3
ip route get 10.20.30.40       # confirm the RoCE path uses the right interface

tcpdump -ni enp1s0f0 'udp port 4791' -c 20

Port 4791 is RoCEv2. On the NIC side, compare CNP counters and adaptive retransmission counters against a known-good baseline - a rising adp_retrans count with no PFC pauses usually means ECN is firing too aggressively and DCQCN is cutting rates too hard, which shows up as throughput loss with no errors in syslog at all.

Useful per-device counters to pull on the host:

ethtool -S enp1s0f0 | grep -Ei 'ecn|cnp|pfc|pause|prio3'
cat /sys/class/infiniband/mlx5_0/ports/1/counters/port_rcv_pause_duration
cat /sys/class/infiniband/mlx5_0/ports/1/counters/port_xmit_pause

Both sides need to be checked. RoCE congestion is often asymmetric: a leaf may be pausing continuously while its neighbour's counters look clean, because the pause is generated by the receiver of the traffic, not the sender. If you only read counters on the switch you suspected, you will usually be looking at the wrong end.

Walking the Counters End to End

A disciplined walk from the host outwards resolves most tickets in a few minutes:

  1. Host NIC: confirm rdma link is Active, the ECN bit is enabled for the priority, and CNP counters are advancing when you load the link.
  2. Access switch: read per-queue counters for the no-drop priority - ECN-marked packets, pause frames transmitted, pause frames received, and current queue depth.
  3. Spine and uplinks: repeat the same per-queue read. Congestion is usually on the fabric uplink, not the access port, because that is where many-to-one traffic converges.
  4. Correlate time. A pause burst that lines up with a collective job's all-reduce phase is expected; continuous pausing in steady state is not.

If ECN-marked packets are almost zero while pause frames are climbing, the ECN band is set too high or not applied to the right queue. If ECN marks are high and pause frames are zero but throughput is poor, DCQCN on the host is overreacting and its parameters need review.

Failure Signatures Worth Memorising

  • Throughput drops, no PFC pauses visible: ECN over-tuned, DCQCN too conservative.
  • Intermittent pauses on one link only: headroom too small for the cable length/RTT, or a slow consumer.
  • Half the GPUs in a job are slow: hash polarisation on ECMP, not a lossless tuning issue.
  • Link flaps with clean optics: thermal or connector problems, check pre-FEC BER before touching QoS.
  • Pause storm with one port saturating: a NIC ignoring its pause budget, or a mis-marked flow landing in the no-drop class.

Common Misconfigurations

  • No-drop on the wrong priority. The switch pauses priority 3 while the host marks DSCP 26 into priority 4 - the two never meet and nothing is protected.
  • Trust boundary resetting DSCP. An access port that re-marks ingress traffic silently destroys the end-to-end class mapping, so congestion control never engages.
  • ECN above PFC. The single most common cause of a lossless fabric that pauses constantly under load.
  • Headroom left at default. Defaults sized for a different port count and cable length are the usual reason a "correct" configuration misbehaves on one rack.
  • Watchdog disabled. Removes the one mechanism that can break a deadlock automatically.

A Tuning Workflow That Scales

  1. Pick one priority for storage or RDMA traffic and mark it consistently from the host NIC all the way through.
  2. Set PFC no-drop only on that priority, and only on the ports that actually carry it.
  3. Size headroom from the real BDP of the longest path, not the shortest.
  4. Place the ECN band below the PFC threshold with margin, then walk it up while watching pause counters.
  5. Enable the watchdog before you enable PFC, so a mistake during tuning cannot deadlock the fabric.
  6. Record a baseline of per-queue counters after tuning so future tickets have a reference point.

FAQ

Should PFC ever be disabled entirely? On a well-designed fabric running priority-based flow control with correctly tuned ECN and DCQCN, PFC is rarely exercised - but leaving it configured is still the norm, because it is the only guarantee against silent drops during an incast burst. Some operators disable it deliberately because of PFC-storm risk; that trade-off only makes sense with a watchdog and strict headroom control.

Why does my fabric work in a lab and fail in production? Lab traffic is a few flows; production is all-to-all. ECMP hash collisions and incast do not exist in a two-host test, so a configuration that looks perfect in the lab can pause continuously under a real collective job.

Is ECN a per-queue or per-flow mechanism? The marking is per-queue on the switch - the queue depth crosses the threshold and every packet in that queue can be marked - but the reaction is per-flow on the sender, because the returned CNP identifies a specific flow.

Related reading: InfiniBand vs RoCE vs Ethernet for AI clusters, RoCEv2 lossless Ethernet: PFC and ECN configuration, SONiC PFC watchdog and storm mitigation, Arista EOS QoS configuration and IOS XR class-map and policy shaping.

原文链接:Lossless Network: PFC + ECN + DCQCN, the lossless trick