Lossless RoCE on MLNX-OS: PFC, ECN and DSCP Trust - 夜莺博客

Lossless RoCE on MLNX-OS: PFC, ECN and DSCP Trust

RoCEv2 moves RDMA traffic over a routed, loss-sensitive fabric, and the difference between 90 percent and near-100 percent GPU cluster utilisation is usually congestion control configuration rather than bandwidth. Lossless RoCE on MLNX-OS needs three things aligned end to end: PFC to stop drops on the lossless priorities, ECN marking so senders slow down instead of being paused forever, and a consistent DSCP-to-priority mapping on every device in the path, including the hosts. This article shows the switch-side configuration and the checks that prove the whole chain agrees.

The three mechanisms and why you need all of them

  • PFC (Priority Flow Control) — per-priority pause frames. It prevents buffer overflow drops on RoCE priorities, but pause-only designs propagate backpressure and can cause head-of-line blocking and PFC storms.
  • ECN — routers and switches mark congested packets instead of dropping them, and DCQCN-capable NICs respond by reducing their send rate. ECN is what makes PFC a rare exception rather than the normal operating mode.
  • DSCP trust and mapping — the NIC marks RoCE traffic with a DSCP value; every switch must map that value to the same internal priority and to the same PFC-enabled priority, or the traffic lands in a lossy queue.

Step 1 — trust DSCP and check the mapping

switch (config) # interface ethernet 1/10 qos trust dscp
switch (config) # show qos maps dscp-to-pcp
switch (config) # show qos maps dscp-to-switch-priority

Trusting DSCP on server-facing ports is not optional. With the default trust mode, a tagged frame's PCP value or an untagged frame's port default decides the priority, and RoCE traffic ends up in the best-effort queue — invisible in normal monitoring, obvious the first time the fabric is congested.

Step 2 — PFC on the lossless priority

switch (config) # interface ethernet 1/10
switch (config interface ethernet 1/10) # qos pfc priority 3 rx
switch (config interface ethernet 1/10) # exit

switch (config) # dcb priority-flow-control mode on
switch (config) # show dcb priority-flow-control
switch (config) # show interfaces ethernet 1/10 counters priority-flow-control

Enable PFC on the same priority number used by the host NIC and by the DSCP mapping. A mismatch here is the most common reason a "lossless" fabric still shows drops: PFC is protecting priority 3 while RoCE is actually arriving on priority 5.

Step 3 — ECN marking thresholds

switch (config) # interface ethernet 1/10
switch (config interface ethernet 1/10) # traffic-class 3
switch (config interface ethernet 1/10 traffic-class 3) # wrr-weight 1
switch (config interface ethernet 1/10 traffic-class 3) # ecn
switch (config interface ethernet 1/10 traffic-class 3) # min-threshold 153600
switch (config interface ethernet 1/10 traffic-class 3) # max-threshold 1536000
switch (config interface ethernet 1/10 traffic-class 3) # exit

Set the ECN min/max thresholds as a fraction of the port buffer rather than copying values from another site. Too low and you mark normal bursts, making senders throttle unnecessarily; too high and PFC does all the work, which is where latency and PFC-storm risk come from. A common starting point on a 100G port is a mark-min around one buffer unit's worth of bytes and a max several times that, then tune from telemetry.

Step 4 — make the configuration persistent

switch (config) # write memory
switch (config) # show running-config | include "pfc|ecn|qos trust"

On MLNX-OS, configuration changes applied through the CLI live in the running configuration until saved. Confirm both the running and startup views, especially after scripting changes across a fabric.

Verification path: switch, then host

! Switch side
switch # show interfaces ethernet 1/10 counters
switch # show interfaces ethernet 1/10 counters priority-flow-control
switch # monitor counters interface ethernet 1/10

! Host side (ConnectX)
$ mstconfig -d 5e:00.0 q | grep -E "ECN|PFC"
$ ethtool -S eth0 | grep -Ei "prio3|pause|cnp"

The pairing matters: pause frames counted on the switch but no CNP activity on the host means the NIC is not responding to ECN, usually because ECN was never enabled at the driver level or the DSCP value does not match the switch mapping. Cisco's RoCE design guidance makes the same point from the other direction — ECN marking has to happen where congestion occurs, and both PFC-only and ECN-mostly designs exist, with the mixed design being the safest default.

Common production failures

Symptom Likely cause
Zero throughput on one server, fine elsewhere PFC priority mismatch between NIC and switch port
Latency spikes cluster-wide under load ECN thresholds too high, so PFC backpressure dominates
Fabric-wide slow down from one flow PFC storm — no ECN and no headroom buffer, backpressure propagated upstream
Drops only on multi-hop paths DSCP trust or mapping inconsistent on one intermediate switch
RoCE works, TCP storage suffers Buffer shared unevenly; reserve headroom for the storage class too

Deployment checklist

  1. Pick one DSCP value for RoCE and document it; every GPU node, storage node and switch must use it.
  2. Enable trust DSCP on all server-facing ports.
  3. Enable PFC on the matching priority, RX side at minimum, on every port that carries RoCE.
  4. Enable ECN with tuned thresholds, and verify CNP counters on the hosts.
  5. Validate with a single-flow saturation test before scaling to the whole cluster.
  6. Monitor PFC pause counters as a first-class metric — a permanently pausing fabric is a misconfiguration, not a design.

Buffer and headroom planning

PFC only works if the switch has somewhere to put the packets it is protecting. Without reserved headroom, a port that receives a pause frame has to hold the in-flight traffic in the same shared pool as every other queue, and the first burst that exceeds the pool becomes a drop — the exact outcome PFC was meant to prevent.

Setting Purpose Typical approach
Shared pool size Total buffer available across ports Leave the platform default unless you have measurement data
Lossless pool / headroom Buffer reserved for PFC-protected traffic Reserve explicitly for the RoCE priority; do not rely on shared allocation
Per-port reserved Guaranteed buffer for one port Size for the round-trip time the link needs to survive a pause
ECN mark thresholds Point at which senders are told to slow down Set below the point where PFC starts pausing

As a rule of thumb, the headroom for a lossless priority should be at least the product of link speed and round-trip time, plus a margin for bursts. On a 100G port with a 10 microsecond round trip, that is on the order of 125 KB per port — which is why large fabrics reserve buffer in the hundreds of KB per port for the storage class.

Diagnosing a PFC storm

A PFC storm is self-reinforcing: a link that cannot drain receives pauses, pauses cause more buffering upstream, and the backpressure spreads across the fabric until unrelated traffic stalls. The symptoms are distinctive enough to recognise quickly.

  1. Application latency rises cluster-wide, not just on one server.
  2. Pause counters climb on ports that carry no storage traffic at all.
  3. Drops appear in a queue that was previously idle.
  4. Throughput collapses for small flows while large flows partially survive, because the pauses affect everyone.
switch # show interfaces ethernet 1/10 counters priority-flow-control
switch # show interfaces ethernet 1/10 counters
switch # show qos maps dscp-to-switch-priority
switch # monitor counters interface ethernet 1/10

Mitigation follows three steps: identify the port whose pause counters grow first (that is the congested egress), confirm the RoCE priority mapping end to end, and add ECN if it is missing. Some platforms also support a PFC watchdog that detects a port that has stopped transmitting and resets it, which is worth enabling where available because it converts a fabric-wide stall into a single-port event.

Change management for a live RDMA fabric

  • Change one switch at a time, and never change the DSCP mapping without confirming every host NIC uses the same value.
  • Record the queue mapping table before and after the change so a rollback restores the exact previous state.
  • Run a bandwidth test after each switch, not just after the whole fabric.
  • Keep PFC counters on a dashboard that operations can see, so a slow drift is caught before an application team reports it.

Related reading

Base switch configuration including VLANs and routing is covered in Mellanox MLNX-OS VLAN and IP routing configuration. For host-side tuning that affects RDMA throughput, see Linux network tuning with sysctl for TCP buffers and backlog and ethtool Linux network diagnostics.

原文链接:https://enterprise-support.nvidia.com/s/article/lossless-roce-configuration-for-mlnx-os-switches-in-dscp-based-qos-mode--advanced-mode-x