WRED vs Tail Drop: Congestion Avoidance Explained - 夜莺博客

WRED vs Tail Drop: Congestion Avoidance Explained

When a queue fills up, the interface has to drop something, and the only question is how. The default behaviour on most platforms is tail drop: the queue fills completely and then every arriving packet is discarded until congestion clears. Because TCP senders respond to loss by halving their windows at roughly the same moment, tail drop produces global synchronisation — throughput sawtooths, link utilisation collapses, and the pattern repeats. Weighted Random Early Detection attacks the same problem earlier and more gently. This article explains the algorithm, the three parameters that govern it, and how to size them so you drop a few packets on purpose rather than thousands by accident.

What each mechanism actually does

Tail drop is the default congestion avoidance behaviour when WRED is not configured: queues fill during congestion, and once full, packets are dropped until the queue empties. It treats all traffic equally and has no notion of class of service, which is exactly why it interacts badly with mixed traffic — a single aggressive flow can consume the buffer and starve interactive traffic that would have fit in a much smaller queue.

WRED monitors the average queue depth rather than the instantaneous one, and begins dropping a small percentage of packets before the queue is full. Because the drops are early and spread out, TCP senders reduce their rates at different times and the sawtooth flattens.

The three knobs

Parameter Meaning Typical value
Minimum threshold Average queue depth at which random drops begin Half of maximum, at minimum
Maximum threshold Average queue depth above which all packets are dropped Vendor default or sized to BDP
Mark probability denominator Fraction of packets dropped when the average sits at the maximum threshold 10 (i.e. 1 in 10)
Exponential weight constant How heavily the running average reflects history versus the current depth 9 on most Cisco platforms

The asymptotic behaviour is easy to reason about: below the minimum threshold nothing is dropped, between the two thresholds the drop probability rises linearly, and above the maximum threshold everything is dropped. Non-IP traffic is treated as precedence 0, so it is dropped more eagerly than marked IP traffic unless you override it.

Command syntax on Cisco IOS and IOS XE

WRED can be enabled directly on an interface or, more usefully, inside a class of a policy map so that different classes get different thresholds.

! Interface-level with DSCP-based thresholds
interface TenGigabitEthernet0/1
 random-detect dscp-based
 random-detect dscp af11 32 40 10
 random-detect dscp af21 28 40 10
 random-detect dscp af31 24 40 10
 random-detect exponential-weighting-constant 9

! Class-based, inside an MQC policy
policy-map WAN-EDGE
 class VOICE
  priority level 1
 class BULK-DATA
  bandwidth remaining percent 40
  random-detect
  random-detect precedence 0 20 40 10
  random-detect precedence 6 33 40 10

Note the interaction rule that catches people out: attaching a service policy configured to use WRED to an interface disables WRED configured on that interface. If any class in the policy map uses WRED, do not leave interface-level WRED enabled underneath it.

Sizing the thresholds

Two mistakes dominate. Setting the minimum threshold too low drops packets unnecessarily and leaves the link under-utilised — you paid for the bandwidth and then threw traffic away before congestion was real. Setting the window between minimum and maximum too narrow causes many packets to be dropped at once, which reproduces the global synchronisation that WRED was supposed to prevent. As a rule of thumb, keep the maximum threshold well above the minimum, and align the thresholds with the bandwidth-delay product of the link so the buffer can hold at least one flight's worth of data for high-priority classes.

Because WRED drops a percentage of packets on purpose, aggressive settings look alarming in monitoring: random drops rise while tail drops stay at zero. That is the algorithm working. What you want to watch is the ratio of random drops to tail drops — tail drops mean the queue hit the maximum threshold and the protection failed. Export those counters and graph them; the workflow in Flexible NetFlow configuration covers the flow-level side.

Verifying what is really happening

show queueing interface TenGigabitEthernet0/1
show policy-map interface TenGigabitEthernet0/1
show queueing random-detect

The output lists, per precedence or DSCP class, the transmitted, random-drop and tail-drop counters alongside the configured minimum threshold, maximum threshold and mark probability. Two checks are worth doing after every change: the mean queue depth should sit between your thresholds rather than pinned at the maximum, and the random drop count should be nonzero but far smaller than the transmitted count. A mean queue depth that hovers at the maximum means you are effectively back to tail drop.

When not to use WRED

WRED is designed around adaptive transports: TCP backs off when packets are lost, so early drops reduce rates. Non-adaptive flows such as fixed-rate UDP — including some storage replication and legacy voice implementations — will simply retransmit or suffer, so those classes need policing or shaping instead of drop-based congestion avoidance. On lossless Ethernet fabrics the equivalent mechanism is Priority Flow Control plus ECN marking rather than WRED. For shaping on Linux hosts and routers, the token-bucket approach in Linux tc HTB traffic shaping is the parallel tool, and the classification and marking rules that feed either mechanism are described in H3C Comware priority trust and QoS mapping.

原文链接:https://www.cisco.com/c/en/us/td/docs/ios-xml/ios/qos_conavd/configuration/15-mt/qos-conavd-15-mt-book/qos-conavd-cfg-wred.html