TCP BBR Congestion Control: Enable and Tune on Linux - 夜莺博客

TCP BBR Congestion Control: Enable and Tune on Linux

Loss-based congestion control has one fundamental flaw: it treats packet loss as the only signal of congestion, so on a link with random loss or a deep buffer it either crawls or overshoots. BBR models the bottleneck bandwidth and the round-trip time instead, which is why a single sysctl change frequently doubles throughput on long-haul paths. This guide covers what BBR does differently, how to enable it correctly (including the part everyone forgets), and the adjacent TCP settings that decide whether it can actually help.

What BBR Measures Instead of Guessing

Reno and CUBIC probe for bandwidth by increasing the window until something breaks. BBR continuously estimates two values — the bottleneck bandwidth and the minimum RTT — and paces the sender to the product of the two. The consequences in production are visible and predictable:

  • Higher goodput on paths with 1–5% random loss (long-haul, wireless, WAN).
  • Lower queue occupancy, so less bufferbloat and smaller latency under load.
  • Fairness behaviour that differs from CUBIC: a BBR flow can coexist with CUBIC but converges differently, and multiple BBR flows can be more aggressive than expected on shallow buffers.
  • No benefit — sometimes slight loss — on short-RTT LAN paths where the window never was the limit.

Enabling It

# 1. Check availability
sysctl net.ipv4.tcp_available_congestion_control
# e.g. reno cubic bbr

modprobe tcp_bbr
echo "tcp_bbr" > /etc/modules-load.d/bbr.conf

# 2. Apply at runtime
sysctl -w net.core.default_qdisc=fq
sysctl -w net.ipv4.tcp_congestion_control=bbr

# 3. Persist
cat >/etc/sysctl.d/99-bbr.conf <<'EOF'
net.core.default_qdisc=fq
net.ipv4.tcp_congestion_control=bbr
EOF
sysctl --system

# 4. Verify
sysctl net.ipv4.tcp_congestion_control
ss -tin | grep -m3 bbr
tc qdisc show dev eth0

The setting everyone forgets is the qdisc. BBR is designed to work with fq (fair queueing) because fq provides the pacing and per-flow scheduling BBR expects. With the old pfifo_fast default, BBR still functions but paces less precisely and the latency improvement largely disappears. If your distribution uses fq_codel, keep it if it is doing a good job on the LAN — but test BBR with fq first so you know what the combination is capable of.

The Knobs That Decide Whether BBR Helps

# Buffer sizing must follow the bandwidth-delay product (BDP)
sysctl -w net.ipv4.tcp_rmem="4096 262144 16777216"
sysctl -w net.ipv4.tcp_wmem="4096 262144 16777216"
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.wmem_max=16777216
sysctl -w net.ipv4.tcp_window_scaling=1
sysctl -w net.ipv4.tcp_mtu_probing=1        # helps on tunnels with MTU black holes
sysctl -w net.ipv4.tcp_notsent_lowat=131072 # reduces queueing latency for HTTP-like flows

Rule of thumb for the BDP: 1 Gbit/s at 100 ms RTT is about 12.5 MB in flight. If rmem_max is still the 6 MB default, neither BBR nor CUBIC can fill the pipe — the socket buffer is the limit and no congestion-control algorithm can work around it.

Measurement Methodology

# Baseline first, then change one thing at a time
iperf3 -c server -t 30 -P 1 --json > cubic-p1.json
sysctl -w net.ipv4.tcp_congestion_control=bbr
iperf3 -c server -t 30 -P 1 --json > bbr-p1.json
# and with parallel streams
iperf3 -c server -t 30 -P 8 --json > bbr-p8.json

# Observe retransmission and RTT behaviour
ss -tinp
nstat -az | egrep 'TcpRetrans|TcpExtTCPLostRetransmit|TcpExtTCPFastRetrans'

Report retransmits from the iperf3 JSON alongside throughput: a BBR run that gains 30% throughput at the cost of 5% retransmissions is telling you the link itself is the problem. Also test with a lossy path (tc qdisc add dev eth0 root netem loss 1%) — that is where the difference is largest and where the marketing claims come from.

Where BBR Is the Wrong Answer

  • Inside a data centre. Sub-millisecond RTTs mean the window was already large; BBR adds nothing and can increase burstiness into shallow switch buffers.
  • Behind a network that polices with strict token buckets. BBR's pacing probes will be shaped, and the resulting loss pattern can look worse than CUBIC's.
  • On hosts with millions of short connections. The benefit is per-flow throughput; connection-rate-limited workloads see no gain and pay the CPU cost of pacing.

Operational Checklist

  1. tcp_available_congestion_control lists bbr on both ends.
  2. fq is the default qdisc and tc qdisc show confirms it per interface.
  3. Socket buffers are sized for the BDP, not the distribution default.
  4. Before/after numbers captured with identical test parameters, including retransmits.
  5. Rollback is one sysctl: net.ipv4.tcp_congestion_control=cubic — no reboot needed.

Related reading: TCP tuning for high bandwidth-delay products, simulating latency and loss with tc netem and iperf3 bandwidth testing methodology.

原文链接:https://www.man7.org/linux/man-pages/man7/tcp.7.html