NIC Offloads: TSO, GSO, GRO and Checksum Tuning - 夜莺博客

NIC Offloads: TSO, GSO, GRO and Checksum Tuning

Offloads exist because the CPU is bad at things the NIC is good at. Checksum generation, packet segmentation and receive-side aggregation all move work off the kernel and onto hardware, which is why a 25G interface can saturate on a modest CPU. They also change what you see on the wire, which is why so many packet-capture and virtualisation problems are solved by turning one of them off.

The features, in the order they matter

ethtool -k eth0
! Features for eth0:
! rx-checksumming: on          ! verify checksums in hardware
! tx-checksumming: on          ! generate checksums in hardware
!   tx-checksum-ip-generic: on
! scatter-gather: on           ! build packets from multiple buffers (SG)
! tcp-segmentation-offload: on ! TSO - NIC splits large TCP writes
!   tx-tcp-segmentation: on
! generic-segmentation-offload: on  ! GSO - software fallback/complement to TSO
! generic-receive-offload: on  ! GRO - merge received segments
! large-receive-offload: off   ! LRO - hardware merge (use GRO instead)
! rx-vlan-offload: on
! tx-vlan-offload: on
! ntuple-filters: off
! rx-hashing: on               ! RSS/flow hashing
  • TSO / GSO — the kernel hands the NIC a large TCP buffer and the NIC chops it into MTU-sized segments. GSO is the software-generalised equivalent, used when the NIC cannot do it for a particular encapsulation. Together they cut per-packet CPU cost dramatically.
  • GRO / LRO — the receive-side mirror image: merge consecutive segments of the same flow before handing them up the stack. GRO is the modern, software-friendly version; LRO is hardware and can break forwarding, so it is usually off.
  • Checksum offload — the NIC computes or verifies TCP/UDP checksums. Nearly always a win, except when something downstream sees the packets before the NIC has finished.
  • Scatter-gather — allows the NIC to transmit across non-contiguous buffers; TSO generally depends on it.

When to turn them off

# disable just what you need, on the interface in question
sudo ethtool -K eth0 rx off tx off            # checksum offload
sudo ethtool -K eth0 tso off gso off          # segmentation
sudo ethtool -K eth0 gro off lro off          # receive aggregation

# inspect the result
ethtool -k eth0 | grep -E "checksum|segmentation|offload"

The scenarios where disabling pays:

  1. Packet capture looks wrong. tcpdump on a host with TX checksum offload enabled shows "incorrect" checksums on outgoing packets, because the capture point is before the NIC fixes them. This is normal, not a fault — but it makes every capture look broken, so disabling TX checksum on the capture interface removes the ambiguity.
  2. Bridging, virtual switching or NFV. A packet that is handed to another kernel path (a bridge, Open vSwitch, a KVM guest) may be validated before the NIC completes the offload. Symptom: packets dropped with checksum errors only on bridged paths.
  3. VM-to-VM or container-to-container traffic where the offload is not end to end. Virtio advertises the features it supports; enable only those consistent across the path.
  4. LRO with an IP forwarding workload. Merged packets with inconsistent headers break forwarding and are a common cause of throughput that is oddly lower than expected.

Making the change persistent

# systemd-networkd does not expose every ethtool feature, so use a dispatcher script
sudo nano /usr/lib/networkd-dispatcher/routable.d/10-offloads
#!/bin/sh
ethtool -K eth0 rx off tx off
sudo chmod +x /usr/lib/networkd-dispatcher/routable.d/10-offloads

# NetworkManager
nmcli connection modify eth0 ethtool.feature-tso off

Whatever mechanism you use, verify after a reboot: ethtool -k eth0 is the only proof that the setting survived. A dispatcher script in the wrong directory (or without the execute bit) fails silently.

Diagnosing offload-related problems

ethtool -S eth0 | grep -iE "tso|gro|drop|err|restart"
ip -s link show eth0
nstat -az | grep -iE "TcpExt|IpExt" | head
dmesg | grep -i eth0 | tail

Useful signals: rx_errors climbing while rx_dropped stays flat points at the physical layer, not offloads. A large tx_restart_queue or tx_busy count suggests queue contention and a case for adjusting queue counts rather than features. And a capture that disagrees with the receiving application is almost always a capture-point problem.

Related: softnet_stat, RPS and backlog tuning for the receive path once packets reach the stack, and TCP MSS clamping and PMTUD for the MTU problems that segmentation offload tends to mask until a tunnel is introduced.

原文链接:https://michael.mulqueen.me.uk/2018/08/disable-offloading-netplan-ubuntu/