IPv6 Path MTU Discovery and ICMPv6 Packet Too Big - 夜莺博客

IPv6 Path MTU Discovery and ICMPv6 Packet Too Big

IPv6 has no in-path fragmentation: routers never fragment a packet they cannot forward, they drop it and return an ICMPv6 Packet Too Big (type 2) message to the source. That design makes Path MTU Discovery mandatory, and it turns every ICMPv6 filter in your network into a potential outage. The signature failure is unmistakable once you know it — the TCP handshake completes because the SYN segments are small, then the connection hangs the moment real data flows. This guide explains how PMTUD works in IPv6, how to reproduce the black-hole condition, and how to fix it at the host, the firewall and the router.

How IPv6 PMTUD differs from IPv4

In IPv4, a router that hits a smaller MTU can fragment the packet and set the DF bit aside, which hides MTU mismatches until higher layers care. IPv6 removed router fragmentation entirely. A router that cannot forward a packet because it exceeds the outgoing link MTU must discard it and send an ICMPv6 type 2 message containing the MTU of the constricting link. The source reduces its path MTU estimate and retries. Nothing negotiates; everything depends on that message being delivered. RFC 8201 formalises the process, and it is explicit about the consequence: if ICMPv6 is filtered, traffic larger than the smallest hop MTU is silently black-holed.

Reproducing the black hole

Use a tunnel or a link with a deliberately smaller MTU. On Linux you can build the scenario locally:

# create a link with a reduced MTU to simulate a tunnel hop
sudo ip link add name dummy0 type dummy
sudo ip link set dummy0 mtu 1280 up
ip -6 addr add 2001:db8:99::1/64 dev dummy0

# discover the local PMTU to a destination
ping -6 -M do -s 1452 2001:db8:1::10

# watch the ICMPv6 messages arriving (type 2 = Packet Too Big)
sudo tcpdump -ni any 'icmp6 and ip6[40] == 2'

A size that succeeds without the ICMPv6 message reaching the sender will fail when the response is dropped — that is the whole failure mode. To see the PMTU cache on a Linux host:

ip -6 route get 2001:db8:1::10
ip -6 route show cache | grep mtu

On Cisco IOS and IOS-XE, the equivalents are:

show ipv6 interface GigabitEthernet0/0 | include MTU
show ipv6 route 2001:db8:1::10
show ipv6 neighbors
show ipv6 traffic | include "too big|unreachable"

Fix it at the source: canonical MTU values

The most common root cause is an overlay — VXLAN, IPsec or GRE — whose encapsulation overhead makes an interface 50 to 100 bytes smaller than the underlay. Two workable strategies:

  • Raise the underlay MTU (typically 9000+ on data centre fabrics) so the encapsulated frame still fits in the physical 1500-byte path toward the WAN edge.
  • Clamp TCP MSS so endpoints never emit a segment that would exceed the smaller MTU. This does not help UDP or control-plane protocols, so treat it as a complement to jumbo frames, not a replacement.
! Cisco IOS/IOS-XE MSS clamping on the tunnel interface
interface Tunnel100
 ipv6 mtu 1400
 ipv6 tcp adjust-mss 1360
!
! Linux, per-route MSS clamping for forwarded traffic
sudo ip6tables -t mangle -A FORWARD -p tcp --tcp-flags SYN,RST SYN \
  -j TCPMSS --clamp-mss-to-pmtu

Fix it in the filter: let ICMPv6 type 2 through

A firewall policy that permits "ICMPv6 error messages" is usually fine, but hand-written ACLs frequently permit only echo request or neighbour discovery. The minimum viable IPv6 control-plane policy is:

  • ICMPv6 type 1 destination unreachable
  • ICMPv6 type 2 packet too big
  • ICMPv6 type 3 time exceeded
  • ICMPv6 type 4 parameter problem
  • ICMPv6 types 133-137 (router solicitation, advertisement, neighbour solicitation, advertisement, redirect) on-link
! IOS-XE example: permit PMTUD on the edge ACL
ipv6 access-list IPV6-EDGE-IN
 permit icmp any any packet-too-big
 permit icmp any any destination-unreachable
 permit icmp any any time-exceeded
 permit icmp any any parameter-problem
 permit tcp any any established
!
interface GigabitEthernet0/0
 ipv6 traffic-filter IPV6-EDGE-IN in

Once that ACL is applied, confirm the counters increment and that oversized pings get a Packet Too Big instead of timing out. If the message is still not arriving, the filter lives somewhere else in the path — a NAT device that will not translate ICMPv6, a load balancer, or a cloud security group are all common culprits.

PLPMTUD: working without ICMPv6 at all

RFC 4821 defines Packetization Layer Path MTU Discovery, which probes upward with packets of increasing size and detects the ceiling from loss rather than from ICMPv6. Linux supports it through the TCP stack's probe settings, and QUIC implements its own version, which is why some applications work fine on paths where classic PMTUD fails. Enabling it makes an application resilient, but it does not fix the network — keep it as a mitigation, not a design.

# enable PLPMTUD probing on Linux endpoints
sudo sysctl -w net.ipv4.tcp_mtu_probing=2
sudo sysctl -w net.ipv4.tcp_base_mss=1024

Diagnostic recipe for a live incident

  1. Confirm the symptom shape: handshake succeeds, bulk transfer hangs. That alone points at MTU, not routing.
  2. Find the smallest MTU in the path by sending IPv6 pings with the DF-equivalent behaviour from the sender's own address, starting at 1500 bytes and stepping down.
  3. Capture icmp6 type 2 on both the sender and the intermediate router. If the router sends it but the sender never sees it, the filter is between them.
  4. Compare the sender's cached PMTU with the underlay MTU; a cache entry stuck at 1280 means an ICMPv6 message was never accepted.
  5. Check MSS clamping on tunnel interfaces, and make sure the TCP MSS clamping and PMTUD tuning notes match the encapsulation you actually run.

Related failure modes are worth having in your runbook: jumbo frame mismatches between switch hops, MTU problems inside VXLAN overlays, and the router advertisement parameters that set the interface MTU for hosts — covered with the rest of IPv6 neighbour discovery.

Best practices

  • Standardise MTU values per role: 1500 on access and WAN edges, 9000-9216 on fabric links, and document the encapsulation overhead for every overlay.
  • Never filter ICMPv6 error types 1-4 wholesale. Rate-limit them instead if abuse is the concern.
  • Set the IPv6 MTU explicitly on tunnel interfaces rather than relying on derived values.
  • Monitor ICMPv6 type 2 counts on the WAN edge — a spike means an MTU change was introduced somewhere.
  • Test with realistic packet sizes after every overlay change; a 1400-byte ping is not a substitute for a full-size TLS transfer.

原文链接:https://www.rfc-editor.org/rfc/rfc8201.html