Paris Traceroute: Finding Paths in ECMP Networks - 夜莺博客

Paris Traceroute: Finding Paths in ECMP Networks

Traceroute was designed for a network where each destination had exactly one path. Modern networks have dozens, chosen per flow by a hardware hash, and that is why the output you collected can contain hops that no single packet ever traversed — a "false path" that wastes hours of investigation. Paris traceroute fixes the measurement method rather than the network, so that what you see is a path a real flow would follow. This article explains why classic traceroute breaks, how flow-consistent probing works, and how multipath discovery enumerates every path.

Why Classic Traceroute Is Unreliable Under ECMP

Classic traceroute varies something in each probe — the UDP destination port, or the ICMP sequence — and that variation feeds the router's load-balancing hash. A router hashing on the 5-tuple can therefore send probe 1 over path A, probe 2 over path B, and probe 3 with a TTL that expires on path C. The collected hops are a union of paths, and worse, a hop can appear as an apparent loop or a "missing" node when in fact no single flow ever took that sequence.

The visible symptoms of doing this wrong are familiar:

  • An R-hop that seems to loop back to the same router address twice, or a router that appears to forward to itself.
  • Latency jumps that cannot be reproduced by the application.
  • A path that differs between runs, minutes apart, with no network change.

The Paris Traceroute Idea

Flow-consistent probing keeps every field that a load balancer might hash on constant across all probes of one trace, and varies only the TTL. Two implementations exist:

  • UDP-based — fix the source and destination port (and IP addresses) and vary only TTL; or vary both ports together in a way that keeps the hash stable, depending on how routers hash.
  • ICMP-based — keep the ICMP identifier and sequence stable, vary TTL and vary the "payload" that some routers include in the hash.

The result is that all probes of a run follow the same path, and the trace becomes a truthful representation of one flow. In practice this immediately exposes the difference between "the path to this service" — a fiction — and "one of the paths my traffic takes" — a measurable reality.

# Flow-consistent, UDP-style probing (scapy-based examples)
pip install scapy

from scapy.all import IP, UDP, ICMP, sr1
def paris_trace(dst, dport=33434, sport=50000, maxttl=30):
    for ttl in range(1, maxttl + 1):
        pkt = IP(dst=dst, ttl=ttl) / UDP(sport=sport, dport=dport)
        r = sr1(pkt, verbose=0, timeout=2)
        if r is None:
            print(ttl, "*")
            continue
        print(ttl, r.src, r.type if r.haslayer(ICMP) else "")
        if r.src == dst or (r.haslayer(ICMP) and r.type == 3):
            break

Keeping sport and dport constant for the whole trace is the essential part: every probe hits the same ECMP branch on every hop.

Enumerating All Paths

A single flow-consistent trace proves one path works — not that all of them do. Multipath discovery answers the harder question: how many distinct paths exist, and does traffic use all of them?

# Free implementation (Debian/Ubuntu: apt install paris-traceroute)
paris-traceroute --algo=exhaustive 8.8.8.8
paris-traceroute -m 20 -q 3 --first-hop 3 10.20.30.40

# Practical variant: probe with many different flows and group by path signature
for p in $(seq 40000 40015); do
    traceroute -n -q 1 -m 20 -p $p -U 192.0.2.10 2>/dev/null | awk '{print $2, $3}' | tr '
' ' '
    echo
done | sort | uniq -c | sort -rn

The counting trick is the useful one in production: run traces across many flows, collapse each into a hop sequence, then count how many flows took each distinct sequence. If 400 out of 400 sampled flows take path A, you have a serious imbalance long before anyone notices a capacity problem. If a path you expected never appears, you have found a broken ECMP member without needing a maintenance window to investigate.

Reading the Results Sanely

  • Wildcard hops (*) are not proof of loss. Many routers rate-limit ICMP TTL-expired generation, so the last hop before a wildcard might be the real one. Re-probe with a lower rate before concluding anything.
  • Second-to-last hop latency is often the forward path plus a slow control-plane reply. Judge latency from the destination hop and from end-to-end application measurements, not from mid-path ICMP timestamps.
  • Asymmetry is normal. Return paths frequently differ from forward paths; a trace from the server towards the client is a separate measurement worth taking.
  • MPLS clouds hide structure. When the path disappears into a carrier MPLS domain, use ICMP extensions (traceroute -e with --icmp-extensions) or ask the carrier; LSP ping is only available to the operator.

Practical Workflow

  1. Reproduce the complaint with an actual application flow — record the 5-tuple.
  2. Run a flow-consistent trace with matching ports to see the path that flow takes.
  3. Run a multipath enumeration across many flows to see how many paths exist and their distribution.
  4. Identify the bad path, then map it to the underlay with routing and link-state commands rather than more traceroute.
  5. Re-run both measurements after the fix; a changed distribution is the evidence that something actually changed.

Related reading: mtr for packet loss and latency, loopback interface design and iperf3 throughput testing methodology.

原文链接:https://www.kentik.com/blog/the-power-of-paris-traceroute-for-modern-load-balanced-networks/