Anycast DNS: Design, Routing and Failure Modes - 夜莺博客

Anycast DNS: Design, Routing and Failure Modes

Anycast means advertising the same address block from many locations and letting BGP pick the nearest instance for each client. For DNS it is close to an ideal fit: the service is small, stateless per query, and latency-sensitive, and the same address appearing everywhere absorbs a distributed denial-of-service attack by spreading it across the whole footprint. It is also unforgiving of design errors. A withdrawn prefix, an over-eager health check or a route leak can take a whole region offline in ways that are hard to see from any single vantage point.

This article covers the design, the routing mechanics, and the failure modes worth testing before you are relying on it.

Why Anycast for DNS

  • Latency. The nearest instance answers, so a resolver in Frankfurt is not waiting on a server in Virginia.
  • Attack distribution. A volumetric flood is absorbed by every site simultaneously, and no single node sees the full rate.
  • Availability without complex failover. If a site goes down and its prefix is withdrawn, BGP reconverges and clients reach the next site. No DNS-level failover logic is required.
  • Simplified client configuration. One address to publish, everywhere.

The origin of the technique as a published practice is RFC 3258, which describes distributing authoritative name servers via shared unicast addresses - the terminology is worth keeping in mind, because what you are building is a shared unicast address, not a multicast group.

Design Rules

  1. Every node must be able to answer every query. Zone data must be identical, zone transfer must complete before a node takes traffic, and no node-specific data may influence an answer. If one node has stale data, anycast converts a replication problem into a correctness problem that only affects part of the world.
  2. Keep answers idempotent. Do not let a node's answer depend on its own address, its own view of the internet, or the time since it booted. Geographic answers are a legitimate exception, but they add per-node divergence and should be a deliberate decision with monitoring.
  3. Aggregate the prefixes you announce. IPv4 anycast prefixes are generally not accepted longer than /24. IPv6 needs at least a /48 to be routable in the global table. Plan the allocation around those limits rather than discovering them at launch.
  4. Announce from a single origin AS where possible. Announcing the same prefix from multiple ASNs creates path selection behaviour that is harder to reason about and can produce intermittent traffic shifts between sites for the same client.
  5. Do not rely on the prefix alone for load distribution. BGP has no concept of node capacity. Two sites with different capacity will not receive proportional traffic purely because of topology.

Health-Based Injection

The control that makes anycast useful is the ability to stop advertising a prefix when the service behind it is broken. Without it, BGP happily delivers traffic to a node whose DNS daemon is dead.

# a simple injector using ExaBGP: announce only while the local check passes
# /etc/exabgp/anycast.conf
neighbor 192.0.2.1 {
   router-id 203.0.113.10;
   local-address 203.0.113.10;
   local-as 64500;
   peer-as 64501;
   process healthcheck {
      run /usr/local/bin/check-dns.sh;
   }
}

# check-dns.sh - exit 0 means "announce", non-zero means "withdraw"
#!/bin/sh
dig +short +time=1 +tries=1 @127.0.0.1 example.com SOA >/dev/null || exit 1
# also verify the zone version is current
dig +short @127.0.0.1 version.bind chaos txt | grep -q "$EXPECTED_VERSION" || exit 1
exit 0

Two rules make this safe. First, the health check must test the actual service the client will use - query the local daemon over the real listening address, not a management socket. Second, add dampening so a flapping check does not churn the global routing table; a node that withdraws and re-announces every thirty seconds is worse than a node that stays quiet.

Where the health check drives a static route rather than a BGP announcement, the same objective is achieved with route tracking, and the pattern in this path-monitoring test instance guide is a useful model for proving a path before advertising anything that depends on it.

The Catch-All Node

The single most valuable design element: at least one node that answers every anycast address you use, whether or not it is authorised to announce it. If a prefix leaks, or a site is isolated but still advertising, or an IX route is misconfigured, traffic still lands somewhere that can answer it.

Configure the catch-all with a binding for the anycast address on a loopback and let the routing announce it only from the sites that should:

# the service listens on the address everywhere
interfaces {
    lo0 {
        unit 0 {
            family inet { address 203.0.113.8/32; }
            family inet6 { address 2001:db8:dead::8/128; }
        }
    }
}
# only the injector decides whether this site advertises it

Failure Modes Worth Testing

Failure What happens What to test
Partial prefix withdrawal Some upstreams still point at the dead site; the rest move Withdraw from one upstream and measure how long the tail of clients keeps failing
Authoritative service down, prefix still announced Full outage for clients homed to that site Kill the daemon without touching BGP and confirm the health check withdraws within your target
Stale zone data on one node Inconsistent answers depending on where you are Delay a zone transfer deliberately and confirm monitoring notices divergence
MTU and fragmentation Large responses silently dropped; clients retry over TCP Query for a large RRset from every site and check truncation behaviour
UDP truncation and TCP fallback Slower answers, or failure where TCP/53 is blocked Confirm TCP/53 is permitted to every node, not just UDP/53
Route leak / hijack Traffic from a distant region lands on one node Monitor per-node query rates for anomalies that correlate with routing events
DNSSEC key rollover with partial propagation Validating resolvers get SERVFAIL for a subset of queries Roll keys with the same care as any multi-master service; the process is described in this DNSSEC signing and validation guide
Asymmetric capacity One node saturated while others idle Measure per-node QPS and query latency, not just aggregate

Monitoring Anycast Properly

The hard part is that you cannot easily observe which node a given resolver reached. What works:

  • Per-node query counters and latency histograms, exposed locally and scraped on the management network so a data-path problem does not hide the evidence.
  • Probes from many external vantage points to the anycast address, treating a latency shift as a signal that routing changed. Public measurement platforms are useful here precisely because they sit outside your network.
  • Zone-version compliance checks per node, alerting on divergence rather than on a failed transfer.
  • A cache-hit sanity check. A sudden move in cache hit ratio on one node usually means either a query pattern change or stale data.

If the authoritative service is built on a modern stack, much of this is built in - the operational surface for combining a high-performance authoritative server with distributed query handling is covered in this PowerDNS and dnsdist guide, and the encrypted-transport side that clients increasingly expect is in this encrypted DNS deployment guide.

Operating It

Anycast turns every routine change into a global one. A zone edit, a software upgrade or a key roll is applied to every node, and a mistake affects every client simultaneously. Sequence changes: one node, verify, then the rest. And plan the maintenance path deliberately - a node that is being upgraded should not be receiving traffic, which means a coordinated withdrawal and re-announcement. Treating that as a maintenance event rather than an emergency is what keeps the announce/withdraw mechanism from becoming a source of instability, and the same logic that makes graceful shutdown work for ordinary BGP applies here, as described in this BGP graceful shutdown guide.

原文链接:https://www.rfc-editor.org/rfc/rfc3258.html