Tailscale Mesh VPN: WireGuard Without the Config - 夜莺博客

Tailscale Mesh VPN: WireGuard Without the Config

A WireGuard mesh is easy to design and tedious to operate: every node needs a key, every peer needs an endpoint, and every new host multiplies the configuration. Tailscale keeps the WireGuard data plane and replaces the key exchange and peer discovery with a coordination service, adding identity-based ACLs, MagicDNS names and subnet routers that expose an existing LAN to the tailnet. The result is a network you can join a laptop to in thirty seconds and then restrict with rules that read like an access policy rather than a routing table. This guide covers the topology, ACL design, subnet routing, exit nodes and the operational checks that tell you why traffic is or is not flowing.

Topology: how the mesh is built

  • Coordination server — distributes public keys and endpoints. It never sees traffic; data flows directly between nodes when NAT permits.
  • DERP relays — fall back paths for nodes that cannot establish a direct connection (symmetric NAT, restrictive firewalls). Traffic is still end-to-end encrypted.
  • Subnet router — a node that advertises a physical subnet into the tailnet, so devices on the LAN are reachable without installing Tailscale on each one.
  • Exit node — a node that carries all internet traffic for a client, giving you a stable egress IP for a home worker or a remote site.
# bring up a node and inspect the mesh
sudo tailscale up --advertise-routes=10.20.0.0/24 --accept-dns=true
tailscale status
tailscale netcheck          # NAT type, DERP latency, UDP support
tailscale ip -4
tailscale ping server-a     # shows whether the path is direct or relayed

tailscale ping is the fastest diagnostic in the product: it reports via DERP or direct plus the round-trip time. A node that is permanently relayed has a NAT or firewall problem, and latency will be measurably worse in both directions.

Subnet routing for existing infrastructure

# on a Linux host inside the data centre
echo "net.ipv4.ip_forward=1" | sudo tee /etc/sysctl.d/99-tailscale.conf
echo "net.ipv6.conf.all.forwarding=1" | sudo tee -a /etc/sysctl.d/99-tailscale.conf
sudo sysctl -p /etc/sysctl.d/99-tailscale.conf
sudo tailscale up --advertise-routes=10.20.0.0/24,10.20.1.0/24

# on the admin console: approve the routes, or do it via the CLI/API
tailscale route list

Advertised routes must be approved before they take effect — an unapproved route is the most common reason "the subnet router does not work". If the LAN has devices with their own firewall, remember that the subnet router forwards packets but does not bypass host firewalls; you still need the destination to accept traffic from the tailnet range (100.64.0.0/10 by default).

Access control: think in identity, not IP

// ACL policy (simplified)
{
  "groups": {
    "group:ops":    ["alice@example.com", "bob@example.com"],
    "group:devs":   ["group:devs-contractors"],
    "tag:server":   ["tag:prod-web"]
  },
  "tagOwners": {
    "tag:prod-web": ["group:ops"],
    "tag:router":   ["group:ops"]
  },
  "acls": [
    {"action": "accept", "src": ["group:ops"],   "dst": ["*:*"]},
    {"action": "accept", "src": ["group:devs"],  "dst": ["tag:prod-web:443,8080"]},
    {"action": "accept", "src": ["tag:prod-web"], "dst": ["tag:db:5432"]},
    {"action": "accept", "src": ["*"],           "dst": ["tag:router:53"]}
  ],
  "ssh": [
    {"action": "accept", "src": ["group:ops"], "dst": ["tag:server"], "users": ["root", "ubuntu"]}
  ]
}

Two design habits keep ACLs maintainable: tag servers rather than referencing hostnames, and give each service a narrowly scoped destination port list. Tailscale ACLs are default-deny, so a missing rule looks exactly like a broken tunnel — check the policy before you start debugging WireGuard. The same default-deny thinking applies to container workloads; the pattern is described in the NetworkPolicy default-deny guide.

# does the ACL actually allow this path?
tailscale ping db-01
tailscale whois 10.20.1.15        # which node and user owns this address
tailscale debug prefs | jq '.Routes, .ExitNodeID'
tailscale serve status            # if you published a service to the tailnet

MagicDNS and name resolution

With MagicDNS on, every node gets a name under your tailnet domain, which removes the need to remember 100.x addresses. Set your split-DNS rules deliberately: send your internal domain to the domain controller or a private resolver, and keep public names going to the standard resolvers. If a node must not use the tailnet resolver — a DNS server itself, for example — disable DNS for that node rather than fighting with the search domain.

sudo tailscale up --accept-dns=false          # keep the node's own resolver
sudo tailscale up --accept-routes=true        # accept advertised subnets from others
sudo tailscale set --exit-node=exit-gw-01     # route all internet traffic through a node
sudo tailscale set --exit-node=                # turn it off again

Failover paths that matter

  • Run two subnet routers for the same prefix where the LAN supports it, so one host failure does not cut off the site.
  • Keep at least one exit node in a well-connected data centre; home broadband uplinks lose far more traffic than a colocation link, and the client will not notice it is relayed unless you check.
  • Pin a subset of nodes to a specific DERP region only when a regulatory requirement demands it, since forcing a relay usually increases latency.
  • Document a break-glass path: if the coordination service is unreachable and nodes cannot refresh keys, existing sessions keep working but new ones fail. Know which local credentials still get you to a console.

Verification and troubleshooting

tailscale status --json | jq '.Self, .Peer[] | {HostName, Online, Relay, CurAddr}'
tailscale netcheck --format=json | jq '.UDP, .RegionLatency'
sudo tailscale --socket=/var/run/tailscale/tailscaled.sock debug netmap | jq '.Peers | keys | length'
journalctl -u tailscaled --since "10 min ago" | tail -30
  • Node online, no connectivity: ACL denies the port, or the remote host firewall drops the 100.x source range.
  • Traffic never takes the direct path: UDP is blocked on one side. Check netcheck; if UDP is unavailable, expect DERP relay and higher latency.
  • Subnet unreachable from a laptop: the client is not accepting routes (--accept-routes), or the subnet router's routes were never approved.
  • DNS works for some names only: split-DNS or search-domain overlap with the corporate resolver.

Where it fits alongside other VPNs

Tailscale is at its best for development, remote administration and connecting a handful of sites without a full SD-WAN. For permanent site-to-site links carrying production traffic, a classic IPsec or WireGuard tunnel between routers remains cheaper per byte and easier to monitor — the WireGuard site-to-site configuration and the IPsec VTI versus policy-based comparison cover those designs. When MTU problems appear on either, the WireGuard MTU troubleshooting notes explain the arithmetic that underlies both.

Operational checklist

  • Enable key expiry deliberately; long-lived untagged nodes are a standing risk.
  • Tag servers at provisioning time, not manually afterwards — tags cannot be edited once applied to some node types, only re-issued.
  • Review the ACL diff in version control; treat the policy file as production code.
  • Alert on nodes that have been offline longer than expected, and on relays appearing where direct paths used to work.
  • Test the failover of subnet routers quarterly, with a real client, from a different network.

原文链接:https://tailscale.com/kb/1151/what-is-tailscale