Docker Swarm Overlay Networking and Service Deployment - 夜莺博客

Docker Swarm Overlay Networking and Service Deployment

Docker Swarm mode is still the shortest path from three Linux hosts to a self-healing
container cluster: no control-plane manifests, no etcd, and a declarative service model built
into the Docker CLI. Where teams get stuck is networking — overlay networks, the ingress
network, VIP-based load balancing, and the fact that publishing a port behaves differently from
a standalone container. This guide builds a three-node swarm, wires up overlay networking
properly, and covers the firewall rules Swarm silently requires.

Cluster Bring-Up

# On the first manager
docker swarm init --advertise-addr 10.10.30.10
#   Swarm initialized: current node (xxxx) is now a manager
#   Copy the join command it prints

# On the other managers (odd number, so 3 or 5)
docker swarm join-token manager      # gives you the manager join command

# On workers
docker swarm join-token worker
docker swarm join --token SWMTKN-1-xxxx 10.10.30.10:2377

# Verify
docker node ls
docker info | grep -i swarm

Use an odd number of managers. Three managers tolerate one failure; five tolerate two. Never
run two — a two-manager swarm has no quorum margin at all.

What Swarm Creates Automatically

Initialising a swarm creates two networks on every node:

  • ingress — an overlay network that handles published-port traffic and
    routing mesh. When any node receives a request on a published port, the ingress network hands
    it to IPVS for load balancing across the service's tasks.
  • docker_gwbridge — a local bridge that connects the overlay networks to the
    node's physical network.

If you create a service and do not attach it to a user-defined overlay, it joins the ingress
network by default. That is convenient for a single service and undesirable for anything you
want segmented.

Firewall Ports Swarm Needs

# Between swarm nodes
2377/tcp     cluster management (managers only)
7946/tcp     container network discovery
7946/udp     container network discovery
4789/udp     VXLAN overlay data path (ingress included)

# From clients
published service ports, e.g. 80/tcp, 443/tcp

# If nodes are not in the same broadcast domain, you must also allow
# IP protocol 50 (ESP) and 97 (VXLAN) for encrypted overlay networks.

A closed 7946/4789 is the reason a service looks healthy on one node and unreachable on
another. Check with:

docker network inspect ingress
ss -lunp | grep -E '7946|4789'
tcpdump -i any -n udp port 4789 -c 20

Overlay Networks the Right Way

# Application tier — reachable only by other swarm services
docker network create --driver overlay --attachable --subnet 10.20.0.0/24 app-net

# Data tier, encrypted
docker network create --driver overlay --opt encrypted --subnet 10.21.0.0/24 data-net

# Expose to the outside on a specific subnet
docker network create --driver overlay --attachable --subnet 10.22.0.0/24 edge-net

--attachable lets you add a plain standalone container or a debug container to
the overlay — extremely useful for troubleshooting, and the reason to include it even when you
think you do not need it.

The default subnet mask length for overlay networks is /24. If you need a different global
default, set --default-addr-pool-mask-length at swarm init time —
this cannot be changed afterwards without rebuilding the swarm. Plan your address pools before
you initialise.

Deploying a Service

docker service create \
  --name web \
  --replicas 3 \
  --network app-net \
  --publish published=8080,target=80 \
  --update-parallelism 1 \
  --update-delay 10s \
  --update-order stop-first \
  --restart-condition any \
  --constraint 'node.role == worker' \
  --limit-cpu 0.5 --limit-memory 256M \
  nginx:stable

docker service ls
docker service ps web
docker service logs -f web
docker service inspect --pretty web

Port publishing in Swarm is a two-part mapping: published=8080,target=80. Omit
the published port and the service is reachable only from other services on the same overlay —
which is usually what you want for a database tier.

# Deploy a stack instead
cat > stack.yml <<'EOF'
version: "3.8"
services:
  api:
    image: registry.local/api:1.4.2
    networks: [app-net, edge-net]
    deploy:
      replicas: 4
      placement:
        preferences:
          - spread: node.labels.dc
      update_config:
        parallelism: 1
        delay: 10s
        order: start-first
      restart_policy:
        condition: on-failure
        delay: 5s
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:8080/healthz"]
      interval: 10s
      timeout: 3s
      retries: 3
  cache:
    image: redis:7-alpine
    networks: [data-net]
networks:
  app-net:  {external: true}
  edge-net: {external: true}
  data-net: {external: true}
EOF

docker stack deploy -c stack.yml prod
docker stack services prod

update-order: start-first on a stateless service avoids a capacity dip during
rolling updates. Use stop-first when the new version cannot coexist with the old
one — for example a schema migration.

Load Balancing, Both Layers

  • External — routing mesh means every node accepts traffic on the published
    port and forwards it to a task, wherever that task runs.
  • Internal — the embedded DNS returns a virtual IP for the service name, and
    IPVS distributes connections across tasks. Using the service name rather than a container name
    is what makes the design self-healing.
  • DNS round-robin — query the tasks.<service> name to get
    individual task IPs when you need to address a specific replica.
docker exec -it $(docker ps -qf name=web) nslookup web
docker exec -it $(docker ps -qf name=web) nslookup tasks.web

Operational Notes

原文链接:https://docs.docker.com/engine/swarm/networking/