CoreDNS in Kubernetes: Corefile, Plugins and Stub Zones - 夜莺博客

CoreDNS in Kubernetes: Corefile, Plugins and Stub Zones

CoreDNS is the default cluster DNS in Kubernetes and, more often than any other component, the cause of a mystery in production: services resolve slowly, some names fail intermittently, and a rollout stalls because a pod cannot reach the API server by name. It is also the easiest cluster component to reason about once you understand one fact — the order of plugins in the Corefile is not the execution order. This guide covers the Corefile model, the plugins that matter, and the standard customisations.

Where the Configuration Lives

kubectl get configmap coredns -n kube-system -o yaml
kubectl -n kube-system get deploy coredns
kubectl -n kube-system logs deploy/coredns -f

Configuration is a single key named Corefile in the coredns ConfigMap in kube-system. The reload plugin watches it (checksum polled roughly every 30 seconds) and hot-reloads, so a restart is often unnecessary — but allow up to two minutes and verify before concluding the change did nothing.

The Default Corefile

.:53 {
    errors
    health {
        lameduck 5s
    }
    ready
    kubernetes cluster.local in-addr.arpa ip6.arpa {
        pods insecure
        fallthrough in-addr.arpa ip6.arpa
        ttl 30
    }
    prometheus :9153
    forward . /etc/resolv.conf {
        max_concurrent 1000
    }
    cache 30
    loop
    reload
    loadbalance
}

Server Blocks and Zones

Each top-level block declares a zone (with . meaning the DNS root), an optional port, and an ordered plugin list. When a query arrives, the block with the longest matching zone wins. Multiple blocks are how you build split-horizon behaviour.

{$ENV_VAR} substitution is available at parse time, which is handy for templating a ConfigMap across clusters with Helm.

Execution Order Is Compiled In

CoreDNS has a compile-time plugin.cfg that fixes the processing sequence regardless of how you order the names. In the default Kubernetes build it is effectively:

errors → log → rewrite → hosts → kubernetes → autopath
       → forward → cache → loop → loadbalance

So writing cache before forward changes nothing about when caching happens. What matters is which plugins you include; excluding one removes it from the chain entirely.

Plugins That Matter in Practice

  • kubernetes — authoritative for cluster.local, in-addr.arpa and ip6.arpa; watches Services, EndpointSlices and Pods. Enable pods insecure if you need 10-244-1-5.default.pod.cluster.local style records.
  • forward — proxies everything else upstream; up to 15 upstream servers, and it speaks DNS-over-TLS.
  • cache — in-memory response cache, default capacity 9984 entries, separate TTLs for success and NXDOMAIN. Set a shorter window than the TTL when rapid service changes matter.
  • loop — detects CoreDNS forwarding to itself (for example via systemd-resolved) and deliberately exits so the Deployment restarts and the failure is visible.
  • loadbalance — randomises A/AAAA/MX record order so Kubernetes' DNS-based load balancing is not a single fixed answer.
  • prometheus — exposes metrics on :9153; scrape it.

Stub Domains: Send Internal Names to the Right Resolver

This is the most common real-world Corefile edit. Add a server block per corporate suffix:

.:53 {
    errors
    health { lameduck 5s }
    ready
    kubernetes cluster.local in-addr.arpa ip6.arpa {
        pods insecure
        fallthrough in-addr.arpa ip6.arpa
        ttl 30
    }
    prometheus :9153
    forward . /etc/resolv.conf { max_concurrent 1000 }
    cache 30
    loop
    reload
    loadbalance
}

corp.example.com:53 {
    errors
    cache 30
    forward . 10.10.0.53 10.10.0.54
}

consul.local:53 {
    errors
    cache 30
    forward . 10.150.0.1
}

Ordering matters when suffixes overlap: a more specific zone must be its own block, and the query goes to whichever block matches longest.

Per-Pod and Static Overrides

# hosts plugin: pin a name during a migration
hosts /etc/coredns/hosts.db {
    fallthrough
}

# Custom search path and upstreams for one workload
spec:
  dnsPolicy: None
  dnsConfig:
    nameservers: ["10.96.0.10"]
    searches:
      - myapp.svc.cluster.local
      - svc.cluster.local
    options:
      - name: ndots
        value: "2"

Lowering ndots is one of the few legitimate large wins in cluster DNS: with the default of 5, a name with fewer than five dots is tried against every search suffix first, multiplying query volume.

NodeLocal DNSCache: The Scaling Fix

Without a node-local cache, every pod DNS query is DNATed to a CoreDNS pod that may be on another node. At high query rates this produces conntrack table pressure (UDP entries have no close event and sit for 30 seconds), DNAT races and packet drops.

# Pod -> 169.254.20.10 (link-local, node-local) -> node-local-dns DaemonSet
#   cluster.local miss -> TCP -> kube-dns VIP -> CoreDNS
#   external miss      -> upstream nameservers

NodeLocal DNSCache has been GA since Kubernetes 1.18 and is the standard remedy for intermittent DNS timeouts in large clusters.

Troubleshooting Order

kubectl -n kube-system get pods -l k8s-app=kube-dns
kubectl -n kube-system logs -l k8s-app=kube-dns --tail=50
kubectl get configmap coredns -n kube-system -o yaml
kubectl run dnstest --rm -it --image=busybox:1.36 --restart=Never -- nslookup kubernetes.default
kubectl run dnstest2 --rm -it --image=busybox:1.36 --restart=Never -- nslookup example.com

If cluster.local fails, look at the kubernetes plugin and the API-server watch. If only external names fail, look at forward and the pod's /etc/resolv.conf. If the CoreDNS pod is in CrashLoopBackOff, suspect a broken Corefile or a loop detection — the logs name it explicitly.

Related Reading

Deeper dives on the same topics from our archive:

原文链接:https://jorijn.com/en/knowledge-base/kubernetes/networking/coredns-kubernetes-architecture-configuration