Cilium BGP Control Plane in Kubernetes Clusters - 夜莺博客

Cilium BGP Control Plane in Kubernetes Clusters

Most Kubernetes clusters make their pod network reachable by encapsulation: VXLAN or Geneve between nodes, NAT at the edge. That works until something outside the cluster needs to reach a pod address directly — a load balancer in hardware, a monitoring system, or an on-premises network with no route to the overlay. The Cilium BGP control plane takes the other approach: each node runs a BGP speaker and advertises the prefixes it owns to the ToR switch, so the pod network becomes a normal routed network. This article covers enabling the feature, writing the BGP configuration as Kubernetes resources, and verifying sessions and advertisements.

Enable the control plane

helm upgrade cilium oci://public.ecr.aws/eks/cilium/cilium \
  --namespace kube-system \
  --reuse-values \
  --set operator.rollOutPods=true \
  --set bgpControlPlane.enabled=true

kubectl -n kube-system get pods --selector=app.kubernetes.io/part-of=cilium

The roll-out flag matters: if BGP was not previously enabled, the operator must be restarted for the configuration to apply. The control plane is disabled by default, and a cluster upgraded with the flag set but without the operator restart will report no BGP resources and no peering.

Model the configuration as resources

Cilium splits BGP configuration into a cluster-level resource (which nodes peer with which ASN, using what credentials) and per-node or per-peer-group detail. A typical two-tier fabric uses a node-level peer with the ToR switch, and a pod or service advertisement to control what is exported.

apiVersion: cilium.io/v2
kind: CiliumBGPClusterConfig
metadata:
  name: cluster-bgp
spec:
  nodeSelector:
    matchLabels:
      kubernetes.io/os: linux
  bgpInstances:
  - name: "65000"
    localASN: 65000
    peers:
    - name: "tor-switch"
      peerASN: 64512
      peerAddress: 192.0.2.1
      peerConfigRef:
        name: "tor-peer-config"
---
apiVersion: cilium.io/v2
kind: CiliumBGPPeerConfig
metadata:
  name: tor-peer-config
spec:
  timers:
    holdTimeSeconds: 9
    keepAliveTimeSeconds: 3
  gracefulRestart:
    enabled: true
    restartTimeSeconds: 120
apiVersion: cilium.io/v2
kind: CiliumBGPAdvertisement
metadata:
  name: advertise-pod-cidr
  labels:
    advertise: bgp
spec:
  advertisements:
  - advertisementType: "PodCIDR"
    selector:
      matchExpressions:
      - { key: somekey, operator: NotIn, values: ['never-used-value'] }

The idle-looking selector on the advertisement is the documented idiom for "select everything": an empty selector may be interpreted as matching nothing, so a condition that is never true is used to match all pods.

Auto-discovery versus explicit peers

Maintaining a peer entry per node is tedious at scale, so Cilium supports default-gateway auto-discovery: the node learns its BGP peer from the default route, and the switch accepts dynamic peers from a range. On the switch side this looks like a listen range with a peer group, which is a common pattern on Cumulus, FRR and Arista platforms.

router bgp 64512
  neighbor CILIUM peer-group
  neighbor CILIUM local-as 65000 no-prepend replace-as
  bgp listen range 192.0.2.0/24 peer-group CILIUM

Note the local-as statement with no-prepend and replace-as. Without it the switch would prepend its own ASN to advertisements, and the cluster would see unexpected paths. Multi-homing with undiscovered gateways lets a node peer with two switches for redundancy, in which case each node advertises the same prefixes to both, and the fabric decides the path.

Verification

cilium bgp peers
cilium bgp routes advertised ipv4 unicast
cilium bgp routes available ipv4 unicast
kubectl get ciliumbgpclusterconfig

The peers output lists local AS, peer AS, peer address, session state, uptime and route counts per address family. A session stuck in idle or connect usually means the peer address is not reachable from the node, TCP 179 is filtered, or the ASNs disagree. Before blaming BGP, verify basic reachability: if the node cannot reach the management network at all, you will be debugging the wrong layer, and the node-level checks in Kubernetes node NotReady troubleshooting apply.

Operational considerations

  • Route scale: every node advertising a PodCIDR means the switch carries one route per node. At a few hundred nodes this is trivial; at thousands, aggregate or advertise per-rack ranges instead.
  • Failure behaviour: if a node fails, its PodCIDR is no longer advertised. Pods on that node are unreachable regardless, but if the failure only affects a subset of the cluster, traffic may still be routed toward a node whose pods cannot serve it. Combine BGP advertisement with node health checks in the fabric.
  • Service address pools: the same control plane can advertise service addresses as an alternative to a separate load balancer controller, which is compared in MetalLB layer 2 versus BGP address pools.

When the cluster's underlay is itself an overlay, keep an eye on encapsulation overhead and MTU; the same arithmetic applies as in Linux VXLAN VTEP configuration. The payoff for the effort is a cluster that the network team can see, monitor and route with the same tools they use for everything else.

原文链接:https://docs.cilium.io/en/stable/network/bgp-control-plane/bgp-control-plane-configuration/