Linux Network Namespaces and veth Pairs Explained - 夜莺博客

Linux Network Namespaces and veth Pairs Explained

Every container and every Kubernetes pod gets its own network stack because of network namespaces, and the fastest way to stop being mystified by container networking is to build one by hand. A network namespace is a complete copy of the network stack - interfaces, routing tables, firewall rules - and as long as a process runs inside it, it sees only that copy.

Creating namespaces and wire them with a veth pair

A veth device is a virtual Ethernet interface that always comes in pairs: whatever enters one end emerges from the other. That makes it the equivalent of a virtual patch cable between namespaces.

sudo ip netns add ns-alpha
sudo ip netns add ns-beta

sudo ip link add veth-a type veth peer name veth-b
sudo ip link set veth-a netns ns-alpha
sudo ip link set veth-b netns ns-beta

sudo ip netns exec ns-alpha ip addr add 10.0.1.1/24 dev veth-a
sudo ip netns exec ns-beta  ip addr add 10.0.1.2/24 dev veth-b

sudo ip netns exec ns-alpha ip link set lo up
sudo ip netns exec ns-alpha ip link set veth-a up
sudo ip netns exec ns-beta  ip link set lo up
sudo ip netns exec ns-beta  ip link set veth-b up

sudo ip netns exec ns-alpha ping -c 3 10.0.1.2

Two details account for most first-attempt failures. The loopback interface inside a new namespace starts down and must be brought up explicitly; and each namespace has its own routing table, so if the ping fails after the addresses are set, check that a route for the peer subnet exists on both sides.

Inspecting a namespace

ip netns list
sudo ip netns exec ns-alpha ip addr
sudo ip netns exec ns-alpha ip route
sudo ip netns exec ns-alpha ss -lntup
sudo ip netns pids ns-alpha

The last command is the useful one in practice: it maps a namespace back to the processes trapped inside it, which is how you identify which container owns a mysterious interface on the host.

Scaling up with a Linux bridge

Pairing namespaces directly works for two, but grows into a full mesh. A Linux bridge is a virtual switch: create it once, attach one end of each namespace's veth pair to it, and give the bridge the address you want the namespace to route through.

sudo ip link add br-lab type bridge
sudo ip link set br-lab up
sudo ip addr add 10.0.2.1/24 dev br-lab

# wire ns-alpha to the bridge
sudo ip link add veth-alpha type veth peer name veth-alpha-br
sudo ip link set veth-alpha netns ns-alpha
sudo ip link set veth-alpha-br master br-lab
sudo ip link set veth-alpha-br up
sudo ip netns exec ns-alpha ip addr add 10.0.2.10/24 dev veth-alpha
sudo ip netns exec ns-alpha ip link set veth-alpha up
sudo ip netns exec ns-alpha ip route add default via 10.0.2.1

Notice that the host end of the pair gets no IP address at all - it is a Layer 2 port on the bridge, and connectivity is provided by the bridge interface. That is exactly the model Docker uses for its default bridge network, and Kubernetes CNI plugins do the same thing with additional routing policy.

Reaching the outside world

# enable forwarding
sudo sysctl -w net.ipv4.ip_forward=1

# masquerade traffic leaving via the host uplink
sudo iptables -t nat -A POSTROUTING -s 10.0.2.0/24 -o eth0 -j MASQUERADE

# let the return direction through
sudo iptables -A FORWARD -i br-lab -j ACCEPT

Without the masquerade rule, packets from the namespace leave the host with a private source address the outside network cannot route back. Without the FORWARD rule, the host firewall drops the traffic even though forwarding is enabled - a frequent cause of "DNS works from the host but not from the container".

Cleanup behaviour worth knowing

Deleting a namespace deletes every interface inside it, including one end of any veth pair, and the kernel removes the orphaned peer automatically. That is why an interface name can disappear from the host after a container is removed, and why recreating it requires recreating the pair rather than bringing an existing interface up. For the traffic shaping and queuing side of the same stack, see Linux tc HTB traffic shaping and ethtool tuning; for container network plumbing at scale, see SR-IOV, Multus and DPDK on Kubernetes.

原文链接:https://labs.iximiuz.com/tutorials/container-networking-from-scratch