PXE and UEFI HTTP Boot Server Setup with dnsmasq - 夜莺博客

PXE and UEFI HTTP Boot Server Setup with dnsmasq

Bare-metal provisioning still starts with network boot, and the goalposts have moved:
UEFI clients expect HTTP Boot as well as TFTP, architecture-specific boot files must be served
to the right platform, and a server that works for one firmware version fails on the next.
This article builds a PXE and UEFI HTTP Boot environment on Linux with dnsmasq, a TFTP root and
an HTTP root, covering both the integration mode (dnsmasq also serving DHCP) and proxy DHCP mode
(dnsmasq supplementing an existing DHCP server) — the latter being the case in most real
networks.

What a Client Actually Asks For

Understanding the handshake removes most of the guesswork:

  1. The client broadcasts a DHCP request with option 60 (vendor class) set to
    PXEClient — for UEFI, PXEClient:Arch:00007:UNDI:003000 or similar.
  2. The DHCP server (or proxy) returns an IP plus options 66/67 (next-server, boot filename)
    or option 43 sub-options.
  3. The client downloads the boot file over TFTP or, with UEFI HTTP Boot, over HTTP.
  4. For UEFI HTTP Boot, the client caches the URL in NVRAM and may try it again on the next
    boot even if your DHCP answers differently.

That last point explains the classic "it works, then a different image boots" problem:
stale NVRAM boot entries.

Layout

/srv/tftp/                              # TFTP root
    pxelinux.0
    ldlinux.c32
    grubx64.efi                         # UEFI x86-64
    shimx64.efi
    grubaa64.efi                        # UEFI ARM64
    pxelinux.cfg/
        default
       01-aa-bb-cc-dd-ee-ff             # per-MAC boot entry
    images/
        vmlinuz
        initrd.img

/srv/http/                              # HTTP root for UEFI HTTP Boot
    uefi/
        grubx64.efi
        shimx64.efi
    images/
        vmlinuz
        initrd.img
sudo apt install -y dnsmasq tftpd-hpa syslinux syslinux-common grub-efi-amd64-signed nginx

Step 1: TFTP Server

# /etc/default/tftpd-hpa
TFTP_USERNAME="tftp"
TFTP_DIRECTORY="/srv/tftp"
TFTP_ADDRESS=":69"
TFTP_OPTIONS="--secure --create"

sudo systemctl enable --now tftpd-hpa
sudo systemctl status tftpd-hpa

Copy the boot files you intend to serve into /srv/tftp, and verify the path is
readable by the tftp user:

sudo -u tftp cat /srv/tftp/grubx64.efi > /dev/null && echo "tftp can read boot files"
# Firewall on the provisioning server
sudo ufw allow 69/udp  comment "TFTP"
sudo ufw allow 80/tcp  comment "HTTP boot"
sudo ufw allow 67/udp  comment "DHCP/PXE (only if dnsmasq runs DHCP)"
sudo ufw allow 4011/udp comment "Proxy DHCP"

Step 2: dnsmasq in Integration Mode

# /etc/dnsmasq.d/pxe.conf
interface=eth0
bind-interfaces

# Serve DHCP itself
dhcp-range=10.10.30.100,10.10.30.200,12h
dhcp-option=3,10.10.30.1
dhcp-option=6,10.10.30.1

# PXE
enable-tftp
tftp-root=/srv/tftp
dhcp-boot=tag:!ipxe,tag:bios,pxelinux.0
log-dhcp

# Architecture-specific boot files
dhcp-match=set:bios,option:client-arch,0
dhcp-match=set:efi64,option:client-arch,7
dhcp-match=set:efi64,option:client-arch,9
dhcp-match=set:efiarm64,option:client-arch,11

dhcp-boot=tag:bios,pxelinux.0
dhcp-boot=tag:efi64,grubx64.efi
dhcp-boot=tag:efiarm64,grubaa64.efi

The architecture matching is what makes one server work for a mixed fleet. Option 60
architecture codes you will see in practice: 0 BIOS x86, 7 and
9 UEFI x86-64, 11 UEFI ARM64.

sudo systemctl restart dnsmasq
sudo journalctl -u dnsmasq -f      # watch the DHCP exchange live while a test client boots

log-dhcp plus journalctl -f is the single most useful debugging
combination in this build. You will see exactly which tag matched and which boot file was
offered.

Step 3: Proxy DHCP for an Existing Network

If your production DHCP server must keep serving addresses, run dnsmasq in proxy mode so it
only supplies boot information:

# /etc/dnsmasq.d/pxe-proxy.conf
port=0
interface=eth0
bind-interfaces
log-dhcp

dhcp-range=10.10.30.0,proxy

enable-tftp
tftp-root=/srv/tftp

dhcp-match=set:bios,option:client-arch,0
dhcp-match=set:efi64,option:client-arch,7
dhcp-match=set:efi64,option:client-arch,9
dhcp-match=set:efiarm64,option:client-arch,11

dhcp-boot=tag:bios,pxelinux.0
dhcp-boot=tag:efi64,grubx64.efi
dhcp-boot=tag:efiarm64,grubaa64.efi

port=0 stops dnsmasq acting as a DNS server, and dhcp-range with
the proxy keyword makes it a PXE proxy — no address allocation. The existing DHCP
server still needs to hand out an IP; the proxy only supplements options 66/67. Some network
gear (and Windows DHCP) requires option 60 to be present to respond to PXE clients at all, so
verify the existing server is PXE-aware.

Step 4: UEFI HTTP Boot

HTTP Boot removes TFTP's slow transfer and small-file limitations, which matters for large
images. Serve the boot files over HTTP and point the client at them:

# /etc/nginx/sites-available/boot
server {
    listen 80 default_server;
    root /srv/http;
    autoindex on;

    location /uefi/ {
        types { application/octet-stream efi; }
        default_type application/octet-stream;
    }
}

sudo ln -s /etc/nginx/sites-available/boot /etc/nginx/sites-enabled/boot
sudo nginx -t && sudo systemctl reload nginx
# Serve the UEFI boot file over HTTP via dnsmasq option 67 (URL form)
dhcp-option=tag:efi64,option:vendor-class,"HTTPClient"
dhcp-boot=tag:efi64,http://10.10.30.10/uefi/grubx64.efi

Test the HTTP side independently of DHCP before blaming the network:

curl -I http://10.10.30.10/uefi/grubx64.efi
curl -I http://10.10.30.10/images/vmlinuz

Step 5: Boot Menu

# /srv/tftp/pxelinux.cfg/default   (BIOS)
DEFAULT menu.c32
PROMPT 0
TIMEOUT 300
ONTIMEOUT local

LABEL local
  MENU LABEL ^Boot from local disk
  LOCALBOOT 0

LABEL install-rocky9
  MENU LABEL ^Rocky Linux 9 network install
  KERNEL images/vmlinuz
  APPEND initrd=images/initrd.img ip=dhcp inst.repo=http://10.10.30.10/rocky9

LABEL rescue
  MENU LABEL ^Rescue shell
  KERNEL images/vmlinuz
  APPEND initrd=images/initrd.img ip=dhcp inst.rescue
# /srv/http/uefi/grub.cfg   (UEFI, served over HTTP too)
set timeout=10
set default=0

menuentry "Boot from local disk" { exit }

menuentry "Rocky Linux 9 network install" {
    linux  /images/vmlinuz ip=dhcp inst.repo=http://10.10.30.10/rocky9
    initrd /images/initrd.img
}

menuentry "Rescue shell" {
    linux  /images/vmlinuz ip=dhcp inst.rescue
    initrd /images/initrd.img
}

Always include a local-disk entry and set it as the timeout default. A PXE server that
always boots the installer will happily reinstall your production servers after an unattended
reboot.

Failure Modes Worth Knowing

  • Client never gets a boot file — check log-dhcp output. If no
    request appears, the client is on the wrong VLAN or the DHCP relay is not forwarding to the
    proxy. If the request appears but no boot file is offered, the client-arch tag did
    not match.
  • Endless "no boot filename received" — option 67 or dhcp-boot
    is missing for that architecture. Very common on ARM64 clients, which have their own
    architecture code.
  • Boots the wrong image or a stale OS — UEFI NVRAM has a cached HTTP Boot
    entry. Clear it (efibootmgr from a live system, or the firmware setup menu).
  • TFTP timeouts on large files — switch to HTTP Boot, or check MTU and
    intervening firewalls. TFTP is UDP and does not take well to packet loss.
  • Secure Boot enabled and nothing boots — you need a signed shim and a
    signed GRUB; serve shimx64.efi as the boot file and let it chain to
    grubx64.efi.
  • Works in the lab, not on the network — in proxy mode nothing happens
    unless the existing DHCP server is PXE-aware, or it is set to hand option 60 through.

Making It Repeatable

  • Per-MAC boot files (pxelinux.cfg/01-<mac>) let you control exactly which
    installer a server receives, which is how you avoid installing the wrong role on the wrong
    machine.
  • Keep kernel and initrd in one directory referenced by both TFTP and HTTP so a rebuild only
    happens once.
  • Provisioning is the first step of a larger pipeline: once the OS is up,
    Catalyst Center Plug and Play onboarding is the network-device equivalent, and Ansible AWX job templates and workflow configuration is a natural way to run the post-install configuration. If your DHCP layer needs hardening while you are here, check DHCP Option 82 and IP Source Guard.

原文链接:https://documentation.suse.com/sled/15-SP7/html/SLED-all/cha-deployment-prep-uefi-httpboot.html