KVM and libvirt: VM Management with virsh - 夜莺博客

KVM and libvirt: VM Management with virsh

KVM provides the hypervisor, but libvirt provides the management layer that makes it usable: a stable XML definition per guest, storage and network abstractions, and the virsh CLI that ties them together. That structure is what lets you script VM lifecycle operations, take a snapshot before a risky upgrade, and live-migrate a guest off a host you need to patch. This guide covers storage pools, bridged networking, the lifecycle commands you will use daily, snapshots, migration, and the console and log paths that turn a hung guest into a five-minute diagnosis.

Host preparation and resource model

# verify the host is actually capable of virtualisation
egrep -c '(vmx|svm)' /proc/cpuinfo
sudo apt install -y qemu-kvm libvirt-daemon-system libvirt-clients virtinst bridge-utils
sudo systemctl enable --now libvirtd
sudo usermod -aG libvirt,kvm $USER
virsh version
virsh nodeinfo
virsh capabilities | head -40

virsh nodeinfo reports the CPU model, memory and NUMA topology libvirt sees — the model it reports determines the guest CPU you can safely expose. On a mixed fleet, use host-model rather than host-passthrough so a guest can migrate between hosts with different CPU generations.

Storage pools: stop hand-managing image paths

virsh pool-define-as vmvol dir --target /var/lib/libvirt/images/vmvol
virsh pool-build vmvol
virsh pool-start vmvol
virsh pool-autostart vmvol
virsh pool-list --all
virsh vol-create-as vmvol web01.qcow2 60G --format qcow2
virsh vol-list vmvol
virsh vol-info --pool vmvol web01.qcow2

Pools give you virsh vol-* commands instead of raw qemu-img paths, and they are required if you ever want libvirt to handle allocation, deletion and resizing consistently. For an LVM pool the name changes but the interface does not:

virsh pool-define-as vgpool logical --source-name vg_vms --target /dev/vg_vms
virsh pool-start vgpool && virsh pool-autostart vgpool

Networking: a real bridge, not NAT defaults

# /etc/netplan/60-bridge.yaml  (Ubuntu, bridged on a physical NIC)
network:
  version: 2
  ethernets:
    eno1: { dhcp4: false }
  bridges:
    br0:
      interfaces: [eno1]
      addresses: [10.30.0.10/24]
      routes: [{ to: default, via: 10.30.0.1 }]
      parameters:
        stp: false
        forward-delay: 0
sudo netplan apply
ip -br addr show br0
bridge link show
virsh net-list --all

The default default network gives guests NAT addresses, which is fine on a laptop and wrong in a data centre: you lose inbound reachability and every guest hides behind the host's address. A Linux bridge puts guests directly on the LAN with their own addresses and MACs. Make sure the upstream switch port is a trunk carrying the guest VLANs, and that the bridge VLAN configuration matches — the same failure mode applies whether the host runs libvirt or Proxmox.

Create, start and console a guest

virt-install --name web01 --memory 8192 --vcpus 4 \
  --cpu host-model --machine q35 --os-variant ubuntu24.04 \
  --disk vol=vmvol/web01.qcow2,bus=virtio \
  --network bridge=br0,model=virtio \
  --graphics none --console pty,target_type=serial \
  --location 'https://mirror.example.com/ubuntu/dists/noble/main/installer-amd64/' \
  --extra-args 'console=ttyS0,115200n8 serial'

virsh list --all
virsh console web01          # exit with Ctrl+] 
virsh dominfo web01
virsh domiflist web01
virsh domblklist web01

Serial console access is not optional on a headless host — without it, a guest that fails to boot its own networking is unreachable and you are reduced to guessing. For windows guests use a VNC or SPICE display plus virtio-win drivers.

Lifecycle and configuration editing

virsh start web01
virsh shutdown web01          # graceful ACPI
virsh reboot web01
virsh suspend web01 && virsh resume web01
virsh destroy web01           # hard power off, last resort
virsh autostart web01

# edit the XML safely
virsh edit web01              # validates on save
virsh dumpxml web01 > web01.xml
virsh define web01.xml        # apply after external editing
virsh domstats web01 --balloon --vcpu

virsh define applies a new definition but does not change a running guest's CPU or memory device layout; those require a stop and start unless the change is supported live. Keep each guest's XML in version control so a rebuild is reproducible.

Snapshots: what they do and do not protect

virsh snapshot-create-as web01 snap-before-upgrade --disk-only --atomic
virsh snapshot-list web01
virsh snapshot-info web01 snap-before-upgrade
virsh blockcommit web01 vda --active --pivot    # merge back and remove dependence

Disk-only snapshots capture the virtual disk chain, not the guest's memory or the application's consistency. For anything with a database, quiesce the application or use the guest agent so the snapshot represents a recoverable state; a crash-consistent snapshot of a running database is a source of subtle corruption. An internal snapshot stores the delta inside the qcow2 file, which grows quickly and complicates backups — prefer external snapshots plus a proper backup from inside the guest.

Live migration and maintenance

# shared storage or a migration network is required
virsh migrate --live --verbose web01 qemu+ssh://node02/system
virsh migrate --live --persistent --undefinesource web01 qemu+ssh://node02/system

# host maintenance: drain guests without downtime
for vm in $(virsh list --name); do virsh migrate --live $vm qemu+ssh://node02/system; done
virsh node-memory-tune

Migration bandwidth is the limiting factor: a guest whose dirty page rate exceeds the link speed will never converge. Use a dedicated 10G+ migration network, enable multiqueue virtio when appropriate, and consider post-copy migration as a last resort since it adds a dependence on the destination host. Inside the guest, bonded interfaces keep the network path redundant across host NICs, which is what stops a single uplink failure from becoming a migration failure.

Troubleshooting

virsh dominfo web01
virsh dumpxml web01 | grep -A3 "<devices>"
journalctl -u libvirtd --since "15 min ago"
tail -50 /var/log/libvirt/qemu/web01.log
virsh blockjob web01 vda
virsh domstats web01 | grep -E "cpu|balloon"
virsh qemu-monitor-command web01 --hmp "info block"
  • Guest will not start, "permission denied" on the disk: SELinux or AppArmor label on the image. Relabel and retry rather than disabling the policy.
  • No network after boot: bridge STP forwarding delay, or the guest NIC is attached to the wrong bridge. Check bridge fdb show.
  • Very slow disk inside the guest: cache=none with io=native on a host with a battery-backed controller, or the qcow2 chain has grown too deep after many snapshots — blockcommit and flatten.
  • Migration never converges: reduce dirty page rate (throttle the workload), use a faster link, or allow post-copy.

Operations checklist

  • Pin guest XML in Git and name guests consistently with their DNS record.
  • Alert on host memory pressure and on any guest that has survived a reboot inconsistency.
  • Test a restore of a guest from backup, not just the backup itself.
  • Keep one spare host's worth of capacity so a failed node can be drained — the same N+1 thinking that applies to network devices.
  • Document console access for every guest; a VM you cannot reach without a working network is not manageable.

原文链接:https://libvirt.org/formatdomain.html