NVMe over TCP: Target and Initiator Configuration - 夜莺博客

NVMe over TCP: Target and Initiator Configuration

NVMe over Fabrics used to mean RDMA, which meant a lossless Ethernet configuration and expensive NICs before you could store a single byte. NVMe/TCP removes that barrier: standard 25/100 GbE adapters, standard TCP, and a queue model that still gives you most of the latency benefit of NVMe. What it does not remove is the need to size TCP buffers, MTU and multipath correctly. This article walks a working target/initiator pair on Linux and the checks that separate "connected" from "performing".

The Model in Four Objects

  • Subsystem — the exported block device set, identified by an NQN such as nqn.2026-10.io.example:subsystem-storage-01.
  • Namespace — the block device inside the subsystem (a partition, LV, ZFS zvol or file-backed device).
  • Port — a listening TCP endpoint (address + port, default 4420).
  • Allowed host — the initiator NQN permitted to connect, optionally with DH-HMAC-CHAP authentication.

Target Side (the storage server)

modprobe nvmet
modprobe nvmet-tcp
mkdir -p /sys/kernel/config
mount -t configfs none /sys/kernel/config 2>/dev/null || true

# 1. Create the subsystem
mkdir /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01
echo 1 > /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01/attr_allow_any_host

# 2. Add a namespace backed by an LV (or a file for testing)
mkdir /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01/namespaces/1
echo -n /dev/vg_data/nvme_vol >   /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01/namespaces/1/device_path
echo 1 > /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01/namespaces/1/enable

# 3. Create a port and link it
mkdir /sys/kernel/config/nvmet/ports/1
echo 10.10.20.10 > /sys/kernel/config/nvmet/ports/1/addr_traddr
echo tcp > /sys/kernel/config/nvmet/ports/1/addr_trtype
echo 4420 > /sys/kernel/config/nvmet/ports/1/addr_trsvcid
echo ipv4 > /sys/kernel/config/nvmet/ports/1/addr_adrfam
ln -s /sys/kernel/config/nvmet/subsystems/nqn.2026-10.io.example:storage-01       /sys/kernel/config/nvmet/ports/1/subsystems/nqn.2026-10.io.example:storage-01

# 4. Verify on the target
dmesg | tail -20
ss -tlnp | grep 4420
nvmetcli save /etc/nvmet/config.json      # persist across reboots

Note the difference between attr_allow_any_host and explicit host allow-listing. For production, create the allowed-host entry with the initiator's NQN instead, and consider CHAP authentication — an NVMe/TCP port is an unauthenticated block device export otherwise.

Initiator Side (the server consuming storage)

modprobe nvme-fabrics
apt install nvme-cli

# Discovery tells you which subsystems the target exports
nvme discover -t tcp -a 10.10.20.10 -s 4420

# Connect
nvme connect -t tcp -a 10.10.20.10 -s 4420 -n nqn.2026-10.io.example:storage-01

# Inspect, then set up multipath
nvme list
nvme list-subsys
lsblk -o NAME,SIZE,TYPE,MODEL,TRAN
ls -l /dev/disk/by-id/ | grep nvme

Persisting the connection

# /etc/nvme/discovery.conf
--transport=tcp --traddr=10.10.20.10 --trsvcid=4420

# systemd unit that connects at boot
cat >/etc/systemd/system/nvme-connect.service <<'EOF'
[Unit]
Description=NVMe/TCP connect
After=network-online.target
Wants=network-online.target

[Service]
Type=oneshot
RemainAfterExit=yes
ExecStart=/usr/sbin/nvme connect-all -t tcp -a 10.10.20.10 -s 4420
ExecStop=/usr/sbin/nvme disconnect-all
EOF
systemctl enable --now nvme-connect.service

Multipath: The Part That Provides Redundancy

# Two target ports on different subnets, from the initiator:
nvme connect -t tcp -a 10.10.20.10 -s 4420 -n nqn...:storage-01
nvme connect -t tcp -a 10.10.21.10 -s 4420 -n nqn...:storage-01

nvme list-subsys       # both controllers should appear under one subsystem
multipath -ll          # native NVMe multipath (or dm-multipath if configured that way)

Modern kernels use native NVMe multipath (/sys/module/nvme_core/parameters/multipath=Y), which is simpler than device-mapper and does not require /etc/multipath.conf entries. Confirm which mode you are in before writing any udev or multipath rules — configuring dm-multipath on top of native multipath is a documented way to get two device nodes for one LUN.

Performance Tuning and Bottlenecks

# TCP buffers must match the bandwidth-delay product
sysctl -w net.ipv4.tcp_rmem="4096 262144 33554432"
sysctl -w net.ipv4.tcp_wmem="4096 262144 33554432"
sysctl -w net.core.rmem_max=33554432
sysctl -w net.core.wmem_max=33554432

# Jumbo frames where the whole path supports them
ip link set dev ens1f0 mtu 9000

# Queue and interrupt distribution
ethtool -l ens1f0
ethtool -S ens1f0 | egrep 'drop|error'
  • Single-flow throughput capped around 10 Gbit/s on a 100 GbE link. TCP socket buffers, not the network. Raise rmem_max/wmem_max and retest.
  • Latency far above RDMA expectations. Expected — NVMe/TCP trades roughly 10–30 µs of additional latency for not needing lossless Ethernet. If your application cannot tolerate that, it needs RDMA, not TCP tuning.
  • Commands queue but never complete after a network blip. Set a controller loss timeout: nvme connect ... --ctrl-loss-tmo 600, and make sure the timeout is shorter than the application's own retry window.
  • Authentication failures after enabling CHAP. Host NQN must match exactly, including the trailing controller index if the initiator uses one.

Verification Checklist

  1. nvme discover from every initiator subnet you plan to use.
  2. nvme list-subsys shows all intended paths to one subsystem.
  3. Failover tested by taking a target port down — I/O continues, and dmesg shows path failure and recovery.
  4. Throughput measured with fio against the expected link rate, not assumed.
  5. Target configuration persisted (nvmetcli save) and initiators reconnect after a reboot.

Related reading: FC vs iSCSI vs NVMe-oF, iSCSI multipath tuning and Linux multipath configuration.

原文链接:https://docs.ceph.com/en/reef/rbd/nvmeof-initiator-linux