NVMe over TCP: Target and Initiator Setup Guide - 夜莺博客

NVMe over TCP: Target and Initiator Setup Guide

NVMe over Fabrics does not require RDMA. The TCP transport ships in the mainline kernel (5.0 and later), works on the Ethernet you already own, and turns a remote NVMe namespace into a block device that behaves like a local disk. For labs, database replicas and cheap storage servers, NVMe/TCP is often a better fit than another iSCSI target. This guide sets up both sides: the kernel target via nvmetcli, and the initiator via nvme-cli, including the security and MTU details that decide whether it is production-ready.

Why NVMe over TCP Instead of iSCSI or RDMA

NVMe over Fabrics defines a transport-agnostic command set: the same submission and completion queues a local SSD uses are carried over a fabric. Three transports implement it — Fibre Channel (FC-NVMe), RDMA (RoCEv2 or InfiniBand) and TCP. FC-NVMe needs a Fibre Channel SAN at both ends. RoCEv2 needs lossless Ethernet with PFC and ECN configured identically on every switch in the path, which is exactly where most projects stall. TCP needs none of that: it runs over routed, congested, best-effort Ethernet, crosses subnets, and is carried by the same firewalls, load balancers and monitoring tools you already operate.

The trade is latency. RoCEv2 adds tens of microseconds; NVMe/TCP adds a few hundred microseconds of per-I/O overhead because of TCP stack processing, checksumming and ACK handling. For a database doing 4 KiB random reads at queue depth one, that difference is measurable. For backups, streaming, VM images, replica targets, container persistent volumes and most lab work, it is irrelevant — the application never sees it above the noise of the storage media itself.

Transport Hardware required Where it fits
iSCSI Any Ethernet NIC; carries SCSI commands Legacy compatibility, widest OS support, existing LUN tooling
NVMe/TCP Any Ethernet NIC; kernel 5.0+ New deployments that want NVMe semantics without rebuilding the fabric
RoCEv2 RNICs, lossless Ethernet, PFC/ECN tuning Latency-sensitive tier-1 workloads with a dedicated network team
FC-NVMe Fibre Channel HBA and switch fabric Existing SAN estates needing maximum determinism

The practical rule: if you are building storage now and do not already own Fibre Channel or a lossless Ethernet fabric, start with NVMe/TCP. You keep the option of moving to RoCEv2 later, because the namespace, subsystem NQN and host ACL model are identical across transports — only the trtype changes.

Architecture and Prerequisites

Component Role
Target (storage server) nvmet kernel target exports a block device (whole NVMe disk, LVM LV or partition) as a namespace inside a subsystem
Transport nvmet-tcp / nvme-tcp, listening on TCP 4420 by default
Initiator (client) nvme-cli discovers and connects; the namespace appears as /dev/nvmeXn1

Use a modern kernel (5.15+ is a sensible floor for production TCP support), and keep control traffic symmetrical: firewalls must allow TCP 4420 in both directions for discovery and I/O.

Target Side: Export a Block Device

sudo apt update && sudo apt install -y nvmetcli
sudo modprobe nvmet
sudo modprobe nvmet-tcp

sudo nvmetcli
# inside the interactive shell:
/ > cd /subsystems
/ > create nqn.2026-09.com.example:storage01
/ > cd nqn.2026-09.com.example:storage01
/ > set attr allow_any_host=1            # lab only — see security below
/ > cd namespaces
/ > create 1
/ > cd 1
/ > set device path=/dev/nvme1n1         # or /dev/vg0/lv_data
/ > set enable 1
/ > cd /
/ > cd /ports
/ > create 1
/ > cd 1
/ > set addr trtype=tcp adrfam=ipv4 traddr=10.0.0.11 trsvcid=4420

Then link the subsystem to the port, either in the same nvmetcli session (ls to the ports/1/subsystems directory and ln -s the subsystem) or with the sysfs equivalent. Save the running configuration so it survives a reboot:

# in nvmetcli: save
/ > saveconfig /etc/nvmet/config.json
# and restore after boot:
sudo nvmetcli restore /etc/nvmet/config.json

Host ACLs: Locking Down the Subsystem

allow_any_host=1 means any initiator that can reach port 4420 may attach the namespace — that is unauthenticated block-level access to your data. Before production, create one host object per initiator, keyed by that initiator's NQN, and leave allow_any_host=0. The client's NQN is printed by cat /etc/nvme/hostnqn.

# target side
sudo nvmetcli
/ > cd /hosts
/ > create nqn.2014-08.org.nvmexpress:uuid:9f2c1b74-...   # copy from the client's /etc/nvme/hostnqn
/ > cd /
/ > cd /subsystems/nqn.2026-09.com.example:storage01
/ > set attr allow_any_host=0
/ > cd allowed_hosts
/ > ln -s ../../hosts/nqn.2014-08.org.nvmexpress:uuid:9f2c1b74-...
/ > saveconfig /etc/nvmet/config.json

If you cannot copy the client NQN by hand, connect once with allow_any_host=1, then read the presenting NQN from dmesg on the target before you close that window. On the initiator the NQN can be regenerated at will (sudo nvme gen-hostnqn > /etc/nvme/hostnqn) — do that before you register it with the target, not after, because a changed NQN stops the connection from coming back on reboot.

Host ACLs are authorization, not authentication. Anyone able to spoof an NQN can present it, so the real control is network reachability: keep TCP 4420 on a dedicated storage VLAN, filter it at the switch or host firewall, and never expose the target port to a user or office segment.

Choosing the Backing Device

The namespace can back any block device the kernel sees, and the choice affects performance more than any tuning knob in this article.

  • Whole NVMe disk (/dev/nvme1n1): lowest latency, no extra layer. Best when the disk is dedicated to the export.
  • LVM logical volume (/dev/vg0/lv_data): thin provisioning, snapshots, easy resizing. Costs one device-mapper layer of latency and hides the disk's own queue limits from the initiator.
  • Partition (/dev/nvme1n1p1): simple, but easy to confuse with the host's own filesystems.
  • ZFS zvol: good snapshots and compression, but volblocksize must align with the workload's I/O size (16 KiB is a common database choice) or read-modify-write amplification destroys throughput.

Whatever you pick, the namespace must not be mounted locally while it is exported. Exporting a mounted, actively-written filesystem over a block fabric is a reliable way to corrupt it.

Initiator Side: Discover, Connect, Persist

sudo apt install -y nvme-cli
sudo modprobe nvme-tcp

# who is advertising?
sudo nvme discover -t tcp -a 10.0.0.11 -s 4420

# connect to the subsystem by NQN
sudo nvme connect -t tcp -a 10.0.0.11 -s 4420 -n nqn.2026-09.com.example:storage01

sudo nvme list              # the remote namespace appears, e.g. /dev/nvme1n1
sudo nvme list-subsys       # transport, state and address of each connection

Make connections persistent across reboots with a discovery entry and the autoconnect service:

# /etc/nvme/discovery.conf
-t tcp -a 10.0.0.11 -s 4420

sudo systemctl enable --now nvmf-autoconnect

Production Hardening

  • Do not leave allow_any_host=1. On any untrusted segment, create host objects for each initiator NQN and allow only those. An open subsystem is block-level storage accessible to anyone who can reach port 4420.
  • Jumbo frames pay off. NVMe blocks are typically 4 KiB, which a 1500-byte MTU splits into three frames. Setting MTU 9000 on switch ports, target NICs and initiator NICs lets one storage block travel in one frame. Mismatched MTU between target and initiator is a classic cause of "connected but writes hang".
  • Use dedicated interfaces/NICs for storage traffic and, where available, pin interrupts and queues for consistent latency.
  • Decide on multipath up front. Two target addresses plus nvme native multipath (or dm-multipath for older kernels) gives you a path failover story before the first cable is pulled.

Multipath and Failover in Practice

A single TCP connection is a single point of failure. NVMe native multipath, built into the kernel, presents multiple paths to the same namespace as one block device and fails over without an administrator. Set it up before the first production attach, not after the first outage.

# /etc/nvme/discovery.conf — one line per target address
-t tcp -a 10.0.0.11 -s 4420
-t tcp -a 10.0.0.12 -s 4420

sudo nvme connect-all -t tcp -a 10.0.0.11 -s 4420
sudo nvme list-subsys          # both paths listed under one subsystem
sudo nvme list-subsys -o json  # machine-readable path state
cat /sys/module/nvme_core/parameters/multipath   # must be 'Y'

To prove multipath works, block one target address or pull one cable: the namespace must stay readable while the failed path shows state=resetting in nvme list-subsys, then recovers. If I/O stalls instead, the kernel is waiting out a timeout — check nvme_core.io_timeout and whether the controller was connected with natively-managed multipath or silently fell back to a single path. On older kernels or with shared enterprise arrays, dm-multipath remains the supported option and needs a matching multipath.conf stanza.

Where NVMe/TCP Fits in the Real World

Three deployments account for most NVMe/TCP production use, and they share a common trait: the workload is throughput-hungry and tolerant of a few hundred microseconds of latency.

  • Storage server for virtualisation. A host with six or eight NVMe drives exports every namespace directly; hypervisors attach them and skip the iSCSI session layer entirely. LUNs appear as local NVMe devices and the kernel's own multipath handles failover.
  • Database replicas and backup targets. Replication streams and backup windows are bound by the target's write throughput, not by network latency. Removing the SCSI translation layer typically buys 20 to 40 percent more IOPS at the same CPU cost.
  • Replacing an ageing SAN. Where a Fibre Channel fabric is reaching end of support, NVMe/TCP over the existing 10G or 25G Ethernet gives comparable throughput without new HBAs, switches, zoning or a separate operations skill set.

Sizing Expectations Before You Deploy

Set the performance target from the component that will actually limit it. On a healthy 25G link with a modern CPU, NVMe/TCP typically delivers 1.5 to 2.5 GB/s per connection and 200 to 400 thousand 4 KiB random-read IOPS per target CPU core. Two failure modes explain almost every disappointing result: the target process is pinned to one core and saturates it, or a single TCP connection's window limits throughput before the fabric does. Adding a second connection, a second target port, or spreading interrupts across cores fixes both.

Measure with fio on the initiator and read the target's top and mpstat output at the same time. If the target is not busy while throughput is low, the bottleneck is the network or the initiator; if it is pinned at 100 percent of one core, the fix is in the target configuration, not the fabric.

Verify and Benchmark

cat /sys/class/nvme/nvme1/transport       # should report tcp
dmesg | grep -i nvmet_tcp                 # target side: port configuration messages
sudo nvme id-ctrl /dev/nvme1 | head -20

# quick performance check, aligned to 4k
sudo fio --name=nvmetcp --filename=/dev/nvme1n1 --direct=1 \
         --rw=randread --bs=4k --iodepth=32 --numjobs=4 --runtime=60 --time_based

If throughput is far below expectations, check in this order: MTU consistency end to end, CPU single-core saturation on the target, and whether the initiator is using multiple queues (nvme list-subsys -o json shows the path count).

Troubleshooting Matrix

Symptom Likely cause What to check
nvme discover returns nothing Port not created, or subsystem not linked to the port nvmetcli listing on the target; dmesg | grep -i nvmet_tcp
Discovery works, connect says "no such subsystem" Subsystem exists but is unattached The ln -s step under /ports/1/subsystems
Connect says "host not allowed" Host ACL missing while allow_any_host=0 dmesg on the target prints the presenting NQN
Connects but writes hang MTU mismatch along the path, or a firewall dropping large frames Compare MTU on both NICs and every switch port in between
Throughput far below expectation Single queue, CPU-bound target, shared congested link nvme list-subsys -o json, IRQ affinity, ethtool -S, target top
Namespace missing after reboot Discovery entry not persisted nvmf-autoconnect enabled and /etc/nvme/discovery.conf present
Cannot recreate a deleted namespace Subsystem still referenced by a host or port link Remove allowed_hosts and port links before deleting the subsystem

Keep the target configuration in version control as /etc/nvmet/config.json. Because it is plain JSON, a diff between the saved file and the live listing tells you immediately whether a reboot will bring the namespaces back.

相关阅读:Fibre Channel、iSCSI 与 NVMe-oF 对比、Linux multipath.conf 配置与 iSCSI SAN 以及 LIO 与 targetcli 构建 iSCSI 目标。

原文链接:NVMe over TCP: exposing NVMe targets over a standard network