Mellanox MLNX-OS: Back Up Switch Configuration to Server - 夜莺博客

Mellanox MLNX-OS: Back Up Switch Configuration to Server

Regular configuration backups are the cheapest insurance policy for any network device, and Mellanox MLNX-OS switches are no exception. This guide, based on IBM's documentation for the Mellanox MSX1710 (8831-NF2) switch, shows every way to back up and restore an MLNX-OS configuration: downloading the running config from the web GUI, generating a BIN configuration file, creating a text-based configuration file, and uploading either format to an external server with SCP. A backup strategy like this turns a failed switch into a 15-minute replacement instead of a rebuild from memory.

The reason this matters more on Mellanox hardware than on a typical access switch is that MLNX-OS configurations are rarely short. A single uplink port can carry PFC, ECN, QoS maps, LAG membership, VRF assignment and a dozen VLAN memberships — none of which are visible from a photograph of the rack. And because these switches frequently sit in front of storage arrays, an unplanned rebuild has a data-path cost, not just an ops cost.

Three Backup Formats and When to Use Each

  • BIN (binary) configuration file — the complete, machine-loadable configuration. Restores the switch exactly, including settings that have no CLI representation. Use it for disaster recovery.
  • Text configuration — the same configuration rendered as CLI commands. Human-readable, diff-friendly and ideal for change management, peer review and version control.
  • GUI download — the lowest-friction path for a one-off copy, and the one most often used when the switch was configured entirely through the web interface.

A mature routine keeps both a BIN file (the recovery artefact) and a text file (the reviewable artefact) for every switch, dated the same way and stored together.

Backing Up via the Web GUI

Log in to the switch by entering its IP address in a browser (for example, https://192.168.93.35) with the default admin/admin credentials. Click Setup, then Configuration, check the box for the file loading as the running config, click Download and save the BIN file to your desired location.

Two cautions for production switches: change the default admin password before the switch is reachable from anything but a management VLAN, and confirm that "running config" is the file you selected rather than a stale saved file. Downloading the wrong configuration file is worse than having no backup, because it silently validates a wrong state.

Preparing the Switch Side of the Transfer

Before any SCP upload, make sure the switch can resolve and reach your backup server. On a management-VLAN-only switch that means three things: an address, a route and a working name resolution path.

switch (config) # interface mgmt0 ip address 192.168.93.35 /24
switch (config) # ip route default gateway 192.168.93.1
switch (config) # ip name-server 192.168.93.10
switch (config) # show ip route

If your environment forbids DNS for infrastructure devices, use the IP address directly in the SCP URL rather than a hostname — it also removes one failure mode from an already-fragile recovery moment.

Creating a BIN Configuration File from the CLI

The configuration snapshot family of commands writes the active running configuration into a named binary snapshot on the switch's own flash:

switch (config) # configuration snapshot active running save my-backup

The BIN format captures the complete binary configuration. To upload it to an external file server:

switch (config) # configuration snapshot active running file my-backup upload scp://root@my-server/root/tmp/my-backup

The URL contains a username and host, so the switch prompts for the remote password interactively. That is acceptable for a manual backup but it is exactly what blocks unattended automation — see the scheduling section below for the workaround.

Creating and Uploading a Text-Based Configuration

The text variant generates a readable copy of the same configuration and then ships it off-box with the same SCP syntax:

switch (config) # configuration text generate active running save my-filename
switch (config) # configuration text file my-filename upload scp://root@my-server/root/tmp/my-filename

Text-based files are human-readable and diff-friendly, which makes them ideal for change management and version control. The same mechanisms can be used in reverse to restore or merge configurations after a failure.

If you run this through a terminal server or an interactive SSH session, remember that the generate ... save step has already written the file to flash; the file ... upload step is a pure transfer. You can therefore generate once and upload the same artefact to several destinations — for example a jump host, a NAS export and a version-control working copy.

Verifying the Files Before You Trust Them

A backup that has never been inspected is a belief, not a backup. Check three things immediately after every upload:

  • Size is non-zero and plausible. An empty or few-hundred-byte file usually means the SCP transfer authenticated and then failed silently.
  • Timestamps match the switch's clock. If the file is older than the change you just made, you backed up a stale snapshot.
  • The text file contains your most recent change. Grep for a VLAN or a hostname you added in the last change window — this single check catches the most common backup error of all.
# on the backup server
ls -l /root/tmp/my-backup /root/tmp/my-filename
grep -n "interface ethernet 1/49" /root/tmp/my-filename
sha256sum /root/tmp/my-filename

Restoring from a Backup

The reverse path uses the same primitives. Fetch the file from the server, then activate it as the configuration the switch loads:

switch (config) # configuration fetch scp://root@my-server/root/tmp/my-filename
switch (config) # show configuration files
switch (config) # configuration switch-to my-filename

On a live switch, prefer restoring during a maintenance window and be explicit about whether you are replacing or merging configuration. A replace-style restore on a switch that carries storage traffic will drop links and, in the worst case, erase the management address you are connected through — have a console cable and out-of-band access ready before you press Enter. The file-handling mechanics, the no-switch flag and factory reset options are covered in detail in the MLNX-OS configuration management guide.

Automating Backups on a Schedule

Manual backups happen exactly as often as somebody remembers them. Two practical automation patterns work well on MLNX-OS:

Pattern 1 — drive it from a Linux jump host. Use sshpass or key-based access with an expect wrapper to run the two configuration ... upload commands nightly, then keep ten generations with a dated rotation. Because the switch prompts for the SCP password interactively, expect-style automation is usually simpler than trying to embed credentials.

# jump host, nightly cron
0 2 * * * /usr/local/bin/mlnx-backup.sh sw-core-01 192.168.93.35
# inside the script, after login:
#   configuration text generate active running save auto-$(date +%F)
#   configuration text file auto-$(date +%F) upload scp://root@nas/vol/backups/

Pattern 2 — pull instead of push. If your security policy prefers inbound-only access to the backup server, keep a nightly configuration text generate active running save on the switch and pull the generated file from the switch with SCP from the server side. This removes the need for the switch to hold any credentials at all.

Either way, alert on failure. A backup job that fails quietly for three weeks is indistinguishable from no backup job, right up until the moment it matters.

What Else Deserves a Backup

The configuration file is the headline item, but a switch rebuild also needs material the configuration does not contain:

  • Firmware version and image name — so the spare boots the same code. Record the output of the version query next to each backup.
  • Licence information — feature licences and their keys, for platforms that require them.
  • Asset facts — serial number, model, slot layout for modular chassis, and the management address of every interface.
  • Physical documentation — a port map listing what is plugged into every port, plus a photograph of the rack cabling taken at commissioning.

Bundle these into one dated directory per switch. During a real recovery you will read them under pressure, and a directory that already contains everything is worth far more than a perfect structure you have to assemble at 2 a.m.

SCP Versus the Other Transports

MLNX-OS understands more than SCP in upload and fetch URLs, but SCP is the one that works everywhere. FTP has no encryption at all, TFTP has neither encryption nor authentication and is frequently blocked by security policy, and HTTP uploads depend on a server willing to accept a PUT. For any switch carrying production traffic, SCP over SSH is the only sensible default — and on newer builds SFTP is accepted in the same URL syntax, so either form will do.

Whichever transport you pick, settle the host key situation once at commissioning. A switch that has never contacted your backup server will stop and ask you to accept a key, and an unattended script fails on that prompt rather than on anything to do with the configuration. Run one interactive transfer first.

Restore Drills: Proving the Backup Works

Once a year, on a lab switch or a spare, restore a production configuration and confirm three things: the switch boots, the management address matches what your documentation claims, and the data-plane ports come back in the modes and VLANs you expect. That is the entire drill — twenty minutes, and it is the only way to know that the BIN file you have been uploading nightly is loadable and complete.

Keep a written one-page runbook with the result: date, operator, firmware version restored, and any step that differed from the script. Recovery procedures rot silently, and the runbook is what turns a restore from an improvisation into a procedure.

Backup Best Practices

Automate the SCP uploads on a schedule, keep several generations of backups, store them off-box (ideally in version control), and test a restore in the lab at least once per year. Document the switch's current MLNX-OS version alongside the backup so the recovery path is unambiguous.

A few additions that pay off in real incidents: record the switch model and the license state next to each backup, because a BIN file from one platform cannot always be loaded onto a spare of a different model. Keep an out-of-band path to the switch — a console server entry or a serial cable in the rack — since a restore may cost you the management IP. And version your text configurations with the same discipline you apply to application code: a one-line commit message stating what changed turns a restore decision into a two-minute review.

Finally, tie the backup cadence to your change process. If every change ticket ends with "config backed up and committed", the backup set cannot drift more than one change behind reality — and that is the only guarantee that actually holds.

More Mellanox content: getting started with Mellanox switches, MLNX-OS breakout cable link troubleshooting, and ONIE and Onyx MLNX-OS installation. For CLI fundamentals, see the MLNX-OS CLI modes and commands guide, and for connectivity after a restore, VLAN interfaces and IP routing.

原文链接:https://www.ibm.com/docs/en/power8/0000-REF?topic=POWER8_REF/p8ef9/p8ef9_backup_switch_mellanox.html