Storwize v3700集群创建报错解决

Storwize v3700: CMMVC8020Ev3700尝试在机柜已存储集群

IBM Storwize v3700 — today sold as the IBM FlashSystem 5000 family — refuses to build a new system when a cluster ID from a previous life is still present in the enclosure hardware. The service assistant (initialization) interface returns CMMVC8020E: Attempted to create a cluster while a storage cluster ID exists and no new system can be created until that stale identity is removed. This article explains what the cluster ID is, why it survives a system delete, and the three ways to clear it: the enclosure configuration page in the GUI, the satask chenclosurevpd -resetclusterid command over SSH, and a clean re-initialization when the identifier keeps coming back.

What CMMVC8020E Actually Means

Storwize code (and its later FlashSystem derivatives) stores a cluster identifier in the enclosure vital product data (VPD) of each node canister. That identifier is written the first time a system is created, and it is how the hardware knows to which system, and to which cluster, it belongs. During initialization the cluster-creation routine reads the enclosure VPD and checks for an existing ID: if one is found, the code assumes the canisters are already part of a system and stops, which is exactly the condition described by CMMVC8020E — “attempted to create a cluster while a storage cluster ID exists”.

The error is therefore not a hardware fault and not a firmware bug. It is a state problem: the canisters still carry the fingerprint of a system that was deleted, never fully created, or de-initialized by a failed upgrade or a botched earlier initialization. Re-running the wizard a second time without clearing the ID simply produces the same error, because nothing in the wizard removes it.

Why the Cluster ID Survives a Delete

Several common operational sequences leave the cluster ID behind:

  • A system was deleted from the management or service GUI, then the enclosures were kept as spares. The delete removes the cluster definition from the running system but the enclosure VPD is not necessarily rewritten to a blank value.
  • An earlier initialization was interrupted — power loss, network drop, or a canister that was reseated mid-wizard — so the ID was written but the cluster never completed creation.
  • The node canisters came from a different system (purchased second-hand, or recycled after a decommission) and were swapped into this enclosure.
  • Only one of the two canisters was reset, leaving the pair in a mismatched state where one holds a cluster ID and the other does not.

That last point explains the most confusing symptom: the wizard fails, someone resets the enclosure, the wizard works, and a week later lsnodecanister shows a node that never joined. Clearing the identifier on one canister does not necessarily clear the other, and the two must agree before the cluster can be created cleanly.

Before You Touch Anything: Prerequisites and Risk

A cluster-ID reset is a destructive initialization prerequisite, not a repair. Treat it as such:

  • Confirm there is no data to keep. If the arrays contain LUNs with production data, a cluster-ID reset and re-initialization can leave the system unusable without vendor-assisted data recovery. Stop here and open a case with IBM/Lenovo support.
  • Record the current VPD and inventory so that you can prove what the hardware looked like before the change. The service assistant exposes read-only commands for this, and the output is what support will ask for.
  • Check both node canisters and every expansion enclosure, and confirm that the firmware levels of the two nodes match; a mismatched pair is a separate problem that no cluster-ID reset fixes.
  • Have the service network ready. The default service IP addresses of the node canisters are 192.168.70.121 and 192.168.70.122, reachable on the technician/service port or on eth0 depending on the model. Both must answer before you start.
  • Note the SAN identity (node WWPNs and any host zoning) if the array will keep its roles, because a re-initialized system gets a new cluster identity and your zoning and multipath configuration must be revisited. See our Linux multipath configuration for iSCSI SAN guide for the host side.

Method 1: Reset the Cluster ID from the Enclosure Configuration Page

This is the documented, GUI-first path and the one to try first. It is the method that resolved the original case this article is based on.

  1. Connect a laptop to the service port and browse to the service assistant on https://192.168.70.121 (or .122 for the second canister). Accept the self-signed certificate.
  2. Log in as service. The welcome page shows the state of both node canisters; if a canister is not reachable, fix the network or reseat it before continuing.
  3. Open the Enclosure configuration page (the enclosure/机柜 configuration view used during initialization).
  4. Look for the cluster ID reset control on that page and apply it. On some code levels the same function is offered as part of the “reset” or “de-initialize” action for the enclosure.
  5. Start the Create a new system (initialization) wizard again. In the reported case the wizard completed immediately after the identity was reset.
  6. If the error reappears, check the second node canister and repeat on it, then continue with Method 2.

The GUI path is preferred because it performs the reset through the same service layer that the wizard uses, which keeps the two canisters consistent. It is also fully logged, which matters if you later need vendor support.

Method 2: satask chenclosurevpd -resetclusterid over SSH

When the GUI control is missing, greyed out, or the error persists, the underlying service task can be run directly on the canister. Log in to the service IP over SSH and run the task:

ssh service@192.168.70.121
satask chenclosurevpd -resetclusterid

satask chenclosurevpd changes enclosure VPD fields; with -resetclusterid it clears the stored cluster identifier. The command must be run on each node canister that holds a stale ID, and the canister should be restarted (or the service state refreshed) before you retry the initialization wizard. If your code level does not recognise the option, print the task help first and use the equivalent reset action listed there:

satask chenclosurevpd -help
satask lshw
satask lsservicestatus

Two warnings apply. First, this command changes hardware VPD and the vendor documents it as a service action — it is not something to run speculatively on a system with data. Second, if the reset appears to succeed but the wizard still reports CMMVC8020E, the ID is probably present on the other canister or the canisters are running mismatched firmware, and the correct next step is a support case rather than repeated attempts.

Method 3: Clean Re-Initialization When the ID Keeps Returning

If both methods above leave the system in the same state, the practical path is a full, deliberate re-initialization: power both canisters down, verify that the service network is clean (no DHCP conflicts, no duplicate addresses from an old system), power up, wait for both canisters to reach “candidate” or “service” state in the service assistant, reset the cluster ID on both, and then run the Create a new system wizard once and let it finish without interruption. Do not interleave other service actions with the wizard — the cluster-creation step is where the ID is written, and a second action in flight at that moment is a common cause of a half-created system.

When the wizard completes, the new system has a fresh cluster identity. Initialize with a system name, confirm the management IP, and then verify from the CLI before handing the array to the SAN team.

lssystem
lsnodecanister
lsenclosure
lsmdiskgrp

After the Reset: Building the New System

Once the cluster exists, the remaining work is ordinary storage initialization: assign the management IP addresses, confirm both node canisters appear as online in lsnodecanister output, create the storage pool (MDisk group) over the internal drives, attach external MDisks if the v3700 is virtualizing another array, then create volumes and host mappings. Because the cluster identity and node WWPNs change, re-zone the SAN fabric for the new WWPNs and update the host multipath configuration — otherwise hosts will see the new array as a completely different device, which is often mistaken for a LUN-presentation fault.

Avoiding This Situation on the Next Decommission

  • When you remove a system, complete the documented decommission routine rather than powering hardware off: delete the cluster definition, then let the canisters return to uninitialized state.
  • Label enclosures and canisters with the system name and serial number they came from, and store that label with the spare.
  • Before reusing a canister from another system, verify its firmware level and reset the cluster ID on the bench, not in the customer rack.
  • Keep the service-network addresses documented; two canisters with a stale address from a previous system can produce error messages that look like hardware failures.
  • For vendor-supported arrays, keep a support case open for any cluster-ID manipulation on hardware that has ever held data.

Related Storage Guides

If you work with the wider family of IBM and HPE arrays, these articles cover adjacent tasks: IBM DS5020 management port password reset, regaining management access to a Dell PowerVault ME4/ME5, HPE 3PAR administration commands, and adjusting cluster settings on 3PAR controllers.

原文记录(Original Chinese Notes)

问题背景

IBM Storwize v3700 是中端存储阵列(现为 IBM FlashSystem 5000 系列),支持将多台存储设备组成存储集群(Storage Cluster)统一管理。当设备之前加入过集群或集群信息未完全清除时,在服务界面新建系统(初始化集群)会报错 CMMVC8020E: 存在存储集群标识时尝试创建集群,导致无法完成初始化。本文记录了该问题的解决方法。

解决方法

首先在机柜配置界面(Enclosure 配置)重置集群标识,再次尝试新建系统即可成功。如果界面操作后仍然不行,可以尝试通过 SSH 登录存储执行以下命令强制重置集群标识:

v3700从服务界面新建系统时,提示:

CMMVC8020E: 存在存储集群标识时尝试创建集群

无法创建新系统,是因为以前的集群信息尚未完全清除干净,在机柜配置界面重置集群标识,再次尝试新建系统成功。

如果还是不行,可以尝试[1]里面讲的方法,ssh登录进存储,执行以下命令:

satask chenclosurevpd -resetclusterid

没有尝试,或许可以。

References:
[1][SOLVED] Storwize v3700: CMMVC8020E

命令说明

satask chenclosurevpd -resetclusterid 用于重置机柜(Enclosure)的集群标识 VPD 信息,执行后需重启或重新初始化设备。该命令风险较高,执行前请确认设备上无重要业务数据,并建议联系 IBM/联想支持确认后再操作。