HPE MSA 2040 iSCSI Configuration - 夜莺博客

HPE MSA 2040 iSCSI Configuration

原文:HPE MSA 2040 iSCSI Configuration — theDXT (Daniel Keer)

I recently deployed an HPE (Hewlett Packard Enterprise) MSA (Modular Smart Array) 2040 SAN (Storage Area Network) Storage unit in my home lab.

In this post, I will show you step-by-step how to set up the HPE MSA 2040 SAN Storage with iSCSI. The walkthrough below covers the whole chain: putting the management ports and the iSCSI data ports on the network, building disk groups and volumes, registering hosts and initiators, mapping LUNs, and then wiring the VMware ESXi side so the datastore actually appears. I have also added the checks that catch the usual mistakes — wrong controller ownership, a LUN 0 collision with vSphere, jumbo frames that only match on one end and a second path that was never configured.

What the MSA 2040 actually is

The MSA 2040 is a dual-controller SAN array built around a 2U or 4U drive enclosure. Each controller has its own management port, four host ports and a dedicated SAS expansion port for adding drive enclosures. Depending on the model you buy, the host ports are 1GbE or 10GbE iSCSI, 8Gb or 16Gb Fibre Channel, or a mix of protocols, which is why the same SMU screenshots appear with completely different port labels on different forums. A controller failure is not a rebuild event: the surviving controller takes over the pools the failed one owned, so long as the host is multipathed.

That last sentence is the single most important design rule for this array. Everything below — port A/B split, host groups, two volumes instead of one, Round Robin in ESXi — exists to make sure a controller restart costs you a few seconds of path failover and not an outage.

Before you start: the checklist

Do these five things before you open a browser, because half of the "my MSA is broken" threads online are really one of these five things being wrong:

  • Firmware. Check that both controllers run the same firmware version and that it is a version your HPE support contract covers. Mixed firmware between controller A and controller B causes odd, intermittent path loss.
  • Management network. Out of the box the two controllers answer on 10.0.0.2 (controller A) and 10.0.0.3 (controller B) with a 255.255.255.0 mask. Know which port is which before you plug in.
  • Data network. Decide the iSCSI subnet(s) up front. A common home lab choice is 192.168.20.0/24 for path A and 192.168.21.0/24 for path B, with the hosts sitting on both.
  • Jumbo frames. If you want MTU 9000, the switch ports, the storage ports and the ESXi vmkernel interfaces must all agree. Mixed MTU "usually works" and then fails under load with no useful error message.
  • Drives. Count them. The pools and sub-groups you can build are entirely determined by the number of drives you have, and you want two left over for spares.

A simple IP plan, written down before configuration, saves an hour of guessing later. Mine looked like this: management on 10.0.0.2/10.0.0.3, iSCSI path A on 192.168.20.11-12 (controllers A1/A3 and B1/B3), iSCSI path B on 192.168.21.11-12, and the ESXi hosts on 192.168.20.50/192.168.21.50.

The Process

  • Login to the HPE SMU (Storage Management Utility).

Image 1

The default username is manage, and the default password is !manage. Point your browser at the management IP of controller A, log in, and change that password immediately — it is published in every manual on the internet. The newer MSA models ship a completely different web interface, so if your screens look nothing like these, you are on the wrong generation of firmware.

Network Setup

The first part we should do is configure the iSCSI network.

  • Click on System.

Image 2

  • Click on Action > Set Up Host Ports.

Image 3

  • Configure your ports as needed.

In my setup, ports A1, B1, A3, and B3 are iSCSI path A and ports A2, B2, A4, and B4 are iSCSI path B.

Image 4

The pattern worth copying is the one the SMU nudges you into: every port on controller A is cabled to switch A, every port on controller B to switch B, and the two fabrics are never cross-connected. A host then has two independent routes to the same LUN — one through fabric A, one through fabric B. If you cable everything into a single switch, you still have multipathing, but you no longer have redundancy, because the switch is now a single point of failure and the array will happily keep running while nothing can reach it.

Two more details on this screen. First, set the IP address, subnet mask and, if the iSCSI traffic has to leave its subnet, a gateway on each host port. Second, decide about jumbo frames now: if you enable MTU 9000 on the array, enable it on the switch ports and on the ESXi vmkernel adapters in the same change window, and verify with vmkping -s 8972 -d before you trust it.

Pool Setup

Next, we should set up our disk pools.

  • Click on Pools.

Image 5

  • Click on Action > Add Disk Group.

Image 6

  • Select the Type.

I will select Virtual.

Virtual is the modern option and the one you want: the array spreads data and spare capacity across the pool, so a rebuild after a drive failure is faster and less likely to hit a second drive. Linear is the legacy layout carried over from the MSA 2000 series, where each volume maps to a fixed set of physical disks. Pick Linear only if you have inherited volumes from an older array and you are importing them; for new deployments, Virtual is the default for a reason.

  • Select theRAID Level.

I will Select RAID-10.

For a home lab with a handful of large drives, the RAID level decision comes down to capacity versus risk tolerance. RAID-10 gives the best write performance and survives a single drive failure in each mirrored pair, at the cost of 50 per cent of your raw capacity. RAID-5 costs one drive of capacity and rebuilds slowly on large spinning disks; on 4TB and larger drives a rebuild can take a day, and a second failure during it loses the pool. RAID-6 survives two failures and is the safer middle ground for high-capacity nearline drives. The MSA 2040 supports RAID 0, 1, 5, 6, 10 and 50, and virtual pools add distributed sparing, which means you do not have to keep a dedicated hot spare idle inside the group.

  • Select the Pool.

The pool selection is for which controller will be the primary controller for the pool. It’s a good idea to make a disk group per controller to maximize the performance of the MSA.

  • Select the number of Sub-groups.

I will choose 6 for my first disk group on Controller A. On Controller B, I will select 5 disks, leaving me with 2 disks as spares. Whichever width you pick, pick it now: growing a pool later means adding drives to it, and you cannot shrink a sub-group without destroying and rebuilding it.

  • Name the disk group.

I will use the name dgA01. Names matter more than they look like they should. Once you have four pools across two controllers, dgA01 and dgB01 tell you at a glance which controller owns what, and that is exactly the information you need at 2 a.m. when a path is flapping.

  • Select the Disks.
  • Click Add.

Image 7

  • Click OK to confirm that the disk group is created.

Image 8

  • Repeat the process as needed.

I will repeat the process but select Pool B and 5 Sub-Groups.

Image 9

While the group is initialising, the SMU will show the pool in a degraded or initialising state. This is normal and on a large drive count it can run for hours. Do not start building volumes on top of a pool that is still initialising if you can avoid it, and do not reboot the array in the middle of it.

Spares

It’s a good idea to have spare disks in the event of a disk failure. When setting up the pools, I left out 2 disks. I will use those disks as global spares.

  • Click on System.

Image 10

  • Click on Action > Change Global Spares.

Image 11

  • Select the Disk you want to set as the Global Spares and click Change.

Image 12

  • Click OK to confirm that the global spares were set up.

Image 13

A global spare is available to any pool that needs it, which is convenient when your pools are on the same controller. A dedicated spare is reserved for one group only. In a small lab with two pools, global spares are the sane choice; in a large deployment where one pool is production and another is scratch, you may want the spare reserved so that scratch capacity cannot consume it.

Volume

We need to create a volume to store all the data on the disk group pools we created.

  • Click on Volumes.

Image 14

  • Click on Action > Create Virtual Volumes.

Image 15

  • Give your volume a name and specify the size and pool.

I will name my first volume, A-Vol01, make it the same size as the virtual pool, and select Pool A.

Image 16

  • Repeat the process as needed by clicking on Add Row.

I will make a second volume because I have a second pool on controller B. I will name the second volume B-Vol02, and I will set it to be on Pool B.

This one-volume-per-pool approach is the simplest thing that works in a lab, and it is also what most small deployments should do. Carving a pool into many small volumes only makes sense when you want to present different sizes to different hosts, or when you want per-volume snapshots. Every volume you add is another LUN to document, another object in the mapping table and another thing that can be attached to the wrong host group.

  • When ready, click OK to create the volume(s).

Image 17

  • Click OK to confirm the Volume(s) were created successfully.

Image 18

Hosts

For the HPE MSA 2040 to communicate with iSCSI hosts, you need to add them to the MSA as hosts.

  • Click on Hosts.

Image 19

  • Click on Action > Create Initiator.

Image 20

  • Enter the Initiator ID from the system connecting to the MSA and give the initiator a name.

Image 21

The initiator ID is the iSCSI Qualified Name of the host, and it must match character for character. On VMware ESXi you can read it from the software iSCSI adapter; on Linux it lives in /etc/iscsi/initiatorname.iscsi; on Windows it is shown in the iSCSI Initiator control panel under the Configuration tab. Copy and paste it rather than typing it — a single transposed character produces a host that is defined but never logs in, which is a genuinely annoying thing to debug because nothing in the logs says "typo".

  • Click OK to add the initiator.

Image 22

  • Click OKto confirm that the initiator was created successfully.

Image 23

  • Repeat the process for each system connecting to the HPE MSA for iSCSI.

Add Initiators to Host

  • Select one of the initiators you just added and click Action > Add to Host.

Image 24

  • Enter a name for the host and click OK.

I will use the initiator nickname as the host name.

Image 25

  • Click OKto confirm the initiator was added to the host successfully.

Image 26

  • Repeat the process for each initiator.

If a single ESXi host has more than one iSCSI initiator, which it should not in the normal software-iSCSI design, add each one and then attach them all to the same host object. One host object equals one set of LUNs; two host objects pointing at the same volume is how you end up with a datastore that appears twice and then corrupts.

Add to Host Group

  • Select all of your hosts.
  • Click on Action > Add to Host Group.

Image 27

  • Enter a name for the host group.

I will use the name g10-esx.

  • Click OK.

Image 28

  • Click OK to confirm that the hosts were added to the host group.

Image 29

A host group is what makes a volume look identical on every host in the group. That is exactly what you want for a shared VMFS datastore or a cluster: all the hosts see the same LUN number for the same volume, so vMotion and High Availability can move a virtual machine without remapping storage. Do not put an unrelated host, or a host you are about to reinstall, into the production group — it will inherit every mapping the group has.

Mapping

For the hosts to access the volumes, we need to map the hosts to a volume.

  • Click on Mapping.

Image 30

  • Click on Action > Map.

Image 31

  • Select the Group you want to map and the Volume to which you want to map the group.

I will select the Group named g10-esx, and I will select both Volume A-Vol01 and B-Vol2

  • Click Map.

Image 32

By default, the MSA 2040 will want to use LUN 0. You should change it as with VMware vSphere LUN 0 will be used by the MPE MSA storage enclosure.

  • Change the LUNs to make sure they don’t use LUN 0.
  • Click OKto initiate the mappings.

Image 33

  • Click OK to confirm the mapping succeeded.

Image 34

Pick LUN numbers deliberately rather than letting the array assign them. A stable convention such as LUN 1 for the first datastore, LUN 2 for the second and so on means that the mapping table in SMU and the device list in ESXi can be read side by side. It also makes the next rebuild of a host far less exciting, because the LUN numbers it expects are the ones you already documented.

That’s it. Now, all you need to do is set up your hosts to access the HPE MSA.

Bringing the LUNs up in ESXi

The array side is done. On the host side, enable the software iSCSI adapter, add the array as a discovery target, rescan and set the multipathing policy. The commands below are the ones you would type if you prefer the CLI over the vSphere Client:

esxcli iscsi software set --enabled=true
esxcli iscsi adapter list
esxcli iscsi adapter discovery sendtarget add --address=192.168.20.11:3260 --adapter=vmhba64
esxcli iscsi adapter discovery sendtarget add --address=192.168.21.11:3260 --adapter=vmhba64
esxcli iscsi adapter discovery sendtarget list --adapter=vmhba64
esxcli storage core adapter rescan --adapter=vmhba64
esxcli storage core device list
esxcli storage core path list

Add both controller IPs as static discovery targets. Dynamic discovery, where the host asks the array for its own portal list, works fine too, but static targets fail in a much more predictable way: a host that cannot reach its configured target says so, whereas one waiting on a discovery record sometimes just silently sees nothing.

Once the devices appear, check how many paths each one has. A healthy LUN on this array should show two, one per fabric. If you see one, stop and fix it before you put any data on it, because a single-path LUN turns a controller firmware upgrade into downtime. The default policy is Most Recently Used, which keeps all the traffic on one path until it fails; Round Robin spreads load across both paths:

esxcli storage nmp device list
esxcli storage nmp psp roundrobin deviceconfig set --type=iops --iops=1 --device=naa.6xxxxxxxxxxxxxxx

Then create the VMFS datastore on the LUN and, if you have two volumes, decide whether to stripe them across both controllers or to keep them as two datastores. Two separate datastores are usually the better answer in a lab: you can rebuild or evacuate one without touching the other.

End-to-end verification checklist

Work through this list once, top to bottom, and you will find the problem in a minute rather than an evening:

  • SMU shows both controllers present and the firmware versions match.
  • Every host port shows link up at the expected speed, and A-side ports live on switch A while B-side ports live on switch B.
  • The pool is online, not initialising or degraded, and its owning controller is the one you intended.
  • Volumes are online, the expected size, and owned by the pool you created them in.
  • Each initiator shows a logged-in status rather than "not logged in".
  • The mapping table lists every volume you expect, on LUN numbers other than 0.
  • ESXi sees each LUN twice, once per fabric, with no dead paths.

Troubleshooting the usual suspects

  • The host sees no LUNs at all. Almost always the initiator: an IQN typo, or the initiator was created but never added to a host, or the volume was never mapped to the group. Check in that order.
  • The LUN appeared once, then vanished. Look for an MTU mismatch. Set both ends to 1500 and retest; if it becomes stable, jumbo frames were the culprit.
  • Only one path per LUN. One fabric is down, mis-cabled, or one controller IP was never added as a discovery target. Verify with a rescan and check the path list again.
  • The datastore is read-only or the VM will not move. The host is looking at a LUN that belongs to a volume mapped to the wrong host group, or the host is missing from the group. Compare the SMU mapping table against the host's SCSI device list.
  • Everything looks right but performance is poor. Confirm the multipathing policy is Round Robin and that traffic is not all riding one controller. A pool that lives entirely on controller A cannot use controller B's ports.

Maintenance habits worth keeping

Write the mapping table down somewhere outside SMU — a wiki page, a text file in your repository, anywhere. When the array is replaced or the lab is rebuilt, that document is the difference between re-creating the environment in an afternoon and re-discovering it over a week. Take a configuration backup from the SMU on every change, keep both controllers on the same firmware, and run a scrub or at least a scheduled disk check so that a failing drive is reported while you still have a spare in the chassis.

If you want to read more about setting up the HPE MSA 2040, here is the HPE documentation.

Related reading on this site: diagnosing HPE 3PAR storage issues in a VMware environment, setting a static IP on ESXi, backing up an ESXi configuration and VMware home lab licensing.