SUN M4000服务器操作手册笔记

SUN M4000操作手册笔记

The SUN M4000 (marketed as the SPARC Enterprise M4000, and built on the same midrange chassis family as the Fujitsu M4000 / M5000) is a rack-mounted SPARC server where almost everything you do as an operator happens through the XSCF — the eXtended System Control Facility. The XSCF is an independent service processor with its own firmware, its own network interface, and its own CLI. It owns power sequencing, the hardware inventory, the fault manager, the environmental monitoring, and the serial console to the host domains. If you understand XSCF you can commission, inspect, and recover the whole machine without ever seeing Solaris boot. This article is a complete operation-manual walkthrough: power and access prerequisites, the panel mode switch dance for the first login, the hardware inventory commands, the firmware version check, fault and error log reading, dropping to the OpenBoot prompt, powering on a domain, and a command reference you can keep open on a second screen.

1. What the XSCF Controls

It helps to draw a hard line between the two computers inside the chassis. The service processor (XSCF) is always alive as long as the server is plugged into power; the host (the SPARC domains running Oracle Solaris) is only alive when you or an auto-power-on policy has turned a domain on. The XSCF therefore answers questions the host cannot answer when it is down: which FRUs are installed, which ones are faulted, what the last fault event was, whether a fan or power supply has degraded, and what the ambient/CPU temperatures look like.

Practically, the XSCF is also the out-of-band management path. You reach it over a serial cable on the front-panel port, over the LAN through its own IP address (SSH or telnet), or by a console server. It never depends on the host being healthy, which is exactly why it is the first thing you connect to on an unfamiliar machine.

2. Prerequisites Before You Touch the Server

  • Serial console or LAN access to the XSCF — a NULL-modem serial cable to the front panel (9600 8N1) is the most reliable first contact; the LAN route needs the XSCF IP address, netmask and gateway, which may only be discoverable from the serial console.
  • The physical key for the operator panel — the M4000 front panel has a mode switch (Locked / Service / Remote) that must be physically turned with the key. You cannot complete a first login without hands on the machine.
  • An account. The factory default login on a virgin system is default with no password; the platform administration account created during installation is usually admin.
  • The product test record / FRU list that shipped with the server — you will use it to compare installed FRUs against what was delivered.
  • Change control. Anything that changes the panel switch position or cycles power on a production domain should be announced; the XSCF will happily let you power off a live domain if you ask it to.

If you are coming from a T-series machine, the philosophy is very close to the ILOM service processor on the T-series, but the command set and the panel switch procedure are different enough that you should not assume muscle memory carries over.

3. First Login and the Panel Mode Switch Procedure

The original notes from the field, kept verbatim so the exact prompts are preserved:

第一次进入,login:default

然后会有如下提示

Change the panel mode switch to Locked and press return...

(此时将钥匙插入旋转到Lock档)

键入回车键

然后系统会有如下提示

Leave it in that position for at least 5 seconds. Change the panel mode switch to Service, and press return...

(至少5秒后,再将钥匙切换到Service档,键入回车键 )

In plain English the sequence is:

  1. Login with the default account when the console presents a login prompt.
  2. The XSCF prints Change the panel mode switch to Locked and press return... — insert the key and turn the front-panel mode switch to the Locked position, then press Enter.
  3. The XSCF then prints Leave it in that position for at least 5 seconds. Change the panel mode switch to Service, and press return... — hold Locked for a minimum of five seconds (do not shortcut this), then turn the switch to Service and press Enter again.

Once that handshake completes you are inside the XSCF command shell. Forgetting the five-second dwell, or turning the key too slowly so the intermediate position is skipped, produces a re-prompt and the sequence starts again — this is the single most common reason people think "the console is broken".

4. Hardware Inventory: showhardconf

First things first: confirm what is actually in the chassis and whether anything is flagged. The primary command is showhardconf.

showhardconf -M
showhardconf -u
version -c xcp            检查XCP版本

The -M flag narrows the output to the main unit view; showhardconf -u gives the unit-by-unit breakdown. A trained eye scans for two things: any FRU that cannot be seen at all, and any FRU carrying an asterisk (*) next to its name — an asterisk means the XSCF knows the slot is occupied but the component is not healthy or is not responding. A clean machine shows every expected FRU with no asterisk.

Then compare the installed count against the "product test record" that came with the server. If the record says four CPU modules and showhardconf -u reports three, stop and resolve that before you put the machine into service — a missing module will change your domain configuration options.

5. Firmware Version: version -c xcp

version -c xcp

This prints the XCP (XSCF Control Package) firmware revision. Record it. Every subsequent conversation with support, every firmware-matrix decision, and every decision about whether you can attach a particular FRU depends on the XCP level. Mixing FRUs that require a newer XCP than the one installed is a classic way to end up with a component that powers up but never appears in showhardconf.

6. Fault Logs: fmdump

The Solaris fault manager inside the XSCF records fault events, and fmdump is the reader. The original captured output:

XSCF> fmdump
TIME UUID MSG-ID
Jul 07 16:28:14.4550 469abb4a-92c8-4257-852c-483f036de9ef SCF-8001-KC
Jul 07 16:28:36.5370 33f6363d-3f14-4c79-890a-e2dea50c87ab SCF-8002-CY
Jun 13 21:42:34.3129 8ee5141b-70e1-4a5f-824a-3d113d51842c SCF-8006-3J
Feb 25 17:19:03.3989 c847ad3c-1ae0-454e-94c1-c7ebf9a89ade SCF-8001-KC

fmdump -V -u

Four columns matter: the timestamp, the UUID that uniquely identifies the event, the message ID (the SCF-8001-KC style code that maps to a documented fault dictionary entry), and — visible with the verbose form — the confidence percentage and the FRU that is implicated.

showlogs -r error

showlogs -r error filters the log to error-severity messages only. Use it as the routine check: fmdump tells you a fault was recorded, showlogs -r error tells you whether the platform is currently complaining about something else as well.

7. Reading a Fault in Detail: fmdump -V

The verbose form is where diagnosis actually happens. The manual's own example, with the accompanying status output:

Using the fmdump Command
The fmdump command can be used to display the contents of any log files associated
with the Oracle Solaris fault manager.
This example assumes there is only one fault.
B.2.4.1 fmdump -V Command
You can obtain more detail by using the -V option, as shown in the following
example.

XSCF> showstatus
FANBP_C Status:Normal;
* FAN_A#0 Status:Faulted;
XSCF>

# fmdump
TIME UUID SUNW-MSG-ID
Nov 02 10:04:15.4911 0ee65618-2218-4997-c0dc-b5c410ed8ec2 SUN4-8000-0Y

# fmdump -V -u 0ee65618-2218-4997-c0dc-b5c410ed8ec2
TIME UUID SUNW-MSG-ID
Nov 02 10:04:15.4911 0ee65618-2218-4997-c0dc-b5c410ed8ec2 SUN4-8000-0Y
100% fault.io.fire.asic
FRU: hc://product-id=SUNW,A70/motherboard=0
rsrc: hc:///motherboard=0/hostbridge=0/pciexrc=0

Read this out loud as a story, because that is how it is designed: showstatus shows a healthy fan backplane and a faulted fan, the fault list gives you a UUID, and fmdump -V -u <uuid> resolves that UUID into a fault class (fault.io.fire.asic), a confidence of 100%, the responsible FRU, and the resource path. The hc:// path is the machine-readable location — motherboard=0, hostbridge=0, pciexrc=0 — which is what you read off the chassis when you go to the rack to replace something.

8. Dropping to the OpenBoot (ok) Prompt and Powering On a Domain

console -d 00
poweron -d 0

console -d 00 attaches your XSCF session to the serial console of domain 0 (the trailing 0 is the domain selector; on a single-domain machine it is always zero, on a split configuration you will see -d 1 and beyond). poweron -d 0 powers domain 0 up. From the OpenBoot prompt you then run the hardware self-checks before letting Solaris boot:

probe-scsi-all
show-devs

9. Command Reference

The original checklist, preserved verbatim, with the prompt each command is issued from:

命令 提示符 说明
showhardconf XSCF Shell 此时将显示服务器中安装的所有组件及其状态。确认任何 FRU 的前面都未显示
星号 (*)。
showhardconf -u XSCF Shell 对照服务器附带的 “产品测试记录”检查服务器中装配的 FRU 数。
probe-scsi-all ok 提示符 确认可以识别服务器中安装的 CD-RW/DVD-RW 驱动器单元和硬盘驱动器。
show-devs ok 提示符 确认可以识别安装的每个 PCIe 卡。

Translated and expanded into a reference table of intent:

  • showhardconf (XSCF shell) — displays every installed component and its status. If any FRU has an asterisk in front of it, that FRU is not healthy.
  • showhardconf -u (XSCF shell) — compare against the product test record to confirm the FRU count the server was delivered with.
  • probe-scsi-all (ok prompt) — confirm the CD-RW/DVD-RW drive unit and the internal disk drives are visible to the domain.
  • show-devs (ok prompt) — confirm every installed PCIe card is enumerated.

10. A Commissioning Walkthrough, In Order

  1. Connect serial to the front panel, confirm 9600 8N1, and press Enter until you see a prompt.
  2. Log in as default, complete the Locked → wait 5 s → Service panel switch sequence.
  3. version -c xcp — record the XCP level.
  4. showhardconf — look for asterisks. showhardconf -u — reconcile counts against the test record.
  5. showlogs -r error and fmdump — get a baseline of existing faults before you change anything, so you cannot be blamed for pre-existing ones.
  6. console -d 00 to reach the ok prompt, then probe-scsi-all and show-devs.
  7. Return to the XSCF shell (#. or the escape sequence), then poweron -d 0.
  8. Watch the domain boot, then leave the panel switch in Locked so no one can accidentally change service state.

11. Mistakes That Cost Real Time

  • Skipping the five-second dwell on the Locked position — the XSCF rejects the mode change and you loop.
  • Reading showhardconf output too fast. The asterisk is one character wide and it is the whole point of the command.
  • Assuming an absent FRU is broken. A missing CPU module also shows up as a configuration shortfall, not a fault; check the test record before opening a support case.
  • Not capturing the UUID. Without the UUID you cannot run fmdump -V -u, and without the verbose output you cannot tell the field engineer which FRU to bring.
  • Leaving the key in Service. Service mode disables some protective interlocks; the machine should be returned to Locked when work is done.

12. Related Reading

If you also run newer Oracle hardware, the reset and access procedures differ: see T5 series machine reset and SUN T7-1 notes. When you are troubleshooting the host rather than the service processor, the same read-a-log-then-map-it-to-hardware discipline applies to the network side of the rack, e.g. Juniper MX hardware issues, damping and holdtime.