APC UPS Network Management Card: SNMP Monitoring Setup - 夜莺博客

APC UPS Network Management Card: SNMP Monitoring Setup

A UPS without network monitoring is a device you only learn about when the rack goes dark. An APC Network Management Card (NMC) — AP9630, AP9631, AP9635 and the current NMC 3 family — turns a Smart-UPS into a monitored node with SNMP, web UI, email and a shutdown-agent integration. This guide covers the network and SNMP configuration that makes the card useful, the trap receivers worth enabling, the OID branch to poll, and the coordination between polling and automatic shutdown.

First, the identity check

Start with one OID to confirm you are talking to the right device and to the right firmware:

snmpget -v2c -c public 10.10.10.30 1.3.6.1.2.1.1.1.0
# SNMPv2-MIB::sysDescr.0 = STRING: APC Web/SNMP Management Card (MB:v4.1.0 PF:v6.8.2 PN:apc_hw05_aos_682.bin ...

1.3.6.1.2.1.1.1.0 is the standard sysDescr object, and on an APC card it returns hardware model, SKU, serial number, firmware levels for the APC OS, the application module and the boot monitor. That single line tells you whether a firmware upgrade is warranted before you spend time chasing a behaviour that was fixed three releases ago.

The vendor-specific branch begins at 1.3.6.1.4.1.318, which is the identifier your monitoring system should use for APC-specific discovery so unknown devices are not misclassified.

Configure SNMP on the card

In the card's web interface, SNMPv1 access is enabled by default — which is also the first thing to change:

  1. Configuration → Network → SNMPv1 → Access: verify whether community-based access is enabled. If you must keep it for legacy tools, restrict the Access Control entries so the community string is accepted only from your monitoring server's subnet, and change the default public string.
  2. Configuration → Network → SNMPv3: create at least one v3 user with authentication and privacy (AES), and set read-only access unless a template requires write.
  3. Configuration → Notification → SNMP Traps → Trap Receivers: add your NMS.

Prefer SNMPv3 for anything crossing an untrusted segment. Community strings travel in clear text and are trivially captured; v3 with auth and priv is the only defensible option on a shared network.

Trap receivers: the part that gives you early warning

Polling tells you the state at the last interval; traps tell you what changed. Configure a trap receiver with the monitoring host's address, enable trap generation, and enable authentication of the traps themselves so a random host cannot spoof an "all clear".

The card's own test function is the fastest way to prove the path works end to end: use the trap test entry for the receiver you just created, then confirm the trap arrived in your NMS. If it did not, the problem is a firewall rule, a wrong community or a wrong v3 user — not the UPS.

Events worth alerting on, in rough priority order:

  • On battery — the moment power is lost; everything else is a consequence.
  • Battery low — the shutdown threshold is approaching.
  • Communications lost between NMC and NMS, and between NMC and the UPS itself. The second is easy to miss and means you have lost visibility of a device that is still running.
  • Input voltage out of range, internal temperature exceeded, battery needs replacement.
  • Minimum redundancy lost on a redundant UPS configuration.

Polling: pick few, alert on fewer

Do not poll the whole UPS MIB at one-minute intervals; you will generate load on a device with modest CPU and a lot of noise. A compact set covers most operational needs: input voltage, output load percentage, battery capacity, battery voltage, runtime remaining, internal temperature, and the output status flags. Trend runtime and battery replacement date — they are the two values that predict a surprise outage months in advance.

Pair polling with graceful shutdown

Monitoring without shutdown automation protects the hardware, not the data. PowerChute Network Shutdown registers with the NMC and initiates an orderly shutdown of connected hosts — including virtualised environments — when a critical event persists. Two configuration details make that reliable:

  • Delay before action. Set the critical event delay long enough that a 20-second brownout does not trigger a shutdown, but shorter than your runtime at expected load.
  • Repeat interval. Configure the trap or notification to repeat, so a missed trap does not mean a missed event.

Then test it: cut utility power with a load you can afford to lose, and watch the shutdown sequence run. An untested shutdown path is a hypothesis.

Where this fits in your monitoring stack

If the NMS is Zabbix, the SNMP template workflow in Zabbix SNMP network device monitoring applies unchanged to the UPS, and the device-specific checks you build are the same style as Prometheus SNMP exporter for switches and routers. If you prefer Prometheus, the SNMP exporter described in Prometheus SNMP exporter for switches and routers polls the same OIDs and gives you alerting through Alertmanager. Either way, add the UPS to the same on-call rotation as the network gear — a rack that loses power but nobody notices is exactly the failure mode monitoring exists to prevent.

原文链接:https://www.se.com/us/en/faqs/FAQ000277838/