UPS Monitoring Over SNMP: Power and Runtime Alerting - 夜莺博客

UPS Monitoring Over SNMP: Power and Runtime Alerting

Servers, switches and storage all report their health, and the UPS that keeps them alive usually does not - until the day it fails during an outage. The UPS MIB (RFC 1628) is old, small and supported by every major vendor, which makes it the cheapest reliability win in a data centre monitoring stack.

What the MIB Actually Gives You

Object OID suffix Why it matters
upsIdentModel / Manufacturer .1.3.6.1.2.1.33.1.1.x Inventory and template matching
upsBatteryStatus .2.1.1.0 unknown / batteryNormal / batteryLow / batteryDepleted
upsEstimatedMinutesRemaining .2.3.0 The number everyone asks for during an outage
upsSecondsOnBattery .2.2.0 How long you have already been running on battery
upsInputLineBads .3.2.0 Mains quality - a rising counter means a failing feed
upsOutputLoad (%) .4.4.0 Headroom for the next hardware install
upsOutputSource .4.1.0 normal / battery / bypass / booster
upsAlarmTable (walk) .6.2.1 Active alarm descriptions in plain text
snmpwalk -v2c -c RO 10.10.10.210 1.3.6.1.2.1.33 | head -40
snmpget  -v2c -c RO 10.10.10.210 1.3.6.1.2.1.33.1.2.3.0   # minutes remaining
snmpget  -v2c -c RO 10.10.10.210 1.3.6.1.2.1.33.1.4.4.0   # output load %

Vendors extend the standard tree heavily (APC has an enterprise MIB with battery pack detail, temperature and per-outlet control; others similar). Start with RFC 1628 so one template covers every brand, then add vendor objects only where they add something - typically battery replacement date, battery temperature and per-outlet state.

The Five Alerts Worth Wiring

  1. On battery. upsOutputSource != normal for more than 30 seconds. This is the alert that starts the incident procedure.
  2. Battery low. upsBatteryStatus = batteryLow. From here you are minutes from load loss.
  3. Runtime below threshold. upsEstimatedMinutesRemaining < X, where X is your design runtime minus a safety margin - not just 'under 10 minutes'.
  4. Load above threshold. upsOutputLoad > 80 sustained for 15 minutes.
  5. Self-test failed / battery needs replacing. Poll the alarm table and the vendor replacement-date object. A UPS with a dead battery is a power strip with extra steps.

Runtime Is Not a Constant

The reason runtime monitoring beats a one-off calculation: upsEstimatedMinutesRemaining is derived from load and battery state, so the same UPS running 40 minutes at 30% load may give 9 minutes at 80%. Baseline it at your real load, re-check after every hardware addition, and alert on the trend as well as the absolute value. Battery ageing shows up as slowly shrinking runtime months before the self-test fails.

Wiring It Into Your Monitoring

The pattern is identical whichever tool you use - Zabbix, Prometheus with the SNMP exporter, LibreNMS or a commercial NMS:

  • One host per UPS with the management card IP (never the serial/USB port, unless the card is the only option).
  • Import or build a template from the RFC 1628 objects above; clone it per vendor only for the extended objects.
  • Set dependency so the UPS alarm triggers before the cascade of host-down alerts; consider suppressing or tagging downstream alerts while on battery.
  • Send traps as well as polling: many cards emit an SNMP trap immediately on mains loss, which is faster than a 60-second poll interval.

Don't Forget the Non-UPS Power Path

Two things bite more often than dead batteries: a PDU or breaker that trips while everything else is healthy, and a UPS whose bypass is active (so it is passing raw mains while reporting 'up'). Read upsOutputSource and the alarm table together, and monitor PDUs that have their own SNMP card for per-outlet current. If your monitoring stack is already collecting device data, the Zabbix SNMP template workflow applies directly, and OID troubleshooting covers the case where a walk returns nothing at all.

原文链接:https://datatracker.ietf.org/doc/html/rfc1628