Nagios Core NRPE Remote Host Monitoring Setup - 夜莺博客

Nagios Core NRPE Remote Host Monitoring Setup

Nagios Core is not a metrics database — it is a state engine. Every check returns OK, WARNING, CRITICAL or UNKNOWN, and everything that drives operations (notifications, escalations, flap detection) hangs off state transitions. That model is why NRPE still matters on Linux hosts: disk space, load average, process counts and logged-in users are signals that only exist on the monitored machine, so the agent runs the plugin locally and returns a status and a performance-data string. Here is the complete setup, plus the tuning that stops alert storms.

Install on Both Sides

# Nagios Core server
sudo apt install nagios4 nagios-nrpe-plugin monitoring-plugins

# Monitored Linux host
sudo apt install nagios-nrpe-server monitoring-plugins-basic monitoring-plugins-standard

# Nagios server hardening basics
sudo htpasswd /usr/local/nagios/etc/htpasswd.users nagiosadmin
sudo systemctl enable --now nagios

NRPE Access Control Is the Security Boundary

### /etc/nagios/nrpe.cfg on the monitored host
allowed_hosts=127.0.0.1,::1,203.0.113.10
dont_blame_nrpe=0

command[check_disk_root]=/usr/lib/nagios/plugins/check_disk -w 20% -c 10% -p /
command[check_load]=/usr/lib/nagios/plugins/check_load -w 5,4,3 -c 10,8,6
command[check_total_procs]=/usr/lib/nagios/plugins/check_procs -w 150 -c 200
command[check_zombie_procs]=/usr/lib/nagios/plugins/check_procs -w 5 -c 10 -s Z
  • allowed_hosts must list the monitoring server's address; the default of 127.0.0.1 is why a fresh install reports "CHECK_NRPE: Error - Could not complete SSL handshake / connection refused".
  • dont_blame_nrpe=0 is non-negotiable: enabling command arguments lets a caller pass shell metacharacters through the check command. Define every acceptable check as a named command instead.
  • Open 5666/tcp only from the monitoring host: sudo ufw allow from 203.0.113.10 to any port 5666.

Object Definitions on the Nagios Server

define command {
  command_name check_nrpe
  command_line $USER1$/check_nrpe -H $HOSTADDRESS$ -c $ARG1$
}

define host {
  use        linux-server
  host_name  web01.example.net
  alias      Web server 01
  address    web01.example.net
}

define service {
  use                    generic-service
  host_name              web01.example.net
  service_description    Disk Usage /
  check_command          check_nrpe!check_disk_root
  check_interval         5
  retry_interval         1
  max_check_attempts     3
}

define service {
  use                    generic-service
  host_name              web01.example.net
  service_description    CPU Load
  check_command          check_nrpe!check_load
}

Do not define the same host_name in two object files — Nagios refuses to start with duplicate object errors, and the validation output names the offending file.

Validate, Test, Reload

sudo nagios4 -v /etc/nagios4/nagios.cfg
$USER1$/check_nrpe -H web01.example.net -c check_disk_root
sudo systemctl reload nagios

Test the NRPE command manually before reloading: a working manual check with a failing GUI check almost always means a missing command definition or a typo in the service's check_command.

Tuning Against Alert Fatigue

  • check_interval vs retry_interval: keep max_check_attempts at 3 and a short retry interval so a transient spike does not page anyone; this single change removes most false CRITICALs.
  • Flap detection: enable it for checks that oscillate between OK and CRITICAL, such as memory checks on hosts with aggressive caches.
  • Performance data: if you need trend graphs rather than current state, export the plugin performance-data string to a metrics system instead of raising the check frequency.
  • NCPA over NRPE: where you need dynamic thresholds or remote command execution, migrate to NCPA with API tokens rather than relaxing the NRPE security model.

Related reading: Alertmanager routing and notification configuration and Linux NFS v4 server configuration.

原文链接:https://simplified.guide/nagios/linux-host-monitor-nrpe