Icinga 2 Monitoring Configuration Guide - 夜莺博客

Icinga 2 Monitoring Configuration Guide

Icinga 2 occupies a useful niche between Zabbix's agent-and-database model and Prometheus's metrics-only world: it is a mature, configuration-driven monitoring system built for check-based alerting, with a proper object model, dependencies, notification escalation and distributed satellite design. This guide covers the object model, a working host-and-service configuration, how to monitor network devices without an agent, and how to scale with satellites.

Where Icinga fits

Prometheus   pull-based metrics, great for trends and SLOs, weak at "is it up?"
Zabbix       agent + SNMP, built-in graphin and templates, heavier
Icinga 2     check-based, object model + templates, satellites, strong
             notification/escalation logic, good at hybrid envs

Most mature shops run both: metrics in Prometheus (see Prometheus SNMP exporter), up/down and threshold alerting in Icinga. Comparisons with the agent-based approach are in Zabbix proxy monitoring and LibreNMS.

Install the master

# Debian/Ubuntu
apt install icinga2 icinga2-ido-mysql monitoring-plugins
mysql -e "CREATE DATABASE icinga2; CREATE USER 'icinga2'@'localhost' IDENTIFIED BY 'pw';"
icinga2 feature enable ido-mysql
icinga2 feature enable command
systemctl restart icinga2

# web interface (Icinga Web 2) plus the director module if you want config in the UI
apt install icingaweb2

Objects and templates

object Host "sw-core-01" {
  import "generic-host"
  address = "10.10.10.1"
  check_command = "hostalive"
  vars.os = "network"
  vars.snmp_community = "monitor"
}

template Host "network-device" {
  import "generic-host"
  check_command = "hostalive"
  vars.snmp_community = "monitor"
  vars.notification["mail"] = { groups = [ "netops" ] }
}
object Host "sw-acc-14" { import "network-device" address = "10.10.10.14" }

Templates are the whole point of Icinga's design: define the pattern once, inherit it everywhere, and override only what differs.

Services and apply rules

apply Service "ping4" {
  import "generic-service"
  check_command = "ping4"
  assign where host.vars.os == "network"
}

apply Service "if-traffic" {
  import "generic-service"
  check_command = "snmp-interface"
  vars.snmp_interface = "GigabitEthernet1/0/1"
  assign where host.name == "sw-core-01"
}

Apply rules remove the need to list services per host, which is what makes the configuration maintainable at hundreds of hosts. For the SNMP check itself, the standard approach is the Icinga snmp plugin family or the check_snmp command with OIDs.

Notifications and escalation

object NotificationCommand "mail-service-notification" {
  command = [ "/etc/icinga2/scripts/mail-notification.sh" ]
}

apply Notification "email-netops" to Service {
  command = "mail-service-notification"
  users = [ "oncall" ]
  interval = 15m
  assign where host.vars.os == "network"
}

# escalation: page only if still broken after 30 minutes
object TimePeriod "24x7" { ranges = { monday = "00:00-24:00" ... } }
apply Notification "page-late" to Service { ... interval = 4h ... }

Two settings prevent most alert fatigue: a re-notification interval longer than the check interval, and dependencies so that a failed switch does not generate an alert for every host behind it.

Dependencies kill the alert storm

apply Dependency "uplink" to Host {
  parent_host_name = "sw-core-01"
  disable_checks = true
  disable_notifications = true
  assign where host.vars.uplink == "sw-core-01"
}

Without dependencies, losing one core switch produces hundreds of host-down alerts. This single feature is usually what justifies Icinga over a simple uptime checker.

Scaling out with satellites and agents

Master zones = ["master"]
Satellite zones = ["site-a", "site-b"]
Agent zones = ["site-a/host-42"]

Master <-> Satellite:  Icinga 2 cluster protocol (port 5665)
Satellite <-> Agent:   same protocol; checks execute close to the target
Benefit: cross-WAN check latency disappears; the master only aggregates
icinga2 node wizard        # on each satellite/agent: ensure to join the right zone
icinga2 daemon -C          # validate config before restart
icinga2 feature list

icinga2 daemon -C is the essential habit: it validates the object graph and prints errors with file and line before you restart a production master.

Pitfalls

Check interval shorter than reality   flaps; align interval with the SLA
No flap detection                     a bouncing link pages all night
Everything in the master zone         WAN latency pollutes check results
SNMP v2 community in plaintext        use v3 (see the SNMPv3 articles)
Director edits vs flat config files   pick one source of truth

FAQ

Q: Can Icinga read Prometheus metrics? Yes, via a check that queries Prometheus, but it duplicates data; prefer Prometheus alerting for metric thresholds and Icinga for availability.
Q: Is the Director required? No, but UI-driven configuration is much easier for a team; flat files plus Git is better for automation-heavy shops.
Q: How many hosts per master? Low thousands with satellites; beyond that add a second master for HA and distribute zones.

原文链接:https://github.com/Icinga/icinga2