SmokePing: Latency and Packet Loss Monitoring - 夜莺博客

SmokePing: Latency and Packet Loss Monitoring

Uptime checks answer "is it up"; they say nothing about a link that has been degrading for three weeks. SmokePing measures round-trip time, its distribution and packet loss on a schedule, stores everything in RRDtool and draws it as latency bands, so a creeping problem is visible long before users complain. It is one of the few monitoring tools that is genuinely cheap to run on a small VM. This guide covers a useful deployment rather than a demo.

Install and lay out the configuration

sudo apt-get install smokeping apache2 fping traceroute
sudo cp /etc/smokeping/config.d/Targets /etc/smokeping/config.d/Targets.orig
sudo systemctl enable --now smokeping

Debian and Ubuntu split the configuration into config.d/: General (owner, contact, mailhost), Alerts, Database, Probes, Slaves and Targets. Set imgurl and cgiurl in General to the addresses clients will use, otherwise the graphs render with broken links.

Define the target hierarchy

*** Targets ***
probe = FPing
menu = Top
title = Network Latency Grapher

+ Core
menu = Core switches
title = Core infrastructure

++ core-sw1
menu = core-sw1
host = 10.10.10.1
alerts = someloss,hostdown

+ Uplinks
menu = Uplinks

++ transit-a
menu = Transit A
host = 203.0.113.1
alerts = bigloss,hostdown

++ transit-b
menu = Transit B
host = 198.51.100.1
alerts = bigloss,hostdown

Each + level creates a menu branch. Group by function - core, uplinks, branches, cloud VPN tunnels - so the first page answers "which path is bad" without clicking.

Probes worth adding

  • FPing - the default, ICMP latency and loss to many hosts from one process.
  • DNS / EchoPingHttp - real application latency, not just ICMP; HTTP probes measure the time to fetch a URL, which is what users experience.
  • TCPConnect - proves a specific port is answerable and times the handshake, useful for firewalled paths where ICMP is blocked.

Where ICMP is filtered, a TCP probe is the difference between a monitoring graph and a flat red line.

Alerts that mean something

*** Alerts ***
to = noc@example.com
from = smokeping@monitor1

+someloss
type = loss
pattern = >0%,*12*,>0%,*12*,>0%
comment = packet loss in 3 of the last 12 samples

+bigloss
type = loss
pattern = ==0%,==0%,==0%,>0%,>0%,>0%
comment = sudden sustained packet loss

+hostdown
type = loss
pattern = ==0%,==0%,==0%,==0%,==0%,>20%
comment = host appears unreachable

Loss patterns are evaluated against consecutive samples, so they ignore single dropped pings. RTT alerts use the same syntax with type = rtt and values in seconds plus edgetrigger if you want one alert per event rather than per sample.

Verify and maintain

  • smokeping --check validates the configuration before a reload; smokeping --reload applies target changes without a full restart.
  • Check the Database section's RRAs: default steps keep five-minute granularity for around a year, which is the right trade-off for capacity planning.
  • For monitoring beyond one location, enable Slaves and keep the shared secret file unreadable; a slave proves a path from a second vantage point, which is the only way to tell "the link is bad" from "my monitor is bad".
  • Combine with mtr Command Guide: Diagnosing Packet Loss and Latency when a graph shows loss but you need per-hop attribution, and with flow data from netflow analysis collector when you need to know whose traffic was affected.

Related reading on this site: mtr Command Guide: Diagnosing Packet Loss and Latency and Zabbix SNMP Switch Monitoring with LLD.

原文链接:https://std.rocks/gnulinux_smokeping.html