rsyslog for Network Device Logs: Reliable Collection - 夜莺博客

rsyslog for Network Device Logs: Reliable Collection

Every device on the network can send syslog, and most teams point it at whatever host answers on UDP 514 and forget it. Then an incident happens, and the logs are incomplete - because a UDP datagram was dropped during the flood of the failure, or because the receiver hit its rate limit, or because the file rotated out three days ago. This is a receiver configuration that does not lose the important lines.

Transport: Move Off UDP Where You Can

Transport Reliability Use when
UDP 514 Lossy, no delivery guarantee, spoofable Legacy devices, non-critical logs
TCP 514 Ordered, retransmitted; device buffers during outages Default choice for network gear
TLS 6514 Encrypted and authenticated Any cross-network path or compliance scope

Devices that support TCP logging will queue messages locally when the link is down, which is exactly the behaviour you need when the failure you are investigating also broke the path to the log server.

# Cisco IOS
logging host transport tcp 10.10.10.80
logging trap informational
logging source-interface Loopback0
logging buffered 65536 informational

# Junos
set system syslog host 10.10.10.80 any info
set system syslog host 10.10.10.80 port 514
set system syslog host 10.10.10.80 source-address 10.0.0.1

A Receiver That Does Not Drop Messages

# /etc/rsyslog.d/10-network.conf
module(load="imtcp")
input(type="imtcp" port="514" ruleset="netdev")

# TLS variant
# module(load="imtcp" StreamDriver="gtls" StreamDriverMode="1"
#        StreamDriverAuthMode="x509/name" StreamDriverPermittedPeers="*.example.com")
# input(type="imtcp" port="6514" ruleset="netdev")

# keep the listener working during bursts
$SystemLogRateLimitInterval 0

ruleset(name="netdev") {
    # one file per device, named by the sending host
    action(type="omfile"
           dynaFile="netlogs/%HOSTNAME%/%$YEAR%-%$MONTH%-%$DAY%.log"
           fileCreateMode="0640"
           dirCreateMode="0750"
           template="RSYSLOG_FileFormat")

    # forward equipment alarms separately for alerting
    if ($msg contains "%LINK-3-UPDOWN" or $msg contains "BGP-5-ADJCHANGE") then {
        action(type="omfile" dynaFile="netlogs/ALERTS/%$YEAR%-%$MONTH%.log")
    }
}

Three deliberate choices there: the rate limit is disabled for the device ruleset (a misconfigured device should not silence itself), files are split per device and per day, and the whole ruleset is scoped so that other inputs on the same server - application logs, for example - keep their own policy.

Retention and Rotation

Logrotate alone will not save you from 'the evidence aged out', and neither will a 30-day retention if your audit needs 12 months. Decide the requirement first, then implement:

  • Keep hot data (search) 30-90 days on the receiver.
  • Ship the same stream to long-term storage (object storage, a SIEM, or a second receiver) from day one - retrofitting is always painful.
  • Rotate daily with compression, and verify that a rotated file is still readable before you rely on it.
  • Monitor free disk space on the receiver itself. A full log server silently stops accepting logs, and the only symptom is missing data.

Time Is the Whole Ball Game

Syslog without synchronised clocks is a pile of text. On every device, configure NTP to the same reliable source as the log server, and check the delta - a two-minute offset between two switches makes a correlation timeline wrong, and a wrong timeline produces wrong conclusions.

# Cisco IOS
! Cisco IOS
ntp server 10.10.10.5 prefer
ntp server 10.10.10.6
clock timezone UTC 0
service timestamps log datetime msec localtime show-timezone year

# Junos
set system ntp server 10.10.10.5
set system time-zone UTC
set system syslog time-format year millisecond

Parsing and Alerting Without a SIEM

You do not need a full pipeline to get value. Two cheap additions to a plain rsyslog receiver deliver most of it: a filter that duplicates link-state and routing-adjacency messages into a separate alert file that your monitoring tool tails, and a daily summary (count by device, count by severity) that shows which devices are noisy. A device that suddenly emits ten times its normal log volume is usually the beginning of an incident, not the aftermath.

Related reading: ingesting network syslog into Splunk, the Filebeat/Logstash pipeline, and Vector as an alternative shipper when you outgrow a single receiver.

原文链接:https://www.rsyslog.com/doc/