Splunk Syslog Ingestion for Network Devices - 夜莺博客

Splunk Syslog Ingestion for Network Devices

Network devices generate a large share of the log volume in a typical enterprise, yet most Splunk deployments ingest them badly: everything lands in one index, timestamps are taken from the collector instead of the device, and severity is buried in the raw text. The result is a search that cannot correlate a BGP flap with the interface event that caused it. This guide covers the ingestion path, the parsing configuration that fixes timestamps and event boundaries, and the field extractions that make vendor logs searchable.

Ingestion architecture

Centralise syslog first, then forward to Splunk. Two patterns are common:

  • Syslog server plus universal forwarder — rsyslog receives on UDP/TCP 514, writes per-host files, and a Splunk universal forwarder monitors them. Cheapest and easiest to debug.
  • Direct to Splunk via syslog input — simplest to deploy, but you lose the intermediate buffer and per-host file organisation that make troubleshooting easy.
# rsyslog: one file per device
template(name="PerHost" type="string"
  string="/var/log/network/%FROMHOST-IP%/%$YEAR%-%$MONTH%-%$DAY%.log")
ruleset(name="network") {
  action(type="omfile" dynaFile="PerHost" template="RSYSLOG_FileFormat")
  stop
}
input(type="imudp" port="514" ruleset="network")
input(type="imtcp" port="514" ruleset="network")

Device-side configuration

! Cisco IOS / IOS-XE
logging host 10.30.0.20
logging trap informational
logging source-interface Loopback0
service timestamps log datetime msec localtime show-timezone year

# Junos
set system syslog host 10.30.0.20 any info
set system syslog host 10.30.0.20 facility-override local5
set system syslog host 10.30.0.20 explicit-priority
set system syslog source-address 10.30.0.1

# Huawei VRP
info-center enable
info-center loghost 10.30.0.20 channel 2
info-center timestamp log date precision-time millisecond
info-center source default channel 2 log level informational

Timestamp precision and timezone are not cosmetic: without milliseconds and an explicit timezone, Splunk cannot order events that occur in the same second, and correlation between device families breaks.

Universal forwarder configuration

# /opt/splunkforwarder/etc/system/local/inputs.conf
[monitor:///var/log/network]
disabled = false
index = network
sourcetype = cisco:ios
crcSalt = <SOURCE>

# /opt/splunkforwarder/etc/system/local/outputs.conf
[tcpout]
defaultGroup = indexers
[tcpout:indexers]
server = idx1.example.net:9997,idx2.example.net:9997

Parsing: props and transforms

# props.conf on the indexers
[cisco:ios]
SHOULD_LINEMERGE = false
LINE_BREAKER = ([\r\n]+)\d{4}-\d{2}-\d{2}
TIME_PREFIX = ^
TIME_FORMAT = %Y-%m-%dT%H:%M:%S.%3N%z
MAX_TIMESTAMP_LOOKAHEAD = 40
TRUNCATE = 10000
EXTRACT-facility_severity = ^\S+\s+%\w+-(?<severity>\d)-(?<facility>\w+)
TRANSFORMS-set_index = network_index
REPORT-cisco_mnemonic = cisco_mnemonic
# transforms.conf
[cisco_mnemonic]
REGEX = %(?:\w+)-(\d)-(\w+):\s+(?<mnemonic>[A-Z0-9_]+):\s+(?<message>.*)
[network_index]
REGEX = .
DEST_KEY = _MetaData:Index
FORMAT = network

SHOULD_LINEMERGE = false plus a LINE_BREAKER that matches the device timestamp keeps multi-line stack traces and config dumps from being split, while still separating events whose timestamp changes. Applied in the wrong order, one mis-parsed event masquerades as thousands.

Searches worth saving

index=network sourcetype=cisco:ios earliest=-1h
| stats count by host, severity, mnemonic
| sort - count

index=network ("%LINK-3-UPDOWN" OR "%LINEPROTO-5-UPDOWN" OR "SNMP_TRAP_LINK")
| transaction host maxspan=2m keepevicted=false
| where eventcount > 4
| table host, _time, eventcount

index=network ("BGP" AND ("reset" OR "down"))
| stats latest(_time) as last_down values(mnemonic) as mnemonics by host
| sort - last_down

Practical guardrails

  • Separate indices for network versus application logs so retention and licensing are controlled independently.
  • Drop or summarise chatty debug facilities at the collector — a single flooding device can consume a licence quota overnight.
  • Monitor forwarder connectivity separately from device reachability, otherwise a silent ingestion gap looks like a quiet network.

Related reading: ELK stack centralised logging for network devices, Fluent Bit to Elasticsearch log pipeline, and rsyslog central log server and forwarding.

原文链接:https://docs.splunk.com/Documentation/Splunk/latest/Data/Usesyslog