Graylog as a Central Syslog Server for Network Devices - 夜莺博客

Graylog as a Central Syslog Server for Network Devices

Syslog is the only telemetry almost every network device supports, and it is the fastest
way to answer "what happened at 03:14" when a link flapped. The problem is scale: fifty devices
writing to fifty local buffers means fifty consoles. Graylog solves it with an input, a parsing
layer, a stream and an index — and the difference between a useful deployment and a noisy one
is entirely in those four pieces. This article covers the Graylog side and the device side, with
the configuration that keeps the signal and drops the rest.

Architecture and Ports

Graylog is three components: Graylog Server (processing and web), MongoDB (configuration
metadata), and OpenSearch or Elasticsearch (the log store). Plan the ports before you build:

  • 9000/tcp — web interface and REST API
  • 1514/tcp and 1514/udp — syslog inputs (use 1514 rather than 514 so you do not need to run
    as root; redirect 514 to 1514 at the firewall)
  • 5044/tcp — Beats input for Filebeat and Winlogbeat
  • 12201/tcp and udp — GELF, for structured application logs
# OpenSearch needs a larger virtual memory map, even in containers
sudo sysctl -w vm.max_map_count=262144
echo "vm.max_map_count=262144" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

sudo ufw allow 9000/tcp comment "Graylog Web"
sudo ufw allow 1514/tcp comment "Graylog Syslog TCP"
sudo ufw allow 1514/udp comment "Graylog Syslog UDP"
sudo ufw allow 5044/tcp  comment "Graylog Beats"
sudo ufw allow 12201/tcp comment "Graylog GELF"
sudo ufw reload

Step 1: Create the Input

Go to System > Inputs > Select input > Syslog UDP (or Syslog TCP) and launch
a new input. Bind it to 0.0.0.0 on port 1514, and set the Override source
option only if you are receiving from a relay that rewrites the source address.

System > Inputs > Launch new input
   Node:        Global
   Title:       Syslog-UDP-1514
   Bind address: 0.0.0.0
   Port:        1514
   Recv buffer size: 1048576

A green 1 RUNNING badge is the minimum evidence the listener is alive. Confirm with
real packets before going further — send one test message from the Graylog host itself:

logger -n 127.0.0.1 -P 1514 -d "graylog smoke test"
# or
nc -u -w1 127.0.0.1 1514 <<< "<14>Mar 12 10:00:00 testhost test: hello graylog"

Step 2: Index Set and Retention

Create one index set per class of device (core, access, firewall) rather than one global
index. Retention policy is where you decide cost: estimate daily volume, multiply by the
retention window, add roughly 30 percent for index overhead. A syslog-only deployment is far
cheaper than an infrastructure-as-code default that keeps everything forever.

System > Indices > Create index set
   Title:             network-devices
   Index prefix:      network_
   Shards:            1     (per Graylog node; keep it simple below 30 GB/day)
   Rotation strategy: Index time — P1D
   Retention:         Delete — max 30 indices

Step 3: Streams and Rules

A stream routes messages to an index set. The useful pattern is one stream per device class,
with a rule that matches on the source address or a custom field.

Streams > Create Stream
   Title:  Network Devices
   Index set: network-devices
   Rules:  Field: source  Type: match regular expression  Value: ^10\.10\.(1|2)\..*$
   Remove matches from Default Stream: yes

Where the device supports it, add a static field at the input — for example
log_type: cisco_ios or log_type: huawei_vrp — and match the stream on
that field. It is far more robust than regex against a hostname.

Step 4: Extractors That Actually Help

Raw syslog from a Cisco IOS device is a single message field. Without
extraction you cannot alert on it. Two low-effort wins:

System > Inputs > [your input] > Manage extractors > Create extractor
   Type:      regular expression
   Source field: message
   Condition:  Regular expression matches
   Expression: %([A-Z0-9_\-]{4,})-(\d)-([A-Z0-9_]+)
   Target field: severity_code

# Then a second extractor for the mnemonic, e.g. LINK-3-UPDOWN > LINK-3-UPDOWN
Extractors > Grok pattern (optional, for structured timestamps and interface names)

Having severity_code as a field is what turns "I see log lines" into "show me
every severity 2 message from the core switches this week".

Step 5: Configure the Devices

! Cisco IOS / IOS-XE
logging host 10.10.9.50 transport udp port 1514
logging trap informational
logging facility local6
logging source-interface Loopback0
logging buffered 64000 informational
service timestamps log datetime msec localtime show-timezone
! Arista EOS
logging host 10.10.9.50 1514
logging trap informational
logging source-interface Loopback0
logging format timestamp high-resolution
! Junos
set system syslog host 10.10.9.50 any info
set system syslog host 10.10.9.50 port 1514
set system syslog host 10.10.9.50 facility-override local6
set system syslog source-address 10.10.9.50
set system time-zone UTC
! Huawei VRP
info-center enable
info-center loghost 10.10.9.50 port 1514
info-center source default channel loghost log level informational
info-center timestamp loghost date

Two habits make the difference between usable and unusable logs: force every device to use
the same timezone (UTC if you operate across regions), and set a
source-interface so the Graylog source field is stable across interface changes.

Alerts Worth Creating on Day One

  • Any severity 2 or lower message from a core device
  • %LINK-3-UPDOWN or vendor equivalent exceeding a threshold per hour
  • Configuration change notifications where the platform supports them
  • Authentication failures and privilege escalation on management planes
  • Absence of expected keepalive syslog — a silent device is also a signal

Graylog pairs well with a proper timestamp discipline: if your devices' clocks drift, the log
pipeline lies to you. Chrony NTP server configuration on Linux covers the server side, and Cisco IOS logging levels, buffer and trap configuration explains the severity scale you will be filtering on.

For querying and long-term correlation at scale, Grafana Loki log aggregation with LogQL is the natural next step if you already run Grafana.

原文链接:https://graylog.org/post/how-to-use-graylog-as-a-syslog-server/