ELK Stack: Centralised Logging for Network Devices - 夜莺博客

ELK Stack: Centralised Logging for Network Devices

Every network device can send syslog, and almost none of them can search it. The ELK stack — Elasticsearch for storage and search, Logstash or Filebeat for ingestion, Kibana for visualisation — turns thousands of unstructured syslog lines per second into something you can query in seconds. For network operations the payoff is concrete: find every interface flap in the last hour across forty switches, alert on BGP neighbour state changes, and keep twelve months of history without a single device having enough flash memory to do it locally. This guide covers the pipeline, the parsing that makes device logs useful, index lifecycle management, and the sizing arithmetic most deployments get wrong.

Pipeline architecture

devices (syslog UDP/TCP 514, TLS 6514)
      |
      v
rsyslog / syslog-ng  <-- buffer, filter, tag by device role
      |
      v
Filebeat (on the collector) --> Logstash (parse, enrich, route)
      |                                  |
      |                                  v
      |                          Elasticsearch (hot/warm/cold + ILM)
      |                                  |
      +--------------------------> Kibana (dashboards, alerting)

Do not point a switch directly at Logstash. Syslog over UDP has no backpressure: a Logstash restart drops everything in flight. Put a durable relay (rsyslog with a disk-assisted queue) in front, and you get retries, per-device tagging and a place to drop unwanted noise before it reaches the cluster.

# /etc/rsyslog.d/10-network.conf
module(load="imudp")
input(type="imudp" port="514" ruleset="network")

template(name="perdevice" type="string"
  string="/var/log/network/%FROMHOST-IP%/%$YEAR%-%$MONTH%-%$DAY%.log")

ruleset(name="network") {
  action(type="omfile" dynaFile="perdevice" template="RSYSLOG_FileFormat")
  action(type="omfwd" target="logstash.example.com" port="5044"
         protocol="tcp" queue.type="disk" queue.size="100000"
         action.resumeRetryCount="-1")
}

Parsing device logs into structured fields

Raw syslog gives you a timestamp and a message blob. The value comes from extracting fields: device, facility, severity, interface, neighbour, and the event that actually matters.

# /etc/logstash/conf.d/network.conf
input {
  beats { port => 5044 }
}
filter {
  # Cisco IOS example: %LINK-3-UPDOWN: Interface GigabitEthernet0/1, changed state to down
  grok {
    match => { "message" => [
      "%{SYSLOG5424PRI}%{NUMBER:seq}: \*?%{WORD:month} %{NUMBER:day} %{TIME:time}: ",
      "%%{DATA:facility}-%{NUMBER:severity}-%{WORD:mnemonic}: %{GREEDYDATA:detail}"
    ] }
  }
  grok {
    match => { "detail" => [
      "Interface %{DATA:interface}, changed state to %{WORD:link_state}",
      "BGP neighbor %{IP:bgp_neighbor} (AS %{NUMBER:bgp_as}) is %{WORD:bgp_state}",
      "%{WORD:iface} (?i)flap"
    ] }
  }
  if [mnemonic] == "UPDOWN" { mutate { add_tag => ["interface-event"] } }
  if [bgp_state] { mutate { add_tag => ["bgp-event"] } }
  date { match => ["[event][original]", "MMM  d HH:mm:ss"] target => "@timestamp" }
  mutate { remove_field => ["message"] }
}
output {
  if "interface-event" in [tags] or "bgp-event" in [tags] {
    elasticsearch { hosts => ["https://es-hot.example.com:9200"]
                    index => "network-events-%{+YYYY.MM}" }
  } else {
    elasticsearch { hosts => ["https://es-hot.example.com:9200"]
                    index => "network-logs-%{+YYYY.MM}" }
  }
}

Route high-value events to their own index. A separate network-events-* index with a longer retention and a smaller shard count is far cheaper to search during an incident than scanning a monolithic log index. Test grok patterns before deploying them — a pattern that fails silently leaves you with unparsed data and no field to alert on.

# validate a grok pattern against a sample line
bin/logstash -e 'filter { grok { match => { "message" => "%{SYSLOG5424PRI}%{NUMBER:seq}" } } }' \
  --config.test_and_exit --pipeline.batch.size 10

Index lifecycle management

{"policy": {"phases": {
  "hot":  {"actions": {"rollover": {"max_primary_shard_size": "30gb", "max_age": "1d"},
                       "set_priority": {"priority": 100}}},
  "warm": {"min_age": "3d", "actions": {"forcemerge": {"max_num_segments": 1},
                                        "set_priority": {"priority": 50}}},
  "cold": {"min_age": "30d", "actions": {"searchable_snapshot": {"snapshot_repository": "s3-repo"}}},
  "delete": {"min_age": "365d", "actions": {"delete": {}}}
}}}

Attach the policy as the index template default so every rolled-over index inherits it. The practical target: hot indices sized so a search across the last 24 hours touches few shards, warm indices force-merged for cheap scanning, and anything older than a month either on cheaper hardware or as a searchable snapshot. Without ILM, an ELK cluster fills up in weeks and then stops accepting writes — which is worse than no logging at all, because you found out during an incident.

Sizing arithmetic

  • Estimate log volume per device: a busy distribution switch with 48 ports can generate 5-20 MB per day, a core router with full BGP feeds far more.
  • Multiply by 30-50 percent for index and field overhead, then add 20 percent headroom. Keep total shard size 30-50 GB and aim for 20-40 tiers of shards per GB of heap.
  • Heap: half the host RAM up to 31 GB per Elasticsearch node, with the rest left for the filesystem cache.
  • Count the collector: a single rsyslog host can comfortably handle a few thousand messages per second with a disk queue; beyond that, shard collectors per site and send both to the cluster.
# cluster health and hot spots
curl -sk https://es-hot.example.com:9200/_cluster/health?pretty
curl -sk https://es-hot.example.com:9200/_cat/indices/network-*?v&s=index
curl -sk https://es-hot.example.com:9200/_cat/shards/network-logs-* | sort -k5 -r | head
curl -sk https://es-hot.example.com:9200/_cat/thread_pool/write?v

Alerting and dashboards that get used

  • Interface flaps: count of link_state:down events per interface per 15 minutes, alerting above a threshold you have tuned for access versus uplink ports.
  • BGP state changes: any bgp-event on a core peer, immediate alert.
  • Device silence: no events from a device for longer than its normal heartbeat — the failure you care about most, because a device that stops logging is either down or unreachable.
  • Authentication failures: burst of failed logins on a management interface.

Build one dashboard per role (core, access, firewall) with the same panel layout so an on-call engineer can move between them without re-learning the view. Keep queries in saved objects and reference them from alerts rather than duplicating logic.

Comparing alternatives

If your team already runs a Grafana-based stack, Loki stores logs with a label-only index and is far cheaper at the volumes syslog produces — the Loki aggregation guide covers the query model. A lightweight Fluent Bit to Elasticsearch pipeline is a good middle ground when you want structure without running Logstash, as shown in the Fluent Bit pipeline walkthrough. And if the primary goal is appliance-style syslog receipt with alerting rather than long-term analytics, Graylog as a central syslog server gets there with less operational effort.

Operational checklist

  • TLS with client certificates for syslog (port 6514) wherever devices support it — management credentials end up in log lines more often than anyone expects.
  • Filter at the collector: drop debug-level noise before it consumes cluster capacity.
  • Back up the snapshot repository and test a restore; Kibana saved objects live in a system index you must include.
  • Reindex or roll over before upgrading Elasticsearch, and check the deprecation API first.
  • Monitor the pipeline itself: if the number of indexed documents per hour drops sharply, you have a collector problem, not a quiet network.

原文链接:https://www.elastic.co/guide/en/elastic-stack/current/index.html