Filebeat and Logstash Pipeline for Device Logs - 夜莺博客

Filebeat and Logstash Pipeline for Device Logs

Filebeat is deliberately thin — it collects, optionally parses, and ships. Logstash is where the hard parsing lives. Putting the split in the wrong place produces an agent that consumes CPU on every server and a Logstash that drops events under load. This guide shows a pipeline layout that scales: Filebeat for transport and framing, Logstash for grok, field normalisation and routing, Elasticsearch for storage with lifecycle management.

Who does what

Stage Responsibility Avoid
Filebeat Read files, detect events, add host metadata, buffer on disk Complex grok; heavy JS processors
Logstash Parsing, enrichment, routing, dropping noise Very large per-event transforms that stall the pipeline
Elasticsearch Storage, retention, search Being the place where parsing "gets fixed" with ingest nodes for everything

Filebeat input and output

# /etc/filebeat/filebeat.yml
filebeat.inputs:
- type: filestream
  id: network-logs
  paths:
    - /var/log/network/*/*.log
  parsers:
    - ndjson:
        overwrite_keys: false
  file_identity.fingerprint: ~

processors:
  - add_host_metadata: ~
  - add_fields:
      target: source
      fields:
        collector: rsyslog-hk-01

output.logstash:
  hosts: ["logstash01.example.net:5044", "logstash02.example.net:5044"]
  loadbalance: true
  worker: 2
  bulk_max_size: 2048
queue.mem:
  events: 8192
  flush.min_events: 512
  flush.timeout: 5s

# Verify the shipper before adding parsing
filebeat test config && filebeat test output

The filestream input with a fingerprint identity keeps position across rotation without depending on inode stability, which matters when logs arrive over a network filesystem or a container path.

Logstash pipeline

# /etc/logstash/conf.d/network.conf
input {
  beats { port => 5044 }
}

filter {
  if [log][file][path] =~ /network/ {
    grok {
      match => { "message" => [
        "%{SYSLOGTIMESTAMP:event_time} %{HOSTNAME:device} %%{WORD:facility}-%{NUMBER:severity}-%{WORD:mnemonic}: %{GREEDYDATA:log_message}",
        "<%{NUMBER:pri}>%{TIMESTAMP_ISO8601:event_time} %{HOSTNAME:device} %{GREEDYDATA:log_message}"
      ] }
      tag_on_failure => ["_grokparsefailure_net"]
    }
    date {
      match => [ "event_time", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "ISO8601" ]
      target => "@timestamp"
    }
    translate {
      source => "severity"
      target => "severity_label"
      dictionary => {
        "0" => "emergency" "1" => "alert" "2" => "critical" "3" => "error"
        "4" => "warning"   "5" => "notice" "6" => "info"    "7" => "debug"
      }
    }
  }

  # drop the usual noise before it reaches the index
  if [log_message] =~ /(CONFIG_I|SPANTREE-7-RECV_1Q_NON_TRUNK)/ { drop {} }
}

output {
  elasticsearch {
    hosts => ["https://es01.example.net:9200"]
    index => "network-%{[device]}"
    user => "logstash_writer" password => "${ES_LS_PW}"
    manage_template => false
  }
}

Index per device keeps retention tunable per device class but multiplies shard counts; for anything beyond a few hundred devices, route to a small set of indices (network-core, network-access, network-fw) and keep the device name as a field instead.

Handling back-pressure honestly

# Logstash: persistent queue so a restart does not lose events
queue.type: persisted
queue.max_bytes: 8gb
pipeline.workers: 4
pipeline.batch.size: 250
pipeline.batch.delay: 50

# monitor the pipeline itself
curl -s localhost:9600/_node/stats/pipelines?pretty | jq '.pipelines.main.events, .pipelines.main.queue'

Watch filtered minus out; a growing gap means the Elasticsearch output is the bottleneck, not the filters. Raising workers will not help if the index is rejecting writes.

First-week checklist

  • Confirm the _grokparsefailure tag rate is near zero before declaring the pipeline done.
  • Check that @timestamp comes from the device, not from the collector's clock.
  • Set an ILM policy before the data volume makes a 30-day retention decision expensive.
  • Alert on the pipeline's event throughput dropping, which catches both collector and shipper failures.

Related reading: ELK stack centralised logging for network devices, Fluent Bit to Elasticsearch log pipeline, and Splunk syslog ingestion for network devices.

原文链接:https://www.elastic.co/guide/en/logstash/current/advanced-pipeline.html