Vector.dev Log Pipelines: Sources, VRL Transforms, Sinks - 夜莺博客

Vector.dev Log Pipelines: Sources, VRL Transforms, Sinks

Vector is a single Rust binary that collects, transforms and ships logs and metrics, and it replaces a stack of agents you would otherwise run separately: a file shipper on every host, a parser, a router, and per-destination forwarders. The pipeline model is a directed graph of sources, transforms and sinks, configured in TOML, YAML or JSON. This guide builds a realistic pipeline: tail application logs, parse them, route by severity, send to two destinations, and archive to object storage with buffering that survives a downstream outage.

The Shape of a Vector Config

data_dir = "/var/lib/vector"

[api]
enabled = true
address = "127.0.0.1:8686"   # vector top

[sources.apache_logs]
type = "file"
include = ["/var/log/apache2/*.log"]
ignore_older_secs = 86400

[transforms.apache_parser]
inputs = ["apache_logs"]
type = "remap"
source = ". = parse_apache_log(.message)"

[sinks.es_cluster]
inputs = ["apache_sampler"]
type = "elasticsearch"
endpoints = ["http://79.12.221.222:9200"]
bulk.index = "vector-%Y-%m-%d"

Component IDs are the graph edges; inputs is what wires them. Wildcards work in sink inputs, so inputs = ["app*"] fans in every component whose ID starts with app.

Sources: Files, Syslog, Journald

[sources.syslog_udp]
type = "syslog"
mode = "udp"
address = "0.0.0.0:514"

[sources.journal]
type = "journald"
include_units = ["nginx.service", "postgresql.service"]

[sources.internal_metrics]
type = "internal_metrics"

Point network devices and Linux hosts at the syslog source and you centralise device logs without running rsyslog on every box. The journald source gives you structured fields instead of re-parsing text.

Transforms: VRL Is the Point

The remap transform runs Vector Remap Language — a small, safe expression language purpose-built for reshaping events. It cannot panic the process and it cannot do I/O, which is exactly why it is safe to hand to a pipeline.

[transforms.parse_syslog]
inputs = ["syslog_udp"]
type = "remap"
source = '''
  . = parse_syslog!(.message)
  .environment = "prod"
  .ingested_at = now()
'''

[transforms.route_by_severity]
inputs = ["parse_syslog"]
type = "route"
[transforms.route_by_severity.route]
urgent = '.severity == "err" || .severity == "emerg"'
normal = 'true'

Two rules of thumb save hours: keep transforms stateless where you can, and never put cache-like semantics in a transform because a transform does not exist to remember things across events.

Routing and Sampling to Control Cost

[transforms.sample_debug]
inputs = ["route_by_severity.normal"]
type = "sample"
rate = 10        # keep 1 in 10

[transforms.scrub]
inputs = ["sample_debug"]
type = "remap"
source = '''
  del(.headers.authorization)
  del(.password)
  .message = replace(.message, r'\d{16}', "[redacted]")
'''

[transforms.archive_only]
inputs = ["route_by_severity.urgent", "scrub"]
type = "filter"
condition = '.severity == "err"'

Sampling 1-in-10 on the chatty stream while keeping 100% of errors is the single biggest lever on observability spend. Scrub credentials in the pipeline, not in the destination — the sink is where data leaks from.

Sinks with Independent Buffers

[sinks.loki]
type = "loki"
inputs = ["scrub"]
endpoint = "http://loki.internal:3100"
labels.environment = "{{ environment }}"
[sinks.loki.buffer]
type = "disk"
max_size = 268435488          # ~256 MiB
when_full = "block"

[sinks.s3_archive]
type = "aws_s3"
inputs = ["archive_only"]
region = "us-east-1"
bucket = "log-archives"
key_prefix = "date=%Y-%m-%d"
compression = "gzip"
framing.method = "newline_delimited"
encoding.codec = "json"
batch.max_bytes = 10000000

A disk buffer with when_full = "block" means that when Loki is down, Vector stops accepting from that sink rather than dropping — and crucially, the S3 archive is unaffected because each sink buffers independently. That is the architectural reason to run one pipeline agent instead of three separate shippers.

Operate It

vector validate --config /etc/vector/vector.toml
vector --config /etc/vector/vector.toml
vector top                       # needs [api] enabled
curl -s localhost:8686/health
curl -s localhost:8686/metrics | grep -E "component_received|component_sent|buffer"

Alert on component_errors_total and on buffer size trending upward, not just on process liveness. A pipeline that is running but dropping events looks perfectly healthy to systemctl status.

Related Reading

Deeper dives on the same topics from our archive:

原文链接:https://vector.dev/docs/reference/configuration/