VictoriaLogs: Lightweight Log Storage and LogsQL Queries - 夜莺博客

VictoriaLogs: Lightweight Log Storage and LogsQL Queries

VictoriaLogs is a single-binary log database from the VictoriaMetrics project, and it attacks the usual problem with log platforms — resource consumption — head on. It is schemaless, indexes every field it ingests, stores compressed data on local disk with automatic retention, and answers queries in LogsQL from a built-in UI. This guide covers deployment, ingestion, retention configuration and the query patterns that replace the ad-hoc grep habit.

Deploy: one binary, one port

curl -L https://github.com/VictoriaMetrics/VictoriaLogs/releases/latest/download/victoria-logs-linux-amd64-v1.26.0.tar.gz | tar xz
./victoria-logs -storageDataPath=/var/lib/victoria-logs -retentionPeriod=30d

The service listens on port 9428 by default and serves both the HTTP API and the web UI at /select/vmui/. No schema planning is required because VictoriaLogs indexes all fields found in the logs, which means a log line with new fields is immediately queryable.

# systemd unit
[Unit]
Description=VictoriaLogs
After=network-online.target

[Service]
ExecStart=/opt/victorialogs/victoria-logs -storageDataPath=/var/lib/victoria-logs -retentionPeriod=30d
Restart=always
RestartSec=5

[Install]
WantedBy=multi-user.target

Ingestion: accept logs from anything

VictoriaLogs speaks the formats you already have: Elasticsearch Bulk API, Loki push API, OpenTelemetry, Datadog, JSON Lines, Journald and Syslog. That flexibility is the point — you do not need to replace existing shippers to adopt the storage layer.

# Loki-compatible push
curl -X POST http://localhost:9428/insert/loki/api/v1/push \
  -H 'Content-Type: application/json' \
  -d '{"streams":[{"stream":{"job":"api"},"values":[["1700000000000000000","login failed for user admin"]]}]}'

# JSON lines
curl -X POST http://localhost:9428/insert/jsonline \
  -H 'Content-Type: application/stream+json' \
  -d '{"log":{"level":"error","message":"database timeout"},"_time":"2026-09-28T09:00:00Z","_stream":"app"}'

Behind a load balancer, /internal/force_flush converts in-memory buffers into searchable blocks so a query immediately after ingestion sees the newest logs. Use it for tests and automation, not on a schedule — it costs CPU and slows ingestion.

Retention: two independent mechanisms

./victoria-logs -retentionPeriod=8w
./victoria-logs -retention.maxDiskSpaceUsageBytes=100GiB
./victoria-logs -retention.maxDiskUsagePercent=80

Data is stored in per-day partition directories identified by UTC calendar day, and retention drops whole partitions rather than individual entries — fast and predictable. The default retention is 7 days, with 1 day as the minimum. The two disk-based limits are mutually exclusive: set either an absolute byte cap or a disk percentage, never both. Combining time-based and disk-based retention is supported and sensible: time removes old partitions, disk space removes the oldest partitions early when the volume approaches its limit.

One consequence of UTC partitioning is worth planning around: if you are not in UTC, a local calendar day spans two partitions, so a query scoped to "yesterday local time" is still correct but touches two directories.

LogsQL: filters plus pipes

Every LogsQL query has a mandatory filter and optional pipes.

# full-text search
_msg:error

# field match with pipes
level:error _time:1h | stats by (host) count() as errors | sort by (errors) desc

# count requests per path
_msg:"GET /api" | stats by (path) count() as hits | sort by (hits) desc | limit 20

Because every ingested field is indexed automatically, the practical workflow is: ingest, then explore in the UI at http://<host>:9428/select/vmui/, then save the queries that proved useful as panels.

Integrations and HA

  • Grafana connects through the LogsQL data source for dashboards and exploration.
  • vmalert plus Alertmanager turns LogsQL queries into alerts, which is how you get "notify me when this log line appears" without writing a custom script.
  • For availability, send the same logs to multiple single-node instances in different failure domains using a shipper that supports multiple destinations, such as vlagent, and put vmauth in front of queries.

Resource-wise, the recommended filesystem is ext4, and for partitions above 1 TB the documentation suggests formatting with the large-file options rather than defaults. If you are already running an appliance-style syslog server, compare the operational cost with Graylog central syslog server before migrating, and if you live in Grafana already, the pipeline in Grafana Loki log aggregation and LogQL shows how Loki handles the same job with different trade-offs.

原文链接:https://victoriametrics.com/blog/victorialogs-architecture-basics/