Grafana Loki LogQL: Selectors, Filters and Parsers - 夜莺博客

Grafana Loki LogQL: Selectors, Filters and Parsers

Loki's query language is deliberately two-part: a stream selector narrows down which log streams to read, then a pipeline parses and filters within them. Loki indexes only labels and timestamps, never the log body, which is why the selector matters so much for performance — an inefficient selector makes Loki read vastly more data than the query needs. This article covers LogQL from the mandatory selector outward: filters, parsers, label extraction and the metric queries that turn logs into graphs.

The Shape of a LogQL Query

{ log stream selector } | log pipeline

The selector is mandatory; the pipeline is optional. Loki indexes the timestamp and the labels, not the contents of the log line.

Stream Selectors

{service_name="nginx", status="500"}

The unique combination of all label pairs defines a stream. Four operators are available:

  • = exactly equal
  • != not equal
  • =~ regex match
  • !~ regex does not match
{name =~ "mysql.+"}
{name !~ `mysql-\d+`}
{app="mysql", name="mysql-backup"}

Note that =~ and !~ are fully anchored, unlike the line filter regex which is a substring search. An empty selector {} is a mistake: it forces Loki to scan everything.

Line Filters

Line filters are substring or regex tests over the raw log line, and they are much cheaper than parsing because they do not require decoding.

{job="mysql"} |= "error"                 # contains
{job="mysql"} != "timeout"               # does not contain
{name="cassandra"} |~  `error=\w+`       # regex match
{app="api"} !~ "healthcheck"             # regex does not match

Filters chain and are applied in order; a line must satisfy every filter to be returned. Putting the most selective filter first lets Loki drop lines earlier, which measurably reduces query time on large streams.

Parsers

Parsers extract labels from the log line so you can filter on values instead of substrings. Loki supports json, logfmt, pattern, regexp and unpack.

{container="query-frontend",namespace="loki-dev"} |= "metrics.go"
  | logfmt
  | duration > 10s and throughput_mb < 500

That single query shows the full pattern: select the container and namespace, keep only lines containing metrics.go, parse them as logfmt, then filter on the parsed numeric fields. Filtering numerically rather than with a substring is the difference between a query that returns the right ten lines and one that returns four thousand.

The JSON parser works the same way for structured logs:

{job="api"} | json | line_format "➡️ {{.request_method}} {{.request_uri}} status {{.status}}"

line_format rewrites the displayed line; label_format rewrites a label. Neither changes the stored data, only the query result, so they are safe to use freely for readability.

Label Filters Versus Line Filters

After parsing, the extracted fields can be filtered with comparison operators:

| json | status >= 500
| logfmt | duration > 2s and method == "POST"
| json | status != 200 or retries > 3

Label filters are evaluated after parsing and are far more precise than a regex over the line. When you find yourself writing an elaborate regex line filter, the better answer is almost always a parser plus a label filter.

Metric Queries

LogQL also produces metrics from log streams, which is how you build rate and error-ratio panels without a separate metrics pipeline:

rate({job="nginx"} |= "500" [5m])
sum by (status) (count_over_time({app="api"} | json [1m]))
quantile_over_time(0.99, {app="api"} | json | unwrap duration_ms [5m]) by (endpoint)

unwrap takes an extracted numeric field and treats it as the sample value, which is what makes percentiles over request durations possible directly from logs.

Performance Rules of Thumb

  • Always provide a selector with at least one equality match; regex-only selectors force full stream enumeration.
  • Keep label cardinality low. A label whose value is unique per request (a trace ID, a user ID) will destroy Loki's index — use a line filter or a parser field instead.
  • Prefer |= over a regex when a substring will do; it is substantially cheaper.
  • Put the narrowest filter first in the pipeline.

Related reading on this site: Grafana Loki: Log Aggregation Without the Index Bill deployment, Graylog as a Central Syslog Server for Network Devices as the alternative centralized syslog approach, and VictoriaLogs: Lightweight Log Storage and LogsQL Queries if you want a Loki-compatible store with a different index model.

原文链接:https://grafana.com/docs/loki/latest/logql