Thanos: Long-Term Prometheus Storage and Global Queries - 夜莺博客

Thanos: Long-Term Prometheus Storage and Global Queries

Prometheus is excellent at collecting and terrible at keeping data forever. Local blocks fill a disk, an HA pair doubles storage for identical series, and every cluster answers queries in isolation. Thanos fixes all three without replacing Prometheus: a sidecar ships completed blocks to object storage, a compactor reduces them, a store gateway serves historic data, and a querier presents every cluster behind one Prometheus-compatible endpoint. This guide covers the deployment topology, the configuration that must be right before you push a single block, and the verification steps that prove your two-year-old metrics are still queryable.

Component roles

  • Sidecar — runs next to each Prometheus, uploads finished blocks plus metadata to object storage, and serves recent data to the querier.
  • Store gateway — serves blocks from object storage, with an index cache that determines query performance.
  • Compactor — merges small blocks, applies retention, and enforces downsampling. Exactly one compactor per bucket; two will corrupt data.
  • Querier — fans a PromQL query out to all stores, deduplicates HA series, and returns a single result.
  • Receiver (optional) — accepts remote-write when Prometheus cannot be given local disk (for example, a multi-tenant SaaS deployment).
Prometheus + sidecar  ----> object storage (S3 / GCS / MinIO)
Prometheus + sidecar  ----> object storage
        |                          |
        +---> Querier <--- Store gateway
                 ^
                 +--- Grafana (one data source)

Bucket layout before you start

s3://metrics-thanos/
├── 01H8.../            # ULID per block, uploaded by the sidecar
│   ├── meta.json
│   ├── chunks/000001
│   └── index
├── thanos.shipper.json # uploader identity and labels
└── debug/metas/        # optional block metadata mirror

Decide the external labels now: cluster, region and replica are the ones that matter. Sidecars attach them to every block, and they are what makes deduplication and multi-cluster queries work later. Changing them after data is uploaded means old blocks can no longer be matched to new ones.

# prometheus.yml
global:
  external_labels:
    cluster: dc1
    region: cn-north
    replica: A

Sidecar and query configuration

# sidecar (one per Prometheus instance)
thanos sidecar \
  --prometheus.url=http://127.0.0.1:9090 \
  --tsdb.path=/var/lib/prometheus \
  --objstore.config-file=/etc/thanos/bucket.yml \
  --grpc-address=0.0.0.0:10901 --http-address=0.0.0.0:10902

# bucket.yml
type: S3
config:
  bucket: metrics-thanos
  endpoint: s3.example.com
  access_key: <access-key>
  secret_key: <secret-key>
  insecure: false

# querier with deduplication across the HA pair
thanos query \
  --http-address=0.0.0.0:19192 \
  --grpc-address=0.0.0.0:10904 \
  --query.replica-label=replica \
  --endpoint=dnssrv+_grpc._tcp.thanos-sidecar.dc1.svc \
  --endpoint=dnssrv+_grpc._tcp.thanos-store.dc1.svc \
  --store.sd-files=/etc/thanos/stores.yaml

--query.replica-label=replica is the deduplication switch. Without it, every HA pair returns two identical series and every graph shows doubled counters — the most common Thanos misconfiguration in production.

Retention, downsampling and compaction

thanos compact \
  --data-dir=/var/thanos/compact \
  --objstore.config-file=/etc/thanos/bucket.yml \
  --retention.resolution-raw=30d \
  --retention.resolution-5m=180d \
  --retention.resolution-1h=3y \
  --wait --wait-interval=5m

Raw resolution stays for a month, five-minute resolution for six, hourly for three years. Downsampled blocks are tiny compared with raw ones, which is what makes multi-year retention on object storage affordable. Give the compactor a persistent data directory with enough disk for the largest compaction group — it needs roughly 30 percent of the bucket size in local scratch space.

# compactor hygiene
curl -s http://compactor:10902/metrics | grep thanos_compact_group_compactions_total
curl -s http://compactor:10902/metrics | grep thanos_compact_blocks_marked_total
thanos tools bucket verify --objstore.config-file=/etc/thanos/bucket.yml --repair

Verification: is your history actually queryable?

# what does the querier see?
curl -s http://querier:19192/api/v1/stores | jq '.data[] | {name, lastCheck, minTime, maxTime}'

# do the blocks line up with the retention policy?
thanos tools bucket ls --objstore.config-file=/etc/thanos/bucket.yml | head -20
thanos tools bucket inspect --objstore.config-file=/etc/thanos/bucket.yml

# query a range that only exists in object storage
curl -s -G http://querier:19192/api/v1/query_range \
  --data-urlencode 'query=rate(node_network_receive_bytes_total{device="eth0"}[5m])' \
  --data-urlencode "start=$(date -d '45 days ago' +%s)" \
  --data-urlencode "end=$(date -d '44 days ago' +%s)" \
  --data-urlencode 'step=300' | jq '.data.result | length'

A non-zero result from that last query means the sidecar uploaded, the compactor downsampled, and the store gateway is serving history. If it returns empty while recent data works, the problem is in the object-store path, not Prometheus.

Common failure modes

  • Queries double every value: deduplication label missing or misnamed in one of the two Prometheus instances.
  • Sidecar upload errors (403/404): bucket policy blocks the debug/metas prefix the sidecar writes before blocks.
  • Store gateway slow queries: the index cache lives in memory on the querier instead of a shared cache such as memcached or Redis — size it against the total index size of recent blocks.
  • Compactor out of disk: retention is longer than local scratch can support for the biggest group. Split buckets by tenant or reduce compaction concurrency.
  • Gaps in long-range graphs: Prometheus was down longer than its local retention before the sidecar uploaded. Monitor prometheus_tsdb_head_series and sidecar upload lag; the alerting patterns for that live in the Alertmanager routing guide.

Scaling and integration notes

  • Use --query.partial-response so one slow store does not fail a dashboard refresh, and surface partial results in Grafana with a flag rather than hiding them.
  • For multi-cluster setups, run one querier per region and a global querier above them; do not let every Grafana instance connect to every store gateway.
  • Retarget your dashboards at the Thanos querier and retire per-Prometheus data sources so historical graphs do not break when a Prometheus instance is rebuilt.
  • Keep the recording rules that generate your expensive dashboards; Thanos makes storage cheap but it does not make a 500-million-sample query fast. The label hygiene that keeps queries efficient is covered in the relabel and target labels guide.
  • Object storage is now part of your monitoring availability — monitor bucket access latency and error rates as a first-class dependency.

原文链接:https://thanos.io/tip/thanos/getting-started.md/