Jaeger v2 Distributed Tracing Deployment Guide - 夜莺博客

Jaeger v2 Distributed Tracing Deployment Guide

Metrics tell you that latency tripled; traces tell you which hop did it. Jaeger v2 rebuilt itself on the OpenTelemetry Collector framework, which means one binary can be an all-in-one lab, a horizontally scaled collector, or a query-only front end — and all of them are configured with the same YAML pipeline model you already know from the Collector. This guide covers the deployment roles, a working configuration, sampling strategy and the storage decisions that determine whether traces are still available during an incident.

Roles in One Binary

Role Purpose Scaling
all-in-one Collector + query + UI in one process, in-memory or Badger storage Development, small labs
collector Receives OTLP/other protocols, processes and writes to storage Stateless — scale horizontally
query Serves the UI and trace query APIs Stateless — scale behind a load balancer
ingester Kafka-based ingestion path For very high ingest volume

Collection and query are stateless, which is what allows rolling upgrades with no trace loss. Upgrade order during a version change is from the end of the pipeline backwards: query first, then ingester, then collector — so that a newer component never receives data from an older one that cannot parse the response.

Quick start: all-in-one

docker run --rm --name jaeger   -p 16686:16686   -p 4317:4317   -p 4318:4318   -p 5778:5778   -p 9411:9411   jaegertracing/jaeger:2.x   --config /etc/jaeger/config.yaml

# UI:      http://localhost:16686
# OTLP gRPC: 4317      OTLP HTTP: 4318

Jaeger v2 requires an explicit config file (--config); there is no environment-variable-only mode as in v1, although the YAML can interpolate environment variables where you need it.

A Production Configuration Skeleton

service:
  extensions: [jaeger_storage, jaeger_query, healthcheckv2]
  pipelines:
    traces:
      receivers:  [otlp]
      processors: [batch, memory_limiter]
      exporters:  [jaeger_storage_exporter]

extensions:
  healthcheckv2:
    use_v2: true
    http: { endpoint: 0.0.0.0:13133 }
  jaeger_query:
    storage:
      traces: some_trace_storage
    http: { endpoint: 0.0.0.0:16686 }
    grpc: { endpoint: 0.0.0.0:16685 }
  jaeger_storage:
    backends:
      some_trace_storage:
        elasticsearch:
          server_urls: ["https://es.internal:9200"]

receivers:
  otlp:
    protocols:
      grpc: { endpoint: 0.0.0.0:4317 }
      http: { endpoint: 0.0.0.0:4318 }

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 75
    spike_limit_percentage: 20
  batch:
    timeout: 5s
    send_batch_size: 8192

exporters:
  jaeger_storage_exporter:
    trace_storage: some_trace_storage

Two processors are non-negotiable in production. memory_limiter prevents the collector from being OOM-killed by an ingest spike (the classic trace-pipeline outage). batch turns thousands of tiny writes into efficient storage calls. The jaeger_storage extension is what distinguishes Jaeger from a plain Collector: the Collector only writes, while Jaeger also reads traces for the UI, so storage configuration must be shared between components.

Sampling: The Decision That Determines Usefulness

  • Head sampling (decide at the start of a trace) is cheap and predictable but blind — it drops exactly the rare failing requests you want.
  • Tail sampling (decide after the trace completes) keeps all errors and all slow traces, at the cost of buffering the whole trace in the collector. Memory requirements grow with trace duration × throughput.
  • Adaptive sampling adjusts rates per operation as traffic changes, which is often the best compromise for high-volume services.

Start with head sampling at 5–10% plus always sample on errors, then introduce tail sampling for the two or three services where debugging cost is highest. Do not attempt to keep 100% of traces and then discover the storage cost in the first monthly bill.

Serving the UI Behind a Path Prefix

extensions:
  jaeger_query:
    base_path: /jaeger
    ui:
      config_file: /etc/jaeger/ui-config.json

base_path describes the path Jaeger sees after your reverse proxy has rewritten it. Getting this wrong produces a UI that loads but never renders data — a symptom that sends people looking at storage instead of at the proxy. Verify with a direct request to the collector before adding the proxy layer.

Verification

# Collector healthy?
curl -s localhost:13133/status
# Send a test span over OTLP HTTP
curl -X POST localhost:4318/v1/traces -H 'Content-Type: application/json' -d '{"resourceSpans":[]}'
# Trace visible in the UI?
open http://localhost:16686/search
# Storage growth sanity check (Elasticsearch backend)
curl -s 'localhost:9200/_cat/indices/jaeger-*?v&s=store.size:desc' | head

Operational Notes

  • Point applications at the collector, never straight at storage; that is the whole benefit of a stateless collector tier.
  • Set retention on the backing store (Elasticsearch ILM or equivalent) from day one — traces are high-cardinality and grow fast.
  • Alert on collector queue depth and dropped spans, not just on process health; a healthy collector that drops data is worse than a dead one.
  • Keep the OTLP endpoint internal; it is unauthenticated by design.

Related reading: Loki and LogQL, Alertmanager routing and Linkerd service mesh with mTLS.

原文链接:https://github.com/jaegertracing/documentation/blob/main/content/docs/v2/2.17/deployment/_index.md