Envoy Proxy Configuration and Traffic Management - 夜莺博客

Envoy Proxy Configuration and Traffic Management

Envoy is the data plane behind most modern proxies — it is what Istio, Gloo and many service meshes run underneath. Learning it directly pays off because you get behaviour that nginx and HAProxy only approximate: first-class health checking, outlier detection, retries with budgets, and a routing configuration that can be changed at runtime without reloading. This article builds a working Envoy configuration from scratch and then walks through the traffic-management knobs that matter operationally.

The core model

Listener   an address/port Envoy accepts traffic on
Filter     HTTP connection manager, TCP proxy, TLS inspector
Route      matches on host/path/header -> picks a cluster
Cluster    a group of upstream endpoints + load balancer policy
Endpoint   an IP:port that actually serves the traffic

Configuration is static (YAML, loaded at start) or dynamic (via xDS APIs, changed at runtime). Static is fine for an edge proxy; anything that discovers endpoints automatically wants dynamic config.

A minimal but realistic edge config

static_resources:
  listeners:
  - name: http_edge
    address: { socket_address: { address: 0.0.0.0, port_value: 8080 } }
    filter_chains:
    - filters:
      - name: envoy.filters.network.http_connection_manager
        typed_config:
          "@type": type.googleapis.com/envoy.extensions.filters.network.http_connection_manager.v3.HttpConnectionManager
          stat_prefix: edge
          route_config:
            name: local_route
            virtual_hosts:
            - name: app
              domains: ["app.example.com"]
              routes:
              - match: { prefix: "/api" }
                route:
                  cluster: api_service
                  timeout: 5s
              - match: { prefix: "/" }
                route: { cluster: web_service }
          http_filters:
          - name: envoy.filters.http.router
            typed_config:
              "@type": type.googleapis.com/envoy.extensions.filters.http.router.v3.Router

  clusters:
  - name: api_service
    connect_timeout: 1s
    type: STRICT_DNS
    lb_policy: LEAST_REQUEST
    load_assignment:
      cluster_name: api_service
      endpoints:
      - lb_endpoints:
        - endpoint:
            address: { socket_address: { address: api.internal, port_value: 9000 } }
    health_checks:
    - timeout: 1s
      interval: 5s
      unhealthy_threshold: 3
      healthy_threshold: 2
      http_health_check: { path: /healthz }

The @type annotations are mandatory in the v3 API — omitting them is the most common first-run failure.

Load balancing and locality

ROUND_ROBIN          default, predictable
LEAST_REQUEST        favours the least-loaded backend (good for variable latency)
RANDOM               cheap, decent at scale
RING_HASH            consistent hashing, few remaps on endpoint change
MAGLEV               Google's algorithm; stable, no explicit hashing config
locality_weighted_lb_config: prefer endpoints in the same zone
                              with a spillover factor

For stateful APIs, ring hash or maglev plus locality weighting usually beats least-request; for stateless HTTP, least-request is the better default.

Retries, timeouts and budgets

route:
  cluster: api_service
  timeout: 5s
  retry_policy:
    retry_on: "5xx,reset,connect-failure,retriable-4xx"
    num_retries: 2
    per_try_timeout: 2s
    retry_host_predicate: [{ name: envoy.retry_host_predicates.previous_hosts }]
    retriable_status_codes: [503]

Unbounded retries are how a small upstream failure becomes a self-inflicted outage. Envoy supports retry budgets (retry_budget) that cap the percentage of requests that may be retried — enable them in any system with more than one layer of retries.

Circuit breaking and outlier ejection

circuit_breakers:
  thresholds:
  - priority: DEFAULT
    max_connections: 1000
    max_pending_requests: 1000
    max_requests: 2000
    max_retries: 3

outlier_detection:
  consecutive_5xx: 5
  interval: 10s
  base_ejection_time: 30s
  max_ejection_percent: 50
  enforcing_consecutive_gateway_failure: 100

Circuit breaking protects the backend; outlier ejection removes a bad instance from the pool. Used together they are the practical difference between "one bad pod degrades everything" and "one bad pod is removed in 30 seconds".

Observability

admin:
  address: { socket_address: { address: 127.0.0.1, port_value: 9901 } }
stats_config:
  stats_tags:
  - tag_name: cluster
    regex: "^cluster\.((.+?)\.)"
curl -s localhost:9901/stats | grep "cluster.api_service.upstream_rq_5xx"
curl -s localhost:9901/clusters | head -40
curl -s localhost:9901/config_dump | jq '.configs | length'

upstream_rq_5xx and upstream_rq_timeout per cluster are the two counters worth alerting on first.

Envoy vs nginx vs HAProxy

nginx       mature, simple, static config; reloads for changes
HAProxy     superb TCP/HTTP lb, excellent observability, static-ish config
Envoy       dynamic config (xDS), richest retry/outlier semantics, sidecar
            friendly, steeper learning curve

Practical guidance: keep a simple edge on nginx or Caddy (see nginx reverse proxy and Caddy), and reach for Envoy when you need dynamic endpoint discovery or per-route resilience policy. Envoy is also the reference data plane for Gateway API migrations.

FAQ

Q: Do I need Istio to use Envoy? No. Standalone Envoy with static config is a perfectly good proxy; the mesh adds control plane and mTLS automation.
Q: How do I reload config without dropping connections? Use the hot restart mechanism or, better, an xDS control plane; a plain restart drops connections.
Q: Where do timeouts live? Route timeout is the total; per_try_timeout bounds a single attempt; connect_timeout on the cluster bounds the TCP handshake. Getting all three consistent is most of the work.

原文链接:https://www.envoyproxy.io/docs/envoy/latest/