RabbitMQ Quorum Queues and Cluster Deployment - 夜莺博客

RabbitMQ Quorum Queues and Cluster Deployment

Quorum queues replaced mirrored classic queues as the recommended replicated queue type in RabbitMQ 3.8 and later: they are built on Raft, they survive the loss of a minority of nodes without split-brain, and their data-safety guarantees are documented rather than incidental. The trade-off is different capacity behaviour and a mandatory-quorum rule that will stop a queue rather than corrupt it. This guide covers a three-node cluster, quorum queue policy, and the failure tests that prove the guarantee.

Cluster prerequisites

  • Odd number of nodes (three is the practical minimum) with low, symmetric latency between them.
  • Fixed hostnames and an /etc/hosts (or DNS) entry for every node on every node — Erlang distribution fails cryptically with inconsistent name resolution.
  • Matched Erlang cookie on all nodes.
# on each node
hostnamectl set-hostname rabbit01.example.net
cat /var/lib/rabbitmq/.erlang.cookie   # must be identical everywhere, mode 400
systemctl enable --now rabbitmq-server

# join node 2 and 3 to node 1
rabbitmqctl stop_app
rabbitmqctl reset
rabbitmqctl join_cluster rabbit@rabbit01
rabbitmqctl start_app

rabbitmqctl cluster_status
rabbitmqctl list_nodes

Quorum queue policy and limits

rabbitmqctl set_policy ha-quorum "^(?!amq\.)" \
  '{"queue-master-locator":"client-local","delivery-limit":5}' \
  --apply-to queues

# node-level cluster size limits (RabbitMQ 3.8+)
rabbitmqctl set_cluster_quorum_queue_leader_locator client-local

# create a quorum queue explicitly - type cannot be changed in place later
rabbitmqadmin declare queue name=jobs.ingest durable=true arguments='{"x-queue-type":"quorum"}'

Key differences from classic mirrored queues:

  • Queues are declared as quorum at creation; existing classic queues cannot be converted in place.
  • Memory use per queue is higher because every replica holds the full Raft log.
  • A queue tolerates the loss of a minority. Lose the majority and the queue refuses to elect a leader — clients get NOT_ENOUGH_REPLICAS style errors until quorum returns. This is the intended behaviour: availability sacrificed to keep data.

Producer-side durability

import pika

params = pika.ConnectionParameters(
    hosts=['rabbit01', 'rabbit02', 'rabbit03'],   # failover list
    heartbeat=30, blocked_connection_timeout=120)

conn = pika.BlockingConnection(params)
ch = conn.channel()
ch.queue_declare(queue='jobs.ingest', durable=True,
                 arguments={'x-queue-type': 'quorum'})
ch.confirm_delivery()                      # publisher confirms
ch.basic_publish(exchange='', routing_key='jobs.ingest',
                 body=payload,
                 properties=pika.BasicProperties(delivery_mode=2))

Without publisher confirms, a publish that RabbitMQ never persisted can look successful. With quorum queues and confirms, a confirmed publish is written to a majority of replicas.

Failure testing

rabbitmq-diagnostics status | head -20
rabbitmqctl list_queues name type members online members
rabbitmq-queues quorum_status jobs.ingest
rabbitmq-queues check_if_node_is_quorum_critical

# simulate: stop one node, publish, restart, confirm no data loss
systemctl stop rabbitmq-server            # on rabbit02
rabbitmq-queues quorum_status jobs.ingest # followers should show 2 remaining
rabbitmq-queues check_if_node_is_quorum_critical

check_if_node_is_quorum_critical is the command that prevents the outage you cannot recover from: if taking a node out of maintenance would drop a queue below majority, it refuses.

Operational habits

  • Use separate vhosts per environment and per application; a single / vhost makes permissions impossible to reason about.
  • Set x-delivery-limit plus a dead-letter exchange so poisoned messages stop looping.
  • Alert on queue depth trend and on unacked-message count, not just depth — a consumer that stops acking looks healthy at the queue level.
  • Keep the management plugin behind authentication and TLS; it exposes message payloads.

Related reading: Kafka KRaft mode setup and operations, Redis Sentinel versus cluster failover design, and gNMI and Kafka streaming telemetry pipeline.

原文链接:https://www.rabbitmq.com/docs/quorum-queues