RabbitMQ Overview¶
RabbitMQ is an open-source message broker written in Erlang. Applications hand it messages, it routes them into queues according to rules you define, and it delivers them to consumers with per-message acknowledgements, so work survives crashes, restarts and slow consumers. It speaks several protocols (AMQP 0-9-1, AMQP 1.0, MQTT, STOMP and its own stream protocol) and runs as a single node or as a cluster with replicated queues.
This page explains how it works and how to run it well. The Operations page is the command-level companion and the Lab is where to try all of it.
What changed since 3.x¶
Most RabbitMQ articles, Stack Overflow answers and generated tutorials still describe 3.x. On 4.3 several of those instructions silently do nothing, and a few now fail outright:
| Old advice | Status in 4.3 | Use instead |
|---|---|---|
Classic mirrored queues (ha-mode, ha-params, ha-sync-mode policies) |
Removed in 4.0; the policy keys have no effect | Quorum queues (or streams) |
Lazy queues (x-queue-mode: lazy) |
Declaring with x-queue-mode fails (CQv1 storage removed in 4.3) |
Nothing; classic queues already page to disk aggressively |
cluster_partition_handling = pause_minority / autoheal |
Accepted but ignored; Raft handles partitions | Odd-sized clusters, quorum queues |
| Mnesia as the metadata store | Removed; Khepri (Raft-based) is the only store | Plan for majority availability |
| Non-durable, non-exclusive ("transient") queues | Rejected by default | Durable queues, exclusive queues, or queue TTL |
vm_memory_high_watermark_paging_ratio |
No longer meaningful | Size memory by the watermark alone |
Global QoS (basic.qos with global=true) |
Denied by default | Per-consumer prefetch |
queue-master-locator |
Denied by default | queue-leader-locator |
rabbitmq-delayed-message-exchange plugin |
Deprecated and archived | TTL tiers, or quorum queue delayed retry for retries |
rabbitmqadmin v1 (Python, downloaded from the node) |
Download endpoint removed | rabbitmqadmin v2 (standalone binary) |
| Consumer timeout applies to every queue | Only quorum queues enforce it | Same setting, now graceful (cancels one consumer, not the channel) |
Upgrades to 4.3 must come from the latest 4.2.x patch, with all stable feature flags enabled first. See upgrades.
Where RabbitMQ fits¶
RabbitMQ is a smart broker: the server decides where each message goes, tracks every delivery and acknowledgement, retries, expires and dead-letters. Clients stay simple. That makes it a good fit for:
- Work distribution: background jobs, task queues, fan-out to worker pools, with fair dispatch, retries and backpressure.
- Service decoupling and integration: events between services, routed by type or content, where each consumer gets its own queue and can be offline for a while.
- Request/reply between services that shouldn't call each other directly.
- IoT and edge ingestion via MQTT, bridged into the same routing model.
- Buffering pipelines, such as log shippers (Vector, Logstash, Fluent Bit all have AMQP inputs and outputs) that need a durable buffer between producers and an indexer.
It's a weaker fit when you need a long, replayable event history consumed by many independent readers at very high throughput. That's Kafka's home turf, although RabbitMQ streams now cover a good part of it.
| RabbitMQ | Apache Kafka | NATS (JetStream) | ActiveMQ Artemis | ZeroMQ | |
|---|---|---|---|---|---|
| Model | Broker with exchanges, queues and streams | Partitioned, replicated log | Subjects; JetStream adds persistence | Broker (JMS-centric) | Library; no broker at all |
| Routing | Rich, broker-side (topic, headers, fanout, ...) | By topic and partition key only | Subject wildcards | Addresses, JMS selectors | You build it |
| Consumption | Per-message ack, redelivery, DLX | Consumer tracks offsets | Ack-based or replay | Per-message ack | N/A |
| Replay | Streams only | Core feature | JetStream | Limited | No |
| Ordering | Per queue | Per partition | Per stream | Per queue | Per socket |
| Typical sweet spot | Task queues, service events, RPC, MQTT | Event sourcing, analytics, CDC | Low-latency service mesh messaging | Java/JMS estates | Embedded, in-process messaging |
Managed equivalents follow the same split: Azure Service Bus and Amazon SQS/SNS are queue-and-topic brokers in RabbitMQ's family; Azure Event Hubs and Amazon Kinesis/MSK sit in Kafka's. Amazon MQ offers managed RabbitMQ itself.
Protocols and ports¶
AMQP 0-9-1 is RabbitMQ's original protocol, and its model (exchanges, queues, bindings) is the broker's model. AMQP 1.0 has been a native core protocol since 4.0 and is where most new client features land; RabbitMQ's newer official clients use it. MQTT (3.1, 3.1.1, 5.0) and STOMP are plugins that map onto the same exchanges and queues, so an MQTT device can publish and an AMQP service can consume. Streams have their own binary protocol for high-throughput consumption.
| Port | Purpose | Expose to |
|---|---|---|
| 5672 / 5671 | AMQP 0-9-1 and 1.0, plain / TLS | Clients |
| 5552 / 5551 | Stream protocol, plain / TLS | Stream clients |
| 1883 / 8883 | MQTT, plain / TLS | Devices |
| 61613 / 61614 | STOMP, plain / TLS | Clients |
| 15672 / 15671 | Management UI and HTTP API, plain / TLS | Operators only |
| 15692 | Prometheus metrics | Monitoring only |
| 15674 / 15675 | Web STOMP / Web MQTT (WebSockets) | Browsers |
| 4369 | epmd, the Erlang port mapper |
Cluster nodes and CLI hosts only |
| 25672 | Erlang distribution (inter-node and CLI traffic) | Cluster nodes and CLI hosts only |
| 35672-35682 | CLI tool connections | Cluster nodes only |
Ports 4369 and 25672 carry full control over the node to anyone holding the Erlang cookie. Never expose them beyond the cluster network.
The core model¶
Producer RabbitMQ (virtual host "/") Consumer
┌────────┐ publish(exchange, key) ┌──────────────┐ binding: "app.#" ┌─────────────────┐ deliver ┌────────┐
│ app │ ───────────────────────▶ │ lab.logs │ ─────────────────▶ │ lab.app-logs │ ────────▶ │ worker │
└────────┘ key="app.error" │ (topic) │ │ (quorum queue) │ ◀──────── └────────┘
└──────┬───────┘ └─────────────────┘ ack
│ binding: "*.error" ┌─────────────────┐
└──────────────────────────▶ │ errors │ ───▶ ...
└─────────────────┘
Connections and channels. A client opens one long-lived TCP connection (optionally TLS) and multiplexes lightweight channels over it. Channels are where publishing, consuming and declaring happen. The usual layout is one connection per process and one channel per thread; opening a connection per message is the classic anti-pattern, because connection setup costs several network round trips and broker-side resources.
Virtual hosts partition one broker into isolated namespaces. Exchanges, queues, bindings, policies and permissions all belong to a vhost; a user's permissions are granted per vhost. Use them to separate applications or environments sharing a cluster.
Exchanges receive published messages and route them. Producers never publish to a queue directly; even "publishing to a queue" goes through the default exchange, a direct exchange that every queue is implicitly bound to under its own name.
| Type | Routes by | Typical use |
|---|---|---|
direct |
Exact match of routing key to binding key | Task queues, routing by type |
topic |
Dot-separated pattern match: * = exactly one word, # = zero or more |
Events by category (order.created, app.web.error) |
fanout |
Ignores the key; copies to every bound queue | Broadcast, cache invalidation |
headers |
Message header values (x-match: all or any) |
Routing on several attributes |
x-consistent-hash (plugin) |
Hash of routing key or header, weighted by binding | Sharding work across queues while keeping per-key ordering |
x-modulus-hash (core as of 4.3) |
Hash modulo the number of bindings | Even distribution across a fixed set of queues |
Bindings are the routing rules linking an exchange to a queue (or to another exchange,
for exchange-to-exchange topologies). A message that matches no binding is unroutable: by
default it's silently dropped. Publish with the mandatory flag to have it returned to the
publisher instead, or give the exchange an alternate exchange that catches everything
unroutable.
Queues hold messages until a consumer acknowledges them. A queue is FIFO, with exceptions
for priorities, redeliveries and delayed retries. Queues are declared by clients or
operators; declaration is idempotent as long as the properties match, and redeclaring with
different arguments fails with 406 PRECONDITION_FAILED and closes the channel. That's why
the lab declares each topology in one shared function.
Messages are an opaque body plus properties the broker understands:
| Property | Meaning |
|---|---|
delivery_mode |
2 = persistent (written to disk); 1 = transient |
content_type, content_encoding |
For consumers; RabbitMQ doesn't inspect the body |
message_id, correlation_id |
Application identifiers; correlation_id matches replies to requests |
reply_to |
Where a reply should go (RPC) |
expiration |
Per-message TTL in milliseconds, as a string |
priority |
Used by priority-capable queues |
timestamp, type, app_id, user_id |
Metadata; user_id is validated against the connection's user |
headers |
Arbitrary key/value table; also where the broker writes x-death, x-delivery-count, etc. |
The maximum message size is 16 MiB by default on 4.x. Large payloads belong in object storage with the message carrying a reference (the claim check pattern).
Queue types¶
RabbitMQ 4.3 has three queue types. Choosing between them is the most consequential design decision you'll make.
Classic queues¶
The original, non-replicated queue type. Each classic queue lives on one node and is served by a single Erlang process. They're fast, cheap and feature-complete for single-node use, and they're the right choice for data you can afford to lose with a node: exclusive reply queues, temporary per-client queues, caches of rebuildable work.
Since 4.3 only the second-generation storage (CQv2) exists. Small messages are embedded in
the queue's own index files and larger ones go to a node-wide shared message store. Classic
queues move messages to disk eagerly, so the old "lazy mode" is gone along with the storage
engine that needed it. Classic queues support priorities (x-max-priority at declaration).
They don't enforce consumer timeouts or delivery limits.
Quorum queues¶
The replicated, durable, data-safety-first queue type, built on the Raft consensus algorithm (via RabbitMQ's Ra library). A quorum queue has a leader and followers on different nodes; publishes and acks go through the leader and are committed once a majority of members has written them to disk. Publisher confirms are only sent after that commit, so a confirmed message survives the loss of any minority of nodes.
Quorum queues stay available as long as a majority of their members is up: a 3-member queue tolerates one node down, a 5-member queue two. If a majority is lost the queue is unavailable (not inconsistent) until enough members return.
Features specific to quorum queues:
- Poison message handling. Each message carries a
delivery-count; past the queue'sdelivery-limit(default 20 since 4.0) it's dead-lettered, or dropped if no DLX is set. - At-least-once dead-lettering (
x-dead-letter-strategy: at-least-once, requiresx-overflow: reject-publish). Dead-lettered messages aren't removed from the source until the target queues confirm them. - Priorities. 32 strict levels as of 4.3 (4.0-4.2 had two). No
x-max-priorityneeded. - Delayed retry (4.3). Returned messages are set aside for a linearly increasing delay
before redelivery:
min(delayed-retry-min × delivery-count, delayed-retry-max). - Consumer timeout (enforced by the queue as of 4.3). A consumer that holds a delivery
too long gets a
basic.cancelfor that one consumer; its messages are returned. - Single active consumer, consumer priorities, per-queue and per-message TTL, length limits.
Quorum queues are always durable and can't be exclusive. Every member stores every message, so a 3-member queue writes three copies; that's the cost of the safety. In 4.3 the per-message memory overhead was roughly halved for messages up to 32 KiB.
Streams¶
A stream is an append-only, replicated log. Consuming doesn't remove anything: each consumer
reads from an offset it chooses (first, last, next, a numeric offset or a timestamp),
and many consumers can read the same data independently and repeatedly. Retention is by size
or age (x-max-length-bytes, x-max-age), applied per segment file.
Streams suit large fan-outs (thousands of consumers reading the same messages), replay, time-travel debugging and very high throughput. Over AMQP 0-9-1 they behave like a queue with an offset argument; the dedicated stream protocol (port 5552) is much faster and adds publisher deduplication, server-side filtering and super streams (a partitioned stream, the closest thing RabbitMQ has to a Kafka topic).
Choosing¶
| Classic | Quorum | Stream | |
|---|---|---|---|
| Replicated | No | Yes (Raft) | Yes |
| Survives node loss | No | Yes, with a majority up | Yes, with a majority up |
| Consumption | Destructive | Destructive | Non-destructive, by offset |
| Exclusive / temporary | Yes | No | No |
| Poison message handling | No | Yes | N/A |
| Priorities | Yes (x-max-priority) |
Yes (32 levels) | No |
| Consumer timeout (4.3) | No | Yes | No |
| Best for | Temporary or reply queues, losable work | Anything you can't lose | Fan-out, replay, high-volume events |
Rule of thumb for 4.x: make durable application queues quorum queues, use classic queues for
exclusive or temporary queues, and reach for streams when consumers need history or there are
a great many of them. You can set a vhost's default_queue_type to quorum so that clients
which don't specify a type get quorum queues.
Queue arguments versus policies¶
Most queue behaviour can be set two ways: as x- arguments at declaration time, or as keys in
a policy applied by name pattern. Arguments are baked in: changing them means deleting and
redeclaring the queue. Policies can be changed at any time and take effect immediately, so
operational settings (length limits, TTLs, dead-lettering, delivery limits, delayed retry) are
usually better as policies, leaving arguments for things that define the queue's identity,
such as its type. When both are set the argument generally wins; operator policies let
administrators impose caps (like a maximum length) that application policies can't override.
| Behaviour | Queue argument | Policy key |
|---|---|---|
| Queue type | x-queue-type |
(argument only) |
| Max messages / bytes | x-max-length, x-max-length-bytes |
max-length, max-length-bytes |
| Overflow behaviour | x-overflow |
overflow |
| Message TTL | x-message-ttl |
message-ttl |
| Delete unused queue after | x-expires |
expires |
| Dead-letter exchange / key | x-dead-letter-exchange, x-dead-letter-routing-key |
dead-letter-exchange, dead-letter-routing-key |
| Dead-letter strategy (quorum) | x-dead-letter-strategy |
dead-letter-strategy |
| Delivery limit (quorum) | x-delivery-limit |
delivery-limit |
| Delayed retry (quorum, 4.3) | x-delayed-retry-type, -min, -max |
delayed-retry-type, -min, -max |
| Consumer timeout (quorum, 4.3) | x-consumer-timeout |
consumer-timeout |
| Initial replica count (quorum) | x-quorum-initial-group-size |
(argument only) |
| Leader placement | x-queue-leader-locator |
queue-leader-locator |
| Single active consumer | x-single-active-consumer |
(argument only) |
Delivery guarantees¶
RabbitMQ gives you at-most-once or at-least-once delivery, depending on how you configure both ends. There's no exactly-once: it would require the broker to know whether your consumer's side effects happened, which it can't. Reliable systems combine at-least-once delivery with idempotent consumers.
At-least-once is a chain, and every link has to hold:
| Link | What it protects against | How |
|---|---|---|
| Durable topology | Queues and exchanges disappearing on restart | durable=True (quorum queues and streams always are) |
| Persistent messages | Messages in a classic queue lost on restart | delivery_mode=2 |
| Publisher confirms | Publisher assuming the broker has a message it never stored | confirm.select, then wait for the ack |
| Mandatory flag / alternate exchange | Messages silently dropped as unroutable | mandatory=True or an alternate-exchange |
| Replication | Loss of the node holding the queue | Quorum queues or streams |
| Manual consumer acks | Consumer crashing mid-message | auto_ack=False, ack after the work is done |
| Idempotent processing | Duplicates caused by all of the above retrying | Dedup by message ID, upserts, unique constraints |
Publisher side¶
With confirms enabled on a channel, the broker acknowledges each publish once it has taken
responsibility for it: written to disk for persistent messages in classic queues, committed by
a majority for quorum queues and streams. A nack means the broker couldn't accept it (for
example the target queue is full and set to reject-publish). Anything neither acked nor
nacked when a connection drops is in an unknown state and must be republished, which is one of
the sources of duplicates.
Waiting for a confirm after every message serializes publishing on network round trips. High-throughput publishers keep a window of unconfirmed messages in flight (hundreds to thousands) and handle confirms asynchronously as they arrive. Lab 5.2 measures the difference.
Consumer side¶
Consumers subscribe with basic.consume and the broker pushes messages to them. Prefetch
(basic.qos) caps how many unacknowledged deliveries each consumer can have outstanding. The
default is unlimited, which lets one consumer hoard an entire backlog in memory, so always set
it. Low values (1-10) give fair dispatch for slow, uneven work; higher values (100-300) keep
fast consumers busy when the network round trip matters.
A delivery ends in one of three ways:
| Consumer action | AMQP 0-9-1 method | Effect |
|---|---|---|
| Processed successfully | basic.ack |
Message removed |
| Failed, try again | basic.reject / basic.nack with requeue=true |
Returned to the queue (or delayed, with delayed retry) |
| Failed permanently | basic.reject / basic.nack with requeue=false |
Dead-lettered if a DLX is set, otherwise discarded |
If the consumer's channel or connection closes, every unacknowledged delivery is requeued
and redelivered with the redelivered flag set.
On 4.3 quorum queues, basic.reject and basic.nack mean different things for poison-message
accounting. basic.reject (like a crash or lost connection) records a failed delivery and
advances delivery-count. basic.nack records a delivery the consumer simply didn't process,
which advances only the new acquired-count, so it never trips the delivery limit. A consumer
that nacks a bad message forever loops forever. Use reject for failures. (Up to 4.2 both
counted.)
Dead-lettering¶
A message is dead-lettered when it's rejected without requeue, its TTL expires, it's pushed out
by a length limit (with drop-head or reject-publish-dlx overflow), or it exceeds a quorum
queue's delivery limit. It's republished to the queue's dead-letter exchange, keeping its
original routing key unless dead-letter-routing-key overrides it, with an x-death header
recording the queue, reason, count and time of each death (plus x-first-death-reason and
friends). Classic queues dead-letter at-most-once: a crash at the wrong moment can lose the
message. Quorum queues can do it at-least-once.
Expiry and limits¶
Per-queue message TTL (message-ttl) expires messages as they reach the head of the queue. A
per-message TTL (expiration property) works the same way, which means a long-TTL message
blocks shorter-TTL messages behind it; that's why the delay pattern in Lab 3.2 uses one queue
per delay. Queue TTL (expires) deletes a queue nobody has used for the given time, which is
how you get self-cleaning temporary queues now that transient non-exclusive ones are disallowed.
Length limits (max-length, max-length-bytes) bound a queue. The overflow setting decides
what happens when it's full: drop-head (default; discard or dead-letter the oldest),
reject-publish (nack new publishes, so publishers with confirms feel backpressure), or
reject-publish-dlx (classic queues only; nack and dead-letter the rejected message). Every
production queue should have a limit: an unbounded queue turns a consumer outage into a broker
outage.
Ordering¶
Messages published on one channel to one queue are delivered in publish order. Redeliveries, priorities, delayed retries and multiple competing consumers all weaken that, because a requeued message goes back near the head while others have already been delivered. If strict per-entity ordering matters, use single active consumer (one consumer at a time, automatic failover to the next), or shard by key with a consistent-hash exchange so each entity always lands in the same queue.
Inside a node¶
Erlang and the process model¶
RabbitMQ runs on the BEAM, Erlang's virtual machine, which schedules millions of lightweight isolated processes across CPU cores. Nearly everything in RabbitMQ is one: each connection has a reader and writer, each channel is a process, each classic queue is one process, and each quorum queue member is a Ra server process. Processes communicate by message passing and fail independently: a crashing channel takes down that channel, not the node.
This model explains a lot of RabbitMQ's performance profile. A single classic queue is served by one process, so it can only use about one CPU core no matter how big the server is. Scaling throughput means using more queues, not bigger ones. Quorum queues pipeline work across several processes, but the same principle holds.
Flow control and resource alarms¶
Inside a node, RabbitMQ uses credit-based flow control between processes. If a queue can't
keep up, the channels publishing into it run out of credit and stop reading from their
connections, which pushes back through TCP to the publisher. A connection in this state shows
as flow in the management UI. Brief flow is normal; constant flow means publishers are
outrunning the queues.
Resource alarms are the hard stop. When memory use crosses vm_memory_high_watermark
(default 0.6 of available memory on current releases, read from cgroup limits in containers)
or free disk falls below disk_free_limit (default 50 MB, which is far too low), the node
blocks all publishing connections cluster-wide. Clients get connection.blocked; consumers
keep running, so the alarm clears once queues drain. Set disk_free_limit to at least a few
GB, or relative to RAM, and alert on alarms before users notice blocked publishers.
rabbitmq-diagnostics memory_breakdown shows where memory goes: queue processes and message
bodies, connection and channel buffers, quorum queue Raft state, the metadata store, and
management-plugin statistics (which can be surprisingly large on brokers with many objects).
Storage¶
Each node has a data directory named after the node (for example
/var/lib/rabbitmq/mnesia/rabbit@node1/ on package installs; the mnesia path name is
historical). It holds classic queue files and the shared message store, the quorum queue
write-ahead log (one WAL per node, shared by all quorum queues) with per-queue segment files
and snapshots, stream segment files, and the Khepri metadata store. Because the directory is
tied to the node name, changing a node's hostname orphans its data.
Quorum queues and streams fsync to disk before confirming, so disk latency directly bounds publish latency. Use SSDs, and on busy nodes give the data directory a dedicated volume.
Metadata: Khepri¶
Definitions (vhosts, users, permissions, exchanges, queues, bindings, policies, runtime parameters) live in Khepri, a Raft-based tree database replicated to every node. It replaced Mnesia, which was fast in good times but resolved network partitions by picking a winner and discarding the other side's changes, the source of most "split brain" war stories. Khepri behaves like the quorum queues: metadata changes need a majority of nodes, and a minority side of a partition can't change topology but also can't diverge.
Feature flags and deprecated features¶
Feature flags let nodes of different versions coexist during rolling upgrades: a new behaviour
is only switched on once every node supports it. They must be enabled explicitly after an
upgrade (rabbitmqctl enable_feature_flag all), and each minor release requires the
previous release's flags to be enabled before you can upgrade to it. Deprecated features go
through phases (permitted by default, denied by default, removed);
rabbitmq-diagnostics check_if_any_deprecated_features_are_used tells you if anything you
depend on is on the way out.
Clustering and high availability¶
A RabbitMQ cluster is a set of nodes sharing metadata through Khepri and connected by Erlang distribution. Every node knows every exchange, binding, user and policy, and a client can connect to any node and reach any queue: the node forwards traffic internally. What the cluster does not do on its own is replicate messages. A classic queue lives on one node and is unavailable while that node is down. Message-level redundancy comes from quorum queues and streams.
Forming a cluster¶
Nodes are named rabbit@<hostname> and must resolve each other's short hostnames. They
authenticate to each other with a shared secret, the Erlang cookie
(/var/lib/rabbitmq/.erlang.cookie, readable only by the rabbitmq user). CLI tools use the same
cookie, which is why rabbitmqctl on a host with a different cookie can't talk to the node.
Nodes join either by hand (rabbitmqctl join_cluster, see
Operations) or automatically at first boot via peer
discovery: a static list (classic_config), DNS A/AAAA records, or a backend for Kubernetes,
AWS, Consul or etcd. The lab uses classic_config.
Cluster size and majorities¶
Because both metadata and replicated queues need a majority, cluster size matters:
| Nodes | Majority | Tolerates | Notes |
|---|---|---|---|
| 1 | 1 | 0 | Fine for development |
| 2 | 2 | 0 | Worse than one: either node down loses the majority |
| 3 | 2 | 1 | The standard production size |
| 5 | 3 | 2 | For tolerating two failures, or one during maintenance |
| 7 | 4 | 3 | Rarely worth the replication overhead |
Quorum queue membership is separate from cluster size. By default a quorum queue gets up to
three members; set x-quorum-initial-group-size to choose. After adding nodes, existing queues
don't spread automatically: rabbitmq-queues grow adds members, and rabbitmq-queues rebalance
quorum redistributes leaders so one node isn't doing all the work.
Partitions and failures¶
When a node dies or is cut off, quorum queues with a majority elsewhere elect a new leader within seconds, and clients connected to surviving nodes carry on. Clients connected to the lost node see their connection drop and must reconnect, through a load balancer or a list of node addresses, then redeclare consumers. Messages they had unacknowledged are redelivered, and anything they published without a confirm must be resent.
On the minority side of a partition, nodes can't change metadata or commit to quorum queues;
they wait. When the partition heals, Raft brings them up to date. There's no
pause_minority/autoheal decision to configure any more, and no manual "pick a winner"
recovery. This is the single biggest operational improvement in 4.x.
Load balancers¶
A TCP load balancer (HAProxy, a cloud NLB, a Kubernetes Service) in front of the cluster
gives clients one address. Two settings matter. Its idle timeouts must comfortably exceed the
AMQP heartbeat interval (60 s by default), or it will kill healthy idle connections. And it
should close sessions to a node it marks down, so clients reconnect immediately rather than
waiting for heartbeats to expire. The lab's haproxy.cfg does both. Don't load-balance the
Erlang distribution ports.
Maintenance and upgrades¶
rabbitmq-upgrade drain puts a node in maintenance mode: it closes client connections,
transfers quorum queue leadership away and stops accepting new work, so the node can be
patched or restarted without surprises. rabbitmq-upgrade revive brings it back. Before
stopping a node, rabbitmq-queues check_if_node_is_quorum_critical tells you whether doing so
would leave any quorum queue without a majority. Rolling upgrades go one node at a time
through the supported upgrade path; blue-green (a new cluster, with consumers draining the old
one, or a shovel moving messages) avoids mixed-version operation entirely.
Across sites¶
A cluster belongs in one low-latency network; Raft and Erlang distribution don't tolerate WAN links well. Between sites or clusters, use the Shovel plugin (moves messages from a source queue to a destination, AMQP 0-9-1, AMQP 1.0, or "local" within a cluster as of 4.2) or Federation (exchanges or queues in one cluster receive messages from another, loosely coupled). The commercial Tanzu edition adds warm-standby replication for disaster recovery.
Kubernetes¶
The RabbitMQ Cluster Operator manages clusters as a RabbitmqCluster custom resource,
handling peer discovery, StatefulSets, persistent volumes and rolling upgrades; the Messaging
Topology Operator manages exchanges, queues, users and policies as further resources, which
suits GitOps workflows. Use them rather than hand-rolled StatefulSets: node identity, storage
and upgrade ordering are easy to get subtly wrong.
Security¶
Users and permissions. Users authenticate per connection and are granted permissions per vhost as three regular expressions over resource names: configure (declare/delete), write (publish to exchanges, bind) and read (consume, bind). Topic permissions can further restrict which routing keys a user may publish to on topic exchanges. Management access is governed by user tags:
| Tag | Grants |
|---|---|
management |
UI/API access to the user's own vhosts |
policymaker |
Above, plus policies and parameters in those vhosts |
monitoring |
Read-only view of all vhosts, connections, channels and node metrics |
administrator |
Everything |
Give each application its own user, scoped to its vhost and, ideally, to its own resource
names. Give monitoring systems a monitoring-tagged user, not an administrator.
The guest user can only connect from localhost by default. Delete it on any real
deployment (the official Docker image relaxes the localhost restriction, which is another
reason to set your own credentials).
TLS. Enable TLS on client listeners (5671, 5551, 8883, 15671) and use certificates your
clients verify. Mutual TLS lets clients authenticate with certificates instead of passwords
(rabbitmq_auth_mechanism_ssl), and Erlang distribution itself can run over TLS for
inter-node traffic.
Authentication backends. Besides the internal database, RabbitMQ can authenticate and authorize against LDAP, an HTTP service, or OAuth 2.0 (JWT access tokens from an identity provider such as Keycloak, Entra ID or Okta, with scopes mapped to RabbitMQ permissions). OAuth 2.0 also covers management UI single sign-on.
Limits contain noisy neighbours: per-vhost max-connections and max-queues, per-user
max-connections and max-channels.
Secrets. Protect the Erlang cookie like a root password. Values in rabbitmq.conf can be
encrypted (encrypted: tagged values, more widely supported as of 4.3). Never publish the
management or distribution ports to the internet.
Performance and capacity¶
Most RabbitMQ performance problems come from how clients use it, not from broker tuning.
Keep queues short. RabbitMQ is fastest when queues are near-empty and messages flow straight through. Large backlogs cost memory and disk I/O, slow down recovery after a restart, and with quorum queues make snapshots and replication heavier. Treat a growing queue as a symptom (consumers too slow or absent) and bound every queue with a length limit.
Parallelise with queues. One queue is served by roughly one core. If you need more throughput than a single queue gives, shard across several queues (a consistent-hash or modulus-hash exchange keeps per-key ordering) and run consumers on each.
Reuse connections and channels. Long-lived connections, one channel per thread, no
connection per publish. Connection churn shows up as rabbitmq_connections_opened_total
climbing steadily and is one of the alerts in the lab.
Set prefetch deliberately. Unlimited prefetch lets one consumer hoard the queue; a prefetch of 1 makes fast consumers wait on round trips. Start around 20-50 for fast handlers and 1-5 for slow, uneven ones, then measure (Lab 5.3 sweeps it).
Batch confirms and acks. Keep many publishes in flight and acknowledge with
multiple=true where the client allows. Both cut round trips dramatically.
Choose persistence and replication on purpose. Persistent messages in a replicated quorum queue cost three fsynced copies. That's the right price for orders and payments; it's wasteful for telemetry you'd drop anyway, which can go to a stream or a classic queue.
Mind message size. Payloads in the low kilobytes are ideal. Hundreds of kilobytes work but eat memory and bandwidth across replication; megabytes belong in object storage.
Provision the host. Fast local SSDs for the data directory, file-descriptor limits of at least 65,536 (every connection uses one), RAM sized so normal operation sits well below the watermark, and CPU proportional to the number of busy queues rather than to message rate alone.
Observability¶
Tools¶
The management plugin provides the web UI on 15672 and the HTTP API underneath it. It's the best way to look around and do one-off operations, but its built-in statistics aren't meant for long-term monitoring or alerting.
The CLI tools run on (or next to) a node and authenticate with the Erlang cookie:
| Tool | Purpose |
|---|---|
rabbitmqctl |
Service management: users, vhosts, permissions, policies, cluster membership, listing objects |
rabbitmq-diagnostics |
Health checks, status, alarms, memory breakdown, environment, logs, observer |
rabbitmq-queues |
Quorum queue and stream membership, leadership rebalancing, quorum-critical checks |
rabbitmq-plugins |
Enable, disable and list plugins |
rabbitmq-upgrade |
Drain/revive (maintenance mode), upgrade-related checks |
rabbitmqadmin (v2) |
HTTP API client for remote use; a standalone binary written in Rust |
Metrics¶
The rabbitmq_prometheus plugin serves metrics on port 15692 at three endpoints. /metrics
returns aggregated, node-level series and is cheap to scrape. /metrics/per-object returns
every series for every queue, connection and channel, which can be enormous on a busy broker.
/metrics/detailed returns per-object series for only the metric families you request (and,
on 4.3, only the queues you filter for), which is the right way to get per-queue depth. Scrape
each node directly, never through the load balancer. The RabbitMQ team publishes Grafana
dashboards for these, notably RabbitMQ-Overview (ID 10991) and Erlang-Distribution
(ID 11352).
The signals worth alerting on:
| Signal | Metric (Prometheus plugin) | Usually means |
|---|---|---|
| Node down | up == 0 |
Node or plugin unreachable |
| Memory / disk / fd alarm | rabbitmq_alarms_memory_used_watermark, ..._free_disk_space_watermark, ..._file_descriptor_limit |
Publishers are blocked now |
| Ready messages growing | rabbitmq_detailed_queue_messages_ready |
Consumers too slow or missing |
| Messages with no consumers | ..._messages_ready > 0 and rabbitmq_detailed_queue_consumers == 0 |
A consumer service is down |
| Unacked growing | rabbitmq_detailed_queue_messages_unacked |
Consumers stuck, or prefetch too high |
| Connection churn | rate(rabbitmq_connections_opened_total) |
A client opening connections per operation |
| Unroutable drops | rate(rabbitmq_global_messages_unroutable_dropped_total) |
Missing binding or wrong routing key |
| FD usage | rabbitmq_process_open_fds / rabbitmq_process_max_fds |
Connection leak, or limits too low |
Health checks¶
rabbitmq-diagnostics provides checks in increasing order of depth: ping (the runtime is
up), check_running (the RabbitMQ application is running), check_local_alarms,
check_port_connectivity, check_virtual_hosts, and check_if_node_is_quorum_critical. Use the
cheapest checks for container or Kubernetes liveness probes; a liveness probe that fails
during a slow but healthy startup or under load restarts nodes for no reason. The HTTP API
offers the same checks under /api/health/checks/ for external monitoring, and the lab's
healthcheck.sh wraps them.
Logs¶
Logs go to files under /var/log/rabbitmq/, the console, syslog, or the systemd journal,
with configurable levels per category (connection, channel, queue, upgrade, ...). JSON output
(log.console.formatter = json / log.file.formatter = json) makes them easy to ship with
Vector, Fluent Bit or Promtail. Connection-level logging is noisy; on busy brokers raise
log.connection.level to warning.
Production checklist¶
| Area | Do this | Why |
|---|---|---|
| Version | Run a supported series; enable all stable feature flags after every upgrade | Required for the next upgrade; old series stop getting fixes |
| Topology | Quorum queues for anything durable; classic only for exclusive/temporary queues | Replication and poison-message handling |
| Topology | Length limits and dead-lettering via policies | Bounded queues; changeable without redeclaring |
| Publishers | Confirms, mandatory or an alternate exchange, reconnect with republish |
No silent loss |
| Consumers | Manual acks, explicit prefetch, idempotent handlers, reject (not nack) for failures |
At-least-once without loops |
| Connections | Long-lived, named connections; heartbeats on; LB timeouts above heartbeat | Avoid churn and phantom disconnects |
| Cluster | 3 or 5 nodes, stable hostnames, peer discovery | Majority-based availability |
| Resources | disk_free_limit of several GB; FD limit ≥ 65,536; SSD data volume |
Alarms before outages |
| Security | Delete guest; per-app users and vhosts; TLS; protect the cookie; firewall 4369/25672 |
Least privilege |
| Monitoring | Prometheus per node, alerts on alarms/backlogs/no-consumers/churn | See problems before publishers block |
| Operations | Export definitions regularly; drain nodes before maintenance; rehearse node replacement | Recoverable topology, uneventful upgrades |
Further reading¶
- RabbitMQ documentation: https://www.rabbitmq.com/docs
- Release information and support timelines: https://www.rabbitmq.com/release-information
- RabbitMQ 4.3 highlights: https://www.rabbitmq.com/blog/2026/04/23/rabbitmq-4.3-release
- Quorum queues: https://www.rabbitmq.com/docs/quorum-queues
- Production checklist: https://www.rabbitmq.com/docs/production-checklist
- Monitoring with Prometheus: https://www.rabbitmq.com/docs/prometheus