Andrew Mercer
on this page

RabbitMQ Overview

RabbitMQ is an open-source message broker written in Erlang. Applications hand it messages, it routes them into queues according to rules you define, and it delivers them to consumers with per-message acknowledgements, so work survives crashes, restarts and slow consumers. It speaks several protocols (AMQP 0-9-1, AMQP 1.0, MQTT, STOMP and its own stream protocol) and runs as a single node or as a cluster with replicated queues.

This page explains how it works and how to run it well. The Operations page is the command-level companion and the Lab is where to try all of it.

What changed since 3.x

Most RabbitMQ articles, Stack Overflow answers and generated tutorials still describe 3.x. On 4.3 several of those instructions silently do nothing, and a few now fail outright:

Old advice Status in 4.3 Use instead
Classic mirrored queues (ha-mode, ha-params, ha-sync-mode policies) Removed in 4.0; the policy keys have no effect Quorum queues (or streams)
Lazy queues (x-queue-mode: lazy) Declaring with x-queue-mode fails (CQv1 storage removed in 4.3) Nothing; classic queues already page to disk aggressively
cluster_partition_handling = pause_minority / autoheal Accepted but ignored; Raft handles partitions Odd-sized clusters, quorum queues
Mnesia as the metadata store Removed; Khepri (Raft-based) is the only store Plan for majority availability
Non-durable, non-exclusive ("transient") queues Rejected by default Durable queues, exclusive queues, or queue TTL
vm_memory_high_watermark_paging_ratio No longer meaningful Size memory by the watermark alone
Global QoS (basic.qos with global=true) Denied by default Per-consumer prefetch
queue-master-locator Denied by default queue-leader-locator
rabbitmq-delayed-message-exchange plugin Deprecated and archived TTL tiers, or quorum queue delayed retry for retries
rabbitmqadmin v1 (Python, downloaded from the node) Download endpoint removed rabbitmqadmin v2 (standalone binary)
Consumer timeout applies to every queue Only quorum queues enforce it Same setting, now graceful (cancels one consumer, not the channel)

Upgrades to 4.3 must come from the latest 4.2.x patch, with all stable feature flags enabled first. See upgrades.

Where RabbitMQ fits

RabbitMQ is a smart broker: the server decides where each message goes, tracks every delivery and acknowledgement, retries, expires and dead-letters. Clients stay simple. That makes it a good fit for:

  • Work distribution: background jobs, task queues, fan-out to worker pools, with fair dispatch, retries and backpressure.
  • Service decoupling and integration: events between services, routed by type or content, where each consumer gets its own queue and can be offline for a while.
  • Request/reply between services that shouldn't call each other directly.
  • IoT and edge ingestion via MQTT, bridged into the same routing model.
  • Buffering pipelines, such as log shippers (Vector, Logstash, Fluent Bit all have AMQP inputs and outputs) that need a durable buffer between producers and an indexer.

It's a weaker fit when you need a long, replayable event history consumed by many independent readers at very high throughput. That's Kafka's home turf, although RabbitMQ streams now cover a good part of it.

RabbitMQ Apache Kafka NATS (JetStream) ActiveMQ Artemis ZeroMQ
Model Broker with exchanges, queues and streams Partitioned, replicated log Subjects; JetStream adds persistence Broker (JMS-centric) Library; no broker at all
Routing Rich, broker-side (topic, headers, fanout, ...) By topic and partition key only Subject wildcards Addresses, JMS selectors You build it
Consumption Per-message ack, redelivery, DLX Consumer tracks offsets Ack-based or replay Per-message ack N/A
Replay Streams only Core feature JetStream Limited No
Ordering Per queue Per partition Per stream Per queue Per socket
Typical sweet spot Task queues, service events, RPC, MQTT Event sourcing, analytics, CDC Low-latency service mesh messaging Java/JMS estates Embedded, in-process messaging

Managed equivalents follow the same split: Azure Service Bus and Amazon SQS/SNS are queue-and-topic brokers in RabbitMQ's family; Azure Event Hubs and Amazon Kinesis/MSK sit in Kafka's. Amazon MQ offers managed RabbitMQ itself.

Protocols and ports

AMQP 0-9-1 is RabbitMQ's original protocol, and its model (exchanges, queues, bindings) is the broker's model. AMQP 1.0 has been a native core protocol since 4.0 and is where most new client features land; RabbitMQ's newer official clients use it. MQTT (3.1, 3.1.1, 5.0) and STOMP are plugins that map onto the same exchanges and queues, so an MQTT device can publish and an AMQP service can consume. Streams have their own binary protocol for high-throughput consumption.

Port Purpose Expose to
5672 / 5671 AMQP 0-9-1 and 1.0, plain / TLS Clients
5552 / 5551 Stream protocol, plain / TLS Stream clients
1883 / 8883 MQTT, plain / TLS Devices
61613 / 61614 STOMP, plain / TLS Clients
15672 / 15671 Management UI and HTTP API, plain / TLS Operators only
15692 Prometheus metrics Monitoring only
15674 / 15675 Web STOMP / Web MQTT (WebSockets) Browsers
4369 epmd, the Erlang port mapper Cluster nodes and CLI hosts only
25672 Erlang distribution (inter-node and CLI traffic) Cluster nodes and CLI hosts only
35672-35682 CLI tool connections Cluster nodes only

Ports 4369 and 25672 carry full control over the node to anyone holding the Erlang cookie. Never expose them beyond the cluster network.

The core model

 Producer                                   RabbitMQ (virtual host "/")                              Consumer
┌────────┐  publish(exchange, key)  ┌──────────────┐  binding: "app.#"  ┌─────────────────┐  deliver  ┌────────┐
│  app   │ ───────────────────────▶ │ lab.logs     │ ─────────────────▶ │ lab.app-logs    │ ────────▶ │ worker │
└────────┘     key="app.error"      │ (topic)      │                    │ (quorum queue)  │ ◀──────── └────────┘
                                    └──────┬───────┘                    └─────────────────┘    ack
                                           │ binding: "*.error"         ┌─────────────────┐
                                           └──────────────────────────▶ │ errors          │ ───▶ ...
                                                                        └─────────────────┘

Connections and channels. A client opens one long-lived TCP connection (optionally TLS) and multiplexes lightweight channels over it. Channels are where publishing, consuming and declaring happen. The usual layout is one connection per process and one channel per thread; opening a connection per message is the classic anti-pattern, because connection setup costs several network round trips and broker-side resources.

Virtual hosts partition one broker into isolated namespaces. Exchanges, queues, bindings, policies and permissions all belong to a vhost; a user's permissions are granted per vhost. Use them to separate applications or environments sharing a cluster.

Exchanges receive published messages and route them. Producers never publish to a queue directly; even "publishing to a queue" goes through the default exchange, a direct exchange that every queue is implicitly bound to under its own name.

Type Routes by Typical use
direct Exact match of routing key to binding key Task queues, routing by type
topic Dot-separated pattern match: * = exactly one word, # = zero or more Events by category (order.created, app.web.error)
fanout Ignores the key; copies to every bound queue Broadcast, cache invalidation
headers Message header values (x-match: all or any) Routing on several attributes
x-consistent-hash (plugin) Hash of routing key or header, weighted by binding Sharding work across queues while keeping per-key ordering
x-modulus-hash (core as of 4.3) Hash modulo the number of bindings Even distribution across a fixed set of queues

Bindings are the routing rules linking an exchange to a queue (or to another exchange, for exchange-to-exchange topologies). A message that matches no binding is unroutable: by default it's silently dropped. Publish with the mandatory flag to have it returned to the publisher instead, or give the exchange an alternate exchange that catches everything unroutable.

Queues hold messages until a consumer acknowledges them. A queue is FIFO, with exceptions for priorities, redeliveries and delayed retries. Queues are declared by clients or operators; declaration is idempotent as long as the properties match, and redeclaring with different arguments fails with 406 PRECONDITION_FAILED and closes the channel. That's why the lab declares each topology in one shared function.

Messages are an opaque body plus properties the broker understands:

Property Meaning
delivery_mode 2 = persistent (written to disk); 1 = transient
content_type, content_encoding For consumers; RabbitMQ doesn't inspect the body
message_id, correlation_id Application identifiers; correlation_id matches replies to requests
reply_to Where a reply should go (RPC)
expiration Per-message TTL in milliseconds, as a string
priority Used by priority-capable queues
timestamp, type, app_id, user_id Metadata; user_id is validated against the connection's user
headers Arbitrary key/value table; also where the broker writes x-death, x-delivery-count, etc.

The maximum message size is 16 MiB by default on 4.x. Large payloads belong in object storage with the message carrying a reference (the claim check pattern).

Queue types

RabbitMQ 4.3 has three queue types. Choosing between them is the most consequential design decision you'll make.

Classic queues

The original, non-replicated queue type. Each classic queue lives on one node and is served by a single Erlang process. They're fast, cheap and feature-complete for single-node use, and they're the right choice for data you can afford to lose with a node: exclusive reply queues, temporary per-client queues, caches of rebuildable work.

Since 4.3 only the second-generation storage (CQv2) exists. Small messages are embedded in the queue's own index files and larger ones go to a node-wide shared message store. Classic queues move messages to disk eagerly, so the old "lazy mode" is gone along with the storage engine that needed it. Classic queues support priorities (x-max-priority at declaration). They don't enforce consumer timeouts or delivery limits.

Quorum queues

The replicated, durable, data-safety-first queue type, built on the Raft consensus algorithm (via RabbitMQ's Ra library). A quorum queue has a leader and followers on different nodes; publishes and acks go through the leader and are committed once a majority of members has written them to disk. Publisher confirms are only sent after that commit, so a confirmed message survives the loss of any minority of nodes.

Quorum queues stay available as long as a majority of their members is up: a 3-member queue tolerates one node down, a 5-member queue two. If a majority is lost the queue is unavailable (not inconsistent) until enough members return.

Features specific to quorum queues:

  • Poison message handling. Each message carries a delivery-count; past the queue's delivery-limit (default 20 since 4.0) it's dead-lettered, or dropped if no DLX is set.
  • At-least-once dead-lettering (x-dead-letter-strategy: at-least-once, requires x-overflow: reject-publish). Dead-lettered messages aren't removed from the source until the target queues confirm them.
  • Priorities. 32 strict levels as of 4.3 (4.0-4.2 had two). No x-max-priority needed.
  • Delayed retry (4.3). Returned messages are set aside for a linearly increasing delay before redelivery: min(delayed-retry-min × delivery-count, delayed-retry-max).
  • Consumer timeout (enforced by the queue as of 4.3). A consumer that holds a delivery too long gets a basic.cancel for that one consumer; its messages are returned.
  • Single active consumer, consumer priorities, per-queue and per-message TTL, length limits.

Quorum queues are always durable and can't be exclusive. Every member stores every message, so a 3-member queue writes three copies; that's the cost of the safety. In 4.3 the per-message memory overhead was roughly halved for messages up to 32 KiB.

Streams

A stream is an append-only, replicated log. Consuming doesn't remove anything: each consumer reads from an offset it chooses (first, last, next, a numeric offset or a timestamp), and many consumers can read the same data independently and repeatedly. Retention is by size or age (x-max-length-bytes, x-max-age), applied per segment file.

Streams suit large fan-outs (thousands of consumers reading the same messages), replay, time-travel debugging and very high throughput. Over AMQP 0-9-1 they behave like a queue with an offset argument; the dedicated stream protocol (port 5552) is much faster and adds publisher deduplication, server-side filtering and super streams (a partitioned stream, the closest thing RabbitMQ has to a Kafka topic).

Choosing

Classic Quorum Stream
Replicated No Yes (Raft) Yes
Survives node loss No Yes, with a majority up Yes, with a majority up
Consumption Destructive Destructive Non-destructive, by offset
Exclusive / temporary Yes No No
Poison message handling No Yes N/A
Priorities Yes (x-max-priority) Yes (32 levels) No
Consumer timeout (4.3) No Yes No
Best for Temporary or reply queues, losable work Anything you can't lose Fan-out, replay, high-volume events

Rule of thumb for 4.x: make durable application queues quorum queues, use classic queues for exclusive or temporary queues, and reach for streams when consumers need history or there are a great many of them. You can set a vhost's default_queue_type to quorum so that clients which don't specify a type get quorum queues.

Queue arguments versus policies

Most queue behaviour can be set two ways: as x- arguments at declaration time, or as keys in a policy applied by name pattern. Arguments are baked in: changing them means deleting and redeclaring the queue. Policies can be changed at any time and take effect immediately, so operational settings (length limits, TTLs, dead-lettering, delivery limits, delayed retry) are usually better as policies, leaving arguments for things that define the queue's identity, such as its type. When both are set the argument generally wins; operator policies let administrators impose caps (like a maximum length) that application policies can't override.

Behaviour Queue argument Policy key
Queue type x-queue-type (argument only)
Max messages / bytes x-max-length, x-max-length-bytes max-length, max-length-bytes
Overflow behaviour x-overflow overflow
Message TTL x-message-ttl message-ttl
Delete unused queue after x-expires expires
Dead-letter exchange / key x-dead-letter-exchange, x-dead-letter-routing-key dead-letter-exchange, dead-letter-routing-key
Dead-letter strategy (quorum) x-dead-letter-strategy dead-letter-strategy
Delivery limit (quorum) x-delivery-limit delivery-limit
Delayed retry (quorum, 4.3) x-delayed-retry-type, -min, -max delayed-retry-type, -min, -max
Consumer timeout (quorum, 4.3) x-consumer-timeout consumer-timeout
Initial replica count (quorum) x-quorum-initial-group-size (argument only)
Leader placement x-queue-leader-locator queue-leader-locator
Single active consumer x-single-active-consumer (argument only)

Delivery guarantees

RabbitMQ gives you at-most-once or at-least-once delivery, depending on how you configure both ends. There's no exactly-once: it would require the broker to know whether your consumer's side effects happened, which it can't. Reliable systems combine at-least-once delivery with idempotent consumers.

At-least-once is a chain, and every link has to hold:

Link What it protects against How
Durable topology Queues and exchanges disappearing on restart durable=True (quorum queues and streams always are)
Persistent messages Messages in a classic queue lost on restart delivery_mode=2
Publisher confirms Publisher assuming the broker has a message it never stored confirm.select, then wait for the ack
Mandatory flag / alternate exchange Messages silently dropped as unroutable mandatory=True or an alternate-exchange
Replication Loss of the node holding the queue Quorum queues or streams
Manual consumer acks Consumer crashing mid-message auto_ack=False, ack after the work is done
Idempotent processing Duplicates caused by all of the above retrying Dedup by message ID, upserts, unique constraints

Publisher side

With confirms enabled on a channel, the broker acknowledges each publish once it has taken responsibility for it: written to disk for persistent messages in classic queues, committed by a majority for quorum queues and streams. A nack means the broker couldn't accept it (for example the target queue is full and set to reject-publish). Anything neither acked nor nacked when a connection drops is in an unknown state and must be republished, which is one of the sources of duplicates.

Waiting for a confirm after every message serializes publishing on network round trips. High-throughput publishers keep a window of unconfirmed messages in flight (hundreds to thousands) and handle confirms asynchronously as they arrive. Lab 5.2 measures the difference.

Consumer side

Consumers subscribe with basic.consume and the broker pushes messages to them. Prefetch (basic.qos) caps how many unacknowledged deliveries each consumer can have outstanding. The default is unlimited, which lets one consumer hoard an entire backlog in memory, so always set it. Low values (1-10) give fair dispatch for slow, uneven work; higher values (100-300) keep fast consumers busy when the network round trip matters.

A delivery ends in one of three ways:

Consumer action AMQP 0-9-1 method Effect
Processed successfully basic.ack Message removed
Failed, try again basic.reject / basic.nack with requeue=true Returned to the queue (or delayed, with delayed retry)
Failed permanently basic.reject / basic.nack with requeue=false Dead-lettered if a DLX is set, otherwise discarded

If the consumer's channel or connection closes, every unacknowledged delivery is requeued and redelivered with the redelivered flag set.

On 4.3 quorum queues, basic.reject and basic.nack mean different things for poison-message accounting. basic.reject (like a crash or lost connection) records a failed delivery and advances delivery-count. basic.nack records a delivery the consumer simply didn't process, which advances only the new acquired-count, so it never trips the delivery limit. A consumer that nacks a bad message forever loops forever. Use reject for failures. (Up to 4.2 both counted.)

Dead-lettering

A message is dead-lettered when it's rejected without requeue, its TTL expires, it's pushed out by a length limit (with drop-head or reject-publish-dlx overflow), or it exceeds a quorum queue's delivery limit. It's republished to the queue's dead-letter exchange, keeping its original routing key unless dead-letter-routing-key overrides it, with an x-death header recording the queue, reason, count and time of each death (plus x-first-death-reason and friends). Classic queues dead-letter at-most-once: a crash at the wrong moment can lose the message. Quorum queues can do it at-least-once.

Expiry and limits

Per-queue message TTL (message-ttl) expires messages as they reach the head of the queue. A per-message TTL (expiration property) works the same way, which means a long-TTL message blocks shorter-TTL messages behind it; that's why the delay pattern in Lab 3.2 uses one queue per delay. Queue TTL (expires) deletes a queue nobody has used for the given time, which is how you get self-cleaning temporary queues now that transient non-exclusive ones are disallowed.

Length limits (max-length, max-length-bytes) bound a queue. The overflow setting decides what happens when it's full: drop-head (default; discard or dead-letter the oldest), reject-publish (nack new publishes, so publishers with confirms feel backpressure), or reject-publish-dlx (classic queues only; nack and dead-letter the rejected message). Every production queue should have a limit: an unbounded queue turns a consumer outage into a broker outage.

Ordering

Messages published on one channel to one queue are delivered in publish order. Redeliveries, priorities, delayed retries and multiple competing consumers all weaken that, because a requeued message goes back near the head while others have already been delivered. If strict per-entity ordering matters, use single active consumer (one consumer at a time, automatic failover to the next), or shard by key with a consistent-hash exchange so each entity always lands in the same queue.

Inside a node

Erlang and the process model

RabbitMQ runs on the BEAM, Erlang's virtual machine, which schedules millions of lightweight isolated processes across CPU cores. Nearly everything in RabbitMQ is one: each connection has a reader and writer, each channel is a process, each classic queue is one process, and each quorum queue member is a Ra server process. Processes communicate by message passing and fail independently: a crashing channel takes down that channel, not the node.

This model explains a lot of RabbitMQ's performance profile. A single classic queue is served by one process, so it can only use about one CPU core no matter how big the server is. Scaling throughput means using more queues, not bigger ones. Quorum queues pipeline work across several processes, but the same principle holds.

Flow control and resource alarms

Inside a node, RabbitMQ uses credit-based flow control between processes. If a queue can't keep up, the channels publishing into it run out of credit and stop reading from their connections, which pushes back through TCP to the publisher. A connection in this state shows as flow in the management UI. Brief flow is normal; constant flow means publishers are outrunning the queues.

Resource alarms are the hard stop. When memory use crosses vm_memory_high_watermark (default 0.6 of available memory on current releases, read from cgroup limits in containers) or free disk falls below disk_free_limit (default 50 MB, which is far too low), the node blocks all publishing connections cluster-wide. Clients get connection.blocked; consumers keep running, so the alarm clears once queues drain. Set disk_free_limit to at least a few GB, or relative to RAM, and alert on alarms before users notice blocked publishers.

rabbitmq-diagnostics memory_breakdown shows where memory goes: queue processes and message bodies, connection and channel buffers, quorum queue Raft state, the metadata store, and management-plugin statistics (which can be surprisingly large on brokers with many objects).

Storage

Each node has a data directory named after the node (for example /var/lib/rabbitmq/mnesia/rabbit@node1/ on package installs; the mnesia path name is historical). It holds classic queue files and the shared message store, the quorum queue write-ahead log (one WAL per node, shared by all quorum queues) with per-queue segment files and snapshots, stream segment files, and the Khepri metadata store. Because the directory is tied to the node name, changing a node's hostname orphans its data.

Quorum queues and streams fsync to disk before confirming, so disk latency directly bounds publish latency. Use SSDs, and on busy nodes give the data directory a dedicated volume.

Metadata: Khepri

Definitions (vhosts, users, permissions, exchanges, queues, bindings, policies, runtime parameters) live in Khepri, a Raft-based tree database replicated to every node. It replaced Mnesia, which was fast in good times but resolved network partitions by picking a winner and discarding the other side's changes, the source of most "split brain" war stories. Khepri behaves like the quorum queues: metadata changes need a majority of nodes, and a minority side of a partition can't change topology but also can't diverge.

Feature flags and deprecated features

Feature flags let nodes of different versions coexist during rolling upgrades: a new behaviour is only switched on once every node supports it. They must be enabled explicitly after an upgrade (rabbitmqctl enable_feature_flag all), and each minor release requires the previous release's flags to be enabled before you can upgrade to it. Deprecated features go through phases (permitted by default, denied by default, removed); rabbitmq-diagnostics check_if_any_deprecated_features_are_used tells you if anything you depend on is on the way out.

Clustering and high availability

A RabbitMQ cluster is a set of nodes sharing metadata through Khepri and connected by Erlang distribution. Every node knows every exchange, binding, user and policy, and a client can connect to any node and reach any queue: the node forwards traffic internally. What the cluster does not do on its own is replicate messages. A classic queue lives on one node and is unavailable while that node is down. Message-level redundancy comes from quorum queues and streams.

Forming a cluster

Nodes are named rabbit@<hostname> and must resolve each other's short hostnames. They authenticate to each other with a shared secret, the Erlang cookie (/var/lib/rabbitmq/.erlang.cookie, readable only by the rabbitmq user). CLI tools use the same cookie, which is why rabbitmqctl on a host with a different cookie can't talk to the node.

Nodes join either by hand (rabbitmqctl join_cluster, see Operations) or automatically at first boot via peer discovery: a static list (classic_config), DNS A/AAAA records, or a backend for Kubernetes, AWS, Consul or etcd. The lab uses classic_config.

Cluster size and majorities

Because both metadata and replicated queues need a majority, cluster size matters:

Nodes Majority Tolerates Notes
1 1 0 Fine for development
2 2 0 Worse than one: either node down loses the majority
3 2 1 The standard production size
5 3 2 For tolerating two failures, or one during maintenance
7 4 3 Rarely worth the replication overhead

Quorum queue membership is separate from cluster size. By default a quorum queue gets up to three members; set x-quorum-initial-group-size to choose. After adding nodes, existing queues don't spread automatically: rabbitmq-queues grow adds members, and rabbitmq-queues rebalance quorum redistributes leaders so one node isn't doing all the work.

Partitions and failures

When a node dies or is cut off, quorum queues with a majority elsewhere elect a new leader within seconds, and clients connected to surviving nodes carry on. Clients connected to the lost node see their connection drop and must reconnect, through a load balancer or a list of node addresses, then redeclare consumers. Messages they had unacknowledged are redelivered, and anything they published without a confirm must be resent.

On the minority side of a partition, nodes can't change metadata or commit to quorum queues; they wait. When the partition heals, Raft brings them up to date. There's no pause_minority/autoheal decision to configure any more, and no manual "pick a winner" recovery. This is the single biggest operational improvement in 4.x.

Load balancers

A TCP load balancer (HAProxy, a cloud NLB, a Kubernetes Service) in front of the cluster gives clients one address. Two settings matter. Its idle timeouts must comfortably exceed the AMQP heartbeat interval (60 s by default), or it will kill healthy idle connections. And it should close sessions to a node it marks down, so clients reconnect immediately rather than waiting for heartbeats to expire. The lab's haproxy.cfg does both. Don't load-balance the Erlang distribution ports.

Maintenance and upgrades

rabbitmq-upgrade drain puts a node in maintenance mode: it closes client connections, transfers quorum queue leadership away and stops accepting new work, so the node can be patched or restarted without surprises. rabbitmq-upgrade revive brings it back. Before stopping a node, rabbitmq-queues check_if_node_is_quorum_critical tells you whether doing so would leave any quorum queue without a majority. Rolling upgrades go one node at a time through the supported upgrade path; blue-green (a new cluster, with consumers draining the old one, or a shovel moving messages) avoids mixed-version operation entirely.

Across sites

A cluster belongs in one low-latency network; Raft and Erlang distribution don't tolerate WAN links well. Between sites or clusters, use the Shovel plugin (moves messages from a source queue to a destination, AMQP 0-9-1, AMQP 1.0, or "local" within a cluster as of 4.2) or Federation (exchanges or queues in one cluster receive messages from another, loosely coupled). The commercial Tanzu edition adds warm-standby replication for disaster recovery.

Kubernetes

The RabbitMQ Cluster Operator manages clusters as a RabbitmqCluster custom resource, handling peer discovery, StatefulSets, persistent volumes and rolling upgrades; the Messaging Topology Operator manages exchanges, queues, users and policies as further resources, which suits GitOps workflows. Use them rather than hand-rolled StatefulSets: node identity, storage and upgrade ordering are easy to get subtly wrong.

Security

Users and permissions. Users authenticate per connection and are granted permissions per vhost as three regular expressions over resource names: configure (declare/delete), write (publish to exchanges, bind) and read (consume, bind). Topic permissions can further restrict which routing keys a user may publish to on topic exchanges. Management access is governed by user tags:

Tag Grants
management UI/API access to the user's own vhosts
policymaker Above, plus policies and parameters in those vhosts
monitoring Read-only view of all vhosts, connections, channels and node metrics
administrator Everything

Give each application its own user, scoped to its vhost and, ideally, to its own resource names. Give monitoring systems a monitoring-tagged user, not an administrator.

The guest user can only connect from localhost by default. Delete it on any real deployment (the official Docker image relaxes the localhost restriction, which is another reason to set your own credentials).

TLS. Enable TLS on client listeners (5671, 5551, 8883, 15671) and use certificates your clients verify. Mutual TLS lets clients authenticate with certificates instead of passwords (rabbitmq_auth_mechanism_ssl), and Erlang distribution itself can run over TLS for inter-node traffic.

Authentication backends. Besides the internal database, RabbitMQ can authenticate and authorize against LDAP, an HTTP service, or OAuth 2.0 (JWT access tokens from an identity provider such as Keycloak, Entra ID or Okta, with scopes mapped to RabbitMQ permissions). OAuth 2.0 also covers management UI single sign-on.

Limits contain noisy neighbours: per-vhost max-connections and max-queues, per-user max-connections and max-channels.

Secrets. Protect the Erlang cookie like a root password. Values in rabbitmq.conf can be encrypted (encrypted: tagged values, more widely supported as of 4.3). Never publish the management or distribution ports to the internet.

Performance and capacity

Most RabbitMQ performance problems come from how clients use it, not from broker tuning.

Keep queues short. RabbitMQ is fastest when queues are near-empty and messages flow straight through. Large backlogs cost memory and disk I/O, slow down recovery after a restart, and with quorum queues make snapshots and replication heavier. Treat a growing queue as a symptom (consumers too slow or absent) and bound every queue with a length limit.

Parallelise with queues. One queue is served by roughly one core. If you need more throughput than a single queue gives, shard across several queues (a consistent-hash or modulus-hash exchange keeps per-key ordering) and run consumers on each.

Reuse connections and channels. Long-lived connections, one channel per thread, no connection per publish. Connection churn shows up as rabbitmq_connections_opened_total climbing steadily and is one of the alerts in the lab.

Set prefetch deliberately. Unlimited prefetch lets one consumer hoard the queue; a prefetch of 1 makes fast consumers wait on round trips. Start around 20-50 for fast handlers and 1-5 for slow, uneven ones, then measure (Lab 5.3 sweeps it).

Batch confirms and acks. Keep many publishes in flight and acknowledge with multiple=true where the client allows. Both cut round trips dramatically.

Choose persistence and replication on purpose. Persistent messages in a replicated quorum queue cost three fsynced copies. That's the right price for orders and payments; it's wasteful for telemetry you'd drop anyway, which can go to a stream or a classic queue.

Mind message size. Payloads in the low kilobytes are ideal. Hundreds of kilobytes work but eat memory and bandwidth across replication; megabytes belong in object storage.

Provision the host. Fast local SSDs for the data directory, file-descriptor limits of at least 65,536 (every connection uses one), RAM sized so normal operation sits well below the watermark, and CPU proportional to the number of busy queues rather than to message rate alone.

Observability

Tools

The management plugin provides the web UI on 15672 and the HTTP API underneath it. It's the best way to look around and do one-off operations, but its built-in statistics aren't meant for long-term monitoring or alerting.

The CLI tools run on (or next to) a node and authenticate with the Erlang cookie:

Tool Purpose
rabbitmqctl Service management: users, vhosts, permissions, policies, cluster membership, listing objects
rabbitmq-diagnostics Health checks, status, alarms, memory breakdown, environment, logs, observer
rabbitmq-queues Quorum queue and stream membership, leadership rebalancing, quorum-critical checks
rabbitmq-plugins Enable, disable and list plugins
rabbitmq-upgrade Drain/revive (maintenance mode), upgrade-related checks
rabbitmqadmin (v2) HTTP API client for remote use; a standalone binary written in Rust

Metrics

The rabbitmq_prometheus plugin serves metrics on port 15692 at three endpoints. /metrics returns aggregated, node-level series and is cheap to scrape. /metrics/per-object returns every series for every queue, connection and channel, which can be enormous on a busy broker. /metrics/detailed returns per-object series for only the metric families you request (and, on 4.3, only the queues you filter for), which is the right way to get per-queue depth. Scrape each node directly, never through the load balancer. The RabbitMQ team publishes Grafana dashboards for these, notably RabbitMQ-Overview (ID 10991) and Erlang-Distribution (ID 11352).

The signals worth alerting on:

Signal Metric (Prometheus plugin) Usually means
Node down up == 0 Node or plugin unreachable
Memory / disk / fd alarm rabbitmq_alarms_memory_used_watermark, ..._free_disk_space_watermark, ..._file_descriptor_limit Publishers are blocked now
Ready messages growing rabbitmq_detailed_queue_messages_ready Consumers too slow or missing
Messages with no consumers ..._messages_ready > 0 and rabbitmq_detailed_queue_consumers == 0 A consumer service is down
Unacked growing rabbitmq_detailed_queue_messages_unacked Consumers stuck, or prefetch too high
Connection churn rate(rabbitmq_connections_opened_total) A client opening connections per operation
Unroutable drops rate(rabbitmq_global_messages_unroutable_dropped_total) Missing binding or wrong routing key
FD usage rabbitmq_process_open_fds / rabbitmq_process_max_fds Connection leak, or limits too low

Health checks

rabbitmq-diagnostics provides checks in increasing order of depth: ping (the runtime is up), check_running (the RabbitMQ application is running), check_local_alarms, check_port_connectivity, check_virtual_hosts, and check_if_node_is_quorum_critical. Use the cheapest checks for container or Kubernetes liveness probes; a liveness probe that fails during a slow but healthy startup or under load restarts nodes for no reason. The HTTP API offers the same checks under /api/health/checks/ for external monitoring, and the lab's healthcheck.sh wraps them.

Logs

Logs go to files under /var/log/rabbitmq/, the console, syslog, or the systemd journal, with configurable levels per category (connection, channel, queue, upgrade, ...). JSON output (log.console.formatter = json / log.file.formatter = json) makes them easy to ship with Vector, Fluent Bit or Promtail. Connection-level logging is noisy; on busy brokers raise log.connection.level to warning.

Production checklist

Area Do this Why
Version Run a supported series; enable all stable feature flags after every upgrade Required for the next upgrade; old series stop getting fixes
Topology Quorum queues for anything durable; classic only for exclusive/temporary queues Replication and poison-message handling
Topology Length limits and dead-lettering via policies Bounded queues; changeable without redeclaring
Publishers Confirms, mandatory or an alternate exchange, reconnect with republish No silent loss
Consumers Manual acks, explicit prefetch, idempotent handlers, reject (not nack) for failures At-least-once without loops
Connections Long-lived, named connections; heartbeats on; LB timeouts above heartbeat Avoid churn and phantom disconnects
Cluster 3 or 5 nodes, stable hostnames, peer discovery Majority-based availability
Resources disk_free_limit of several GB; FD limit ≥ 65,536; SSD data volume Alarms before outages
Security Delete guest; per-app users and vhosts; TLS; protect the cookie; firewall 4369/25672 Least privilege
Monitoring Prometheus per node, alerts on alarms/backlogs/no-consumers/churn See problems before publishers block
Operations Export definitions regularly; drain nodes before maintenance; rehearse node replacement Recoverable topology, uneventful upgrades

Further reading