The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-queue-no-consumers

Operations Guides

RabbitMQ queue with zero consumers: zombie queues that grow until they OOM the node

A queue with zero consumers and a non-zero publish rate is a memory leak with a routing key. Every published message lands in the queue, nothing drains it, and the backlog grows until something stops it: a TTL, a queue length limit, a manual purge, or the node’s memory alarm. If none of the first three happen, you find out when publishers across the whole cluster get blocked.

This failure is invisible in global metrics. Cluster-wide consumer count stays healthy because your other queues have plenty of consumers. Cluster-wide queue depth grows, but if you run batch workloads, growth alone does not page anyone. The only signal that isolates this failure early is the per-queue consumer count.

This guide covers how to confirm a zombie queue, how to distinguish a real consumer outage from an intentional accumulating queue, and how to drain or bound the damage without making things worse.

What this means

When a consumer subscribes with basic.consume, RabbitMQ registers it against the queue and starts delivering messages up to the prefetch limit. When the last consumer disconnects, unsubscribes, or its channel dies, the queue’s consumer count drops to zero. The queue itself does not care. It keeps accepting every message routed to it.

What happens next depends on queue type and version:

  • Classic queues (pre-3.12 behavior): messages sit in RAM until the paging ratio (default 0.5 of the memory watermark) forces them to disk. Deep zero-consumer queues eat memory first, disk second.
  • Lazy and modern classic queues (3.12+): messages go to disk much more aggressively. High depth no longer threatens memory directly; disk free space and disk I/O take the hit instead.
  • Quorum queues: every message is appended to a Raft WAL on each member. A zero-consumer quorum queue grows WAL segments on multiple nodes at once, and disk consumption can outpace snapshot compaction. This is the fastest path to a disk alarm.

Either way, the endgame is a resource alarm. Memory crosses vm_memory_high_watermark (default 0.4 of RAM in 3.x, 0.6 in 4.x) or free disk drops below disk_free_limit (default 50MB), and the alarm blocks all publishers on all nodes. Consumers still work, but nothing new comes in. One forgotten queue can halt ingestion for the entire cluster.

flowchart TD
  A[Consumer app crashes or is undeployed] --> B[Queue consumers = 0]
  B --> C[messages_ready grows with every publish]
  C --> D{Paging / WAL writes to disk}
  D --> E[Memory ratio climbs toward watermark]
  D --> F[Disk free drops toward disk_free_limit]
  E --> G[Memory alarm: all publishers blocked cluster-wide]
  F --> H[Disk alarm: all publishers blocked cluster-wide]
  G --> I[Ingestion halt or OOM kill]
  H --> I

Common causes

CauseWhat it looks likeFirst thing to check
Consumer deployment failureConsumers dropped to zero at a specific time; often coincides with a deploy, crash loop, or scaled-to-zero replica setDeployment and pod/instance history for the consumer app
Queue declared but consumer never deployedNew queue, publish rate > 0 since creation, consumer count has always been 0When the queue was created and by which application
Intentional batch or scheduled-consumer queueDepth grows and drains on a schedule (hourly, nightly); consumer count is 0 between runsHistorical depth pattern: does it ever drain?
Zombie connections hiding a dead consumerConsumer count > 0 but ack rate is 0; looks subscribed but nothing processeslist_consumers and per-queue ack rate
Routing change sends traffic to the wrong queueA previously quiet queue suddenly has publish rate and zero consumersExchange bindings and recent topology changes
Temporary queue that lost its auto-delete triggerNon-auto-delete durable queue left behind by a retired applicationQueue arguments (x-expires, auto-delete flag) and last consumer activity

The distinction that matters operationally: a batch queue that accumulates and drains on a schedule is normal, and alerting on it is noise. A queue whose consumer count silently went to zero and stays there is an incident in progress. Depth history tells you which one you have.

Quick checks

All of these are read-only.

# The core signal: per-queue consumers next to depth
rabbitmqctl list_queues name consumers messages_ready messages_unacknowledged

Any line with consumers = 0 and a positive, growing messages_ready is a zombie queue.

# Same data via the management API, including publish rate per queue
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | select(.consumers == 0 and .messages_ready > 0) |
      {name, vhost, messages_ready, consumers,
       publish_rate: .message_stats.publish_details.rate}'

The publish rate matters: a zero-consumer queue with publish_rate = 0 is a leftover, not a leak. A zero-consumer queue with a positive publish rate is actively filling.

# Check how much runway the backlog is buying you
curl -s -u guest:guest http://localhost:15672/api/nodes | \
  jq '.[] | {name, mem_used, mem_limit,
             ratio: (.mem_used / .mem_limit),
             disk_free, disk_free_limit, mem_alarm, disk_free_alarm}'
# See which exchanges feed the queue, and which policies bound it
rabbitmqctl list_bindings
rabbitmqctl list_policies
# If consumer count is non-zero but you suspect zombies, list actual consumers
rabbitmqctl list_consumers

How to diagnose it

  1. Enumerate zero-consumer queues. Run the list_queues command above and look at messages_ready. The queue with the most messages and zero consumers is your incident.

  2. Check whether it is still being fed. Pull the per-queue publish rate from the management API. If publish rate is zero, the queue is inert: a cleanup task, not an emergency. If publish rate is positive, every minute of delay costs memory or disk.

  3. Decide: outage or by design? Look at depth history. A queue that grows for 23 hours and drains at 01:00 every night is a scheduled batch consumer; the fix is to exclude it from zero-consumer alerting, not to page anyone. A queue whose consumers vanished at deploy time and never came back is an outage. If you have no per-queue history, that is the monitoring gap to close after the incident.

  4. Check for zombie consumers. If the consumer count is non-zero but ack rate is zero, you have a different pattern: consumers connected but not processing. Causes include deadlocked application threads and dead client processes whose TCP connections are still open. Verify consumer-side application health, and check consumer_utilisation for the queue: a value near 0 with a growing backlog confirms consumers are attached but ineffective.

  5. Quantify the blast radius. Compute mem_used / mem_limit and disk_free / disk_free_limit from the node API. Above 0.7 memory ratio, paging is likely active and the alarm is approaching. Below 3x the disk limit, disk is your binding constraint. Estimate runway from the growth rate: (mem_limit - mem_used) / memory_growth_rate or (disk_free - disk_free_limit) / disk_consumption_rate.

  6. Identify the owning application. Bindings tell you which exchange feeds the queue. The queue name, vhost, and the user that declared it usually identify the owning team. If the consumer application was scaled to zero, crashed, or is stuck in a crash loop, restoring it is the fix.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-queue consumer countThe only signal that isolates this failure; global consumer count hides a single queue at zero0 on a queue that should have consumers
Per-queue messages_ready trendTells you how fast the zombie is filling and whether it ever drainsSustained positive growth with consumers = 0
Per-queue publish rateDistinguishes an active leak from an inert leftover queue> 0 on a zero-consumer queue
mem_used / mem_limitRunway to the memory alarm; paging starts at the paging ratio (default 0.5 of the watermark)> 0.7 sustained
disk_free / disk_free_limitRunway to the disk alarm, the usual endgame for lazy and quorum queues< 3.0, and < 1GB absolute
Queue paging countersConfirm the backlog has reached disk; a leading indicator before the alarmPer-queue messages_paged_out / messages_paged_out_bytes in the management API; rabbitmq_detailed_queue_messages_paged_out / _bytes on the Prometheus plugin’s per-object endpoint (absent from the aggregated scrape)
consumer_utilisationDistinguishes “no consumers” from “consumers attached but stuck”Near 0 with backlog growing and consumers > 0
Ack rate per queueGround truth that processing happens at all0 while consumers > 0 (zombie connections)

For alerting, the composite condition is consumers == 0 AND messages_ready > 0 (or publish rate > 0), sustained long enough to cover scheduled batch windows, with a per-queue allowlist for intentional accumulating queues. If you export via the Prometheus plugin, the equivalent expression is rabbitmq_queue_consumers == 0 and rabbitmq_queue_messages > 0. Either way, the alert must be per-queue; a global consumer gauge will not fire.

Fixes

Restore the consumers

If the consumer application crashed or was undeployed, bringing it back is the correct fix. Once consumers reconnect, the queue drains at the rate the consumers can process. Watch messages_ready trend down and confirm ack rate tracks delivery rate. If the backlog is huge, scale consumers up temporarily, but watch the consumer-side downstream dependencies (database, APIs) that the drain surge will hit.

Purge the queue (destructive)

If the messages are expendable (telemetry, cache-warming jobs, anything idempotent or replayable) and the alarm is close, purging is the fastest relief:

# DESTRUCTIVE: deletes all messages in the queue without delivering them
rabbitmqctl purge_queue <queue_name>

This discards every message permanently. Confirm with the owning team that the messages are replayable or worthless before running it. Do not purge a queue whose owner you cannot identify during a live incident without sign-off.

Delete the queue (destructive)

If the queue is an orphan (no consumers, no owner, no value), deleting it removes the leak entirely. Deleting a queue drops its messages. Verify bindings first so you know what stops receiving copies, and get sign-off from the owning team if one exists.

Do not restart the broker

Restarting does not fix a zero-consumer queue. Durable queues reload from disk on boot, the messages come back, and the memory spike during queue recovery can make things worse. Do not treat a broker restart as a first response to a memory-wall incident.

Bound the queue with policy

For queues that are intentionally accumulating, add guardrails so a consumer outage cannot take the node down:

  • Message TTL (x-message-ttl via policy): messages expire and are dropped or dead-lettered after a fixed age.
  • Queue length limit (max-length or max-length-bytes): oldest messages are dropped (or dead-lettered, if a DLX is configured) once the limit is hit. Prefer max-length-bytes for bounding memory and disk directly.
  • Dead-letter exchange: route expired or overflowed messages somewhere inspectable instead of silently discarding them.

Every bound has a tradeoff: TTL and length limits lose data by design. That is acceptable for telemetry and unacceptable for financial transactions. Set bounds per queue according to what the data is worth.

Prevention

  • Alert per-queue, not globally. The zero-consumer alert (consumers == 0 and messages or publish rate > 0) must evaluate per queue, with a documented allowlist for batch and scheduled-consumer queues. Global consumer count will never catch this.
  • Use auto-delete and exclusive queues for ephemeral work. Queues that exist to serve one consumer should die with that consumer instead of lingering.
  • Set queue expiry for forgotten queues. x-expires deletes a queue after it has been unused (no consumers, no redeclaration) for a period. It is a garbage collector for abandoned topology.
  • Bound every accumulating queue. Any queue designed to hold a backlog should have max-length-bytes or a message TTL and a dead-letter exchange. Unbounded durable queues are how one forgotten consumer becomes a cluster-wide publisher halt.
  • Set heartbeats so dead consumers actually disconnect. Clients killed without a clean shutdown (spot instance termination, OOM kill, kill -9) leave connections that look alive until TCP times out. Heartbeats let the broker detect the dead peer, close the channel, and requeue unacked messages. If you terminate consumer instances aggressively, this is not optional.
  • Watch the paging ratio as the early warning. Memory alarms are cliff-edge. Paging to disk (default at 0.5 of the watermark) is the last graduated signal before the cliff, and disk I/O plus queue paging counters are how you see it.
  • Raise disk_free_limit from the 50MB default. If your zombie-queue endgame is disk exhaustion, the default limit gives you almost no runway. See the related guides on the disk alarm.

How Netdata helps

  • Netdata collects per-queue consumer count, messages_ready, and messages_unacknowledged from the management API, so a queue drifting to zero consumers shows up on the queue’s own charts rather than being averaged away in cluster totals.
  • Correlating per-queue publish rate against consumer count on the same dashboard separates “actively filling zombie” from “inert leftover” without manual API queries.
  • Node-level mem_used / mem_limit and disk free charts next to queue depth charts let you read runway visually: which queue is growing, and how much room is left before the alarm.
  • Paging activity and disk I/O on the data partition surface the graduated warning that precedes the cliff-edge memory alarm.
  • Per-queue alerting lets you page on zero consumers for critical queues while suppressing the alert for known batch queues, instead of choosing between noise and blindness.