The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-consumers-not-acking

Operations Guides

RabbitMQ consumers connected but not acknowledging: the consumer black hole

The queue has consumers. The management UI shows consumers: 3. Yet messages_ready keeps climbing, the ack rate is flat at zero, and nothing is draining. Everything looks connected and nothing is working. This is the consumer black hole: dangerous because every surface-level health check passes.

The trap is confusing connection health with consumer health. An open TCP connection and a registered consumer tag do not mean messages are being processed. A consumer can be connected, have channels open, and be completely stuck: deadlocked, blocked on a failed downstream dependency, or a zombie where the application process died but the OS never closed the socket. The broker will hold unacknowledged messages for it, pin them in RAM, and wait. Left alone, this escalates into the memory wall: unacked messages are not paged out, memory grows, the watermark is crossed, and every publisher in the cluster is blocked.

This article is about finding the black hole fast and telling the three root causes apart, because the fix for each is different.

What this means

RabbitMQ tracks consumer reality through three independent signals, and the black hole is defined by their divergence:

  • consumers > 0 on the queue: consumer registrations exist.
  • ack rate at or near zero: no processing is completing.
  • messages_ready rising: work is accumulating.

The deliver rate splits the pattern in two, and this is the single most useful first observation:

flowchart TD
  A[consumers > 0, ack rate ~0] --> B{deliver rate?}
  B -->|near zero| C[Consumers not pulling]
  B -->|non-zero| D[Consumers pulling but not acking]
  C --> C1[Flow control on connection]
  C --> C2[Prefetch already exhausted]
  C --> C3[Zombie: process dead, socket open]
  D --> D1[Stuck on downstream dependency]
  D --> D2[Deadlock / infinite loop]
  D --> D3[Nack-requeue loop, poison message]
  C3 --> E[Unacked pinned in RAM]
  D1 --> E
  D2 --> E
  E --> F[Memory alarm: all publishers blocked]

If deliver is also zero, consumers are not even pulling messages: suspect flow control on the connection, exhausted prefetch slots against a backlog of unacked messages, or dead processes behind live sockets. If deliver is non-zero but ack is zero, consumers receive messages but never finish: suspect a downstream outage, a deadlock, or a nack loop where the same messages are redelivered and rejected forever. Check the redeliver rate to separate the last case from the others.

Common causes

CauseWhat it looks likeFirst thing to check
Downstream dependency failure (database, API, filesystem)Consumer process alive, near-zero CPU, unacked count static at exactly prefetch size, deliver may continue until prefetch fillsConsumer application logs and downstream dependency health
Deadlock or infinite loop in consumer codeSame as above, but dependencies are healthy; thread dump shows no progressThread dump / profiler on the consumer process
Zombie connection (process dead, TCP open)Consumers still registered, deliver rate zero, no consumer-side process exists on the client hostWhether the client host and process actually still exist
Nack/requeue loop (poison message)High deliver rate, near-zero ack rate, elevated redeliver rate, queue depth roughly stablePer-queue redeliver rate and consumer error logs
Flow control throttling the consumer’s connectionConnection shows state: flow; deliver stalls though consumer is healthy; often the consumer also publishes on the same connectionGET /api/connections connection states
Silent channel closure swallowed by client libraryConsumer process running, but its channel is gone on the broker; consumer count lower than the app expectsBroker log for channel.close events vs what the app believes
Prefetch misconfigurationUnacked pinned at exactly prefetch x consumer count; a small number of stuck consumers hoard the entire backlograbbitmqctl list_consumers prefetch values

Quick checks

All read-only. Run against any node; queue state is cluster-global.

# 1. Per-queue view: consumers, ready, unacked, utilisation
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | {name, consumers, messages_ready, messages_unacknowledged, consumer_utilisation}'

# 2. Global rates: is ack actually zero, is deliver moving, is redeliver elevated?
curl -s -u guest:guest http://localhost:15672/api/overview | \
  jq '.message_stats | {publish, deliver, deliver_get, ack, redeliver}'

# 3. Detailed consumer view: which connections/channels, what prefetch, ack mode
rabbitmqctl list_consumers

# 4. Connection states: flow, blocked, blocking, running
curl -s -u guest:guest http://localhost:15672/api/connections | \
  jq 'group_by(.state) | map({state: .[0].state, count: length})'

# 5. Is this already escalating? Memory usage vs limit
curl -s -u guest:guest http://localhost:15672/api/nodes | \
  jq '.[] | {name, mem_used, mem_limit, mem_alarm}'

# 6. Recent channel/connection errors in the broker log
grep -i "channel.close\|connection.close\|error" /var/log/rabbitmq/rabbit@$(hostname).log | tail -30

What to look for:

  • Unacked == prefetch x consumers: classic stuck-consumer signature. Every consumer is holding a full prefetch buffer and not processing.
  • Unacked low, ready growing, deliver near zero: consumers are not pulling at all. Zombie connections or flow control.
  • Redeliver high with stable depth: nack loop, not a stuck consumer. See the poison message loop guide.
  • consumer_utilisation at 0.0 with ready > 0: the queue always has to wait for consumers, confirming consumers are the bottleneck rather than the broker. Note the field is called consumer_capacity in the Management API since RabbitMQ 3.8.13 (formerly consumer_utilisation, still returned as an alias for backwards compatibility) and is absent for queues with no consumers.

How to diagnose it

  1. Split the pattern by deliver rate. From check 2: deliver near zero means “not pulling”; deliver positive with ack zero means “not finishing.” Everything downstream branches on this.
  2. Check redeliver. If redeliver is elevated, you are in a poison-message loop, not a black hole. The consumer is alive and rejecting; the fix is a delivery limit and a dead-letter exchange, not a consumer restart.
  3. Verify the consumers actually exist. rabbitmqctl list_consumers gives you the connection and channel behind each consumer. Take the peer host and port, then confirm the process is alive on that host. If the host was terminated (spot instance reclaim, SIGKILL, hard crash), the TCP connection can linger: the broker only detects the dead peer after the heartbeat timeout expires (default 60 seconds, negotiated with the client, detected after two missed heartbeats). Until then, messages already delivered sit unacked against a corpse.
  4. If the consumer process is alive, find what it is waiting on. Near-zero CPU plus static unacked means blocked on I/O: a database that is down, an HTTP call with no timeout, a filesystem hang. A thread dump or the consumer’s own logs will name the dependency. This is the most common production cause.
  5. Check for flow control on the consumer’s connection. If the same connection both publishes and consumes, credit-based flow control triggered by the publish path throttles the whole connection, including delivery. Look for state: flow on that connection. See connection in flow state and flow control vs resource alarms.
  6. Compare what the app believes with what the broker sees. If the client library swallowed a channel error, the application thinks it is consuming while the broker closed the channel long ago. Broker-side channel.close log events, matched against the app’s consumer count, expose this.
  7. Measure the escalation runway. Unacked messages are not paged out; they hold RAM until acked or requeued, and rising ready messages add their own memory pressure. Watch mem_used / mem_limit: if the ratio keeps climbing with paging active, you are on the path to a cluster-wide publisher halt. Paging to disk is your early warning; see paging messages to disk.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-queue ack rateGround truth that processing completesZero while consumers > 0
Per-queue deliver (deliver_get) rateSplits “not pulling” from “not finishing”Zero with ready growing
Per-queue redeliver rateDistinguishes nack loops from stuck consumersElevated with stable depth
messages_unacknowledgedHolds RAM and is not paged out; number-one precursor to memory alarmsStuck at prefetch x consumers, or growing
consumer_capacity (formerly consumer_utilisation)Consumer effectiveness, not just presence0.0 with messages_ready > 0
Connection statesFlow control or alarm involvement on consumer connectionsConsumer connections in flow
head_message_timestampLatency that depth alone cannot showHead age climbing while depth looks flat; see head message age
mem_used / mem_limitThe escalation pathRising steadily during a black hole
Connection churnReconnect loops from crashing consumersChurn high while consumer count looks stable

Fixes

Stuck on a downstream dependency

Restore the dependency or unblock the consumer (add timeouts, circuit breakers). Once the consumer resumes, it will drain its unacked backlog naturally. If the dependency will be down for a while, stop the consumers cleanly: closing channels requeues their unacked messages back to ready, where other consumers or a later restart can pick them up. A clean stop is far better than leaving a full prefetch buffer pinned per consumer.

Zombie connections

If the client process is dead but the broker still shows the connection, close it via the management API (DELETE /api/connections/{name}) or the management UI. This is disruptive to that connection by design: closing it requeues its unacked messages immediately instead of waiting for heartbeat detection. Longer term, configure AMQP heartbeats on the client: without them, detecting a dead peer depends on TCP timeouts, which can take far longer than the 60-second heartbeat default. Also check any load balancer in front of the broker: if the LB idle timeout is shorter than the heartbeat interval, it can silently kill the pipe while both ends believe they are connected.

Deadlock or infinite loop in consumer code

Capture a thread dump first (you will want it for the postmortem), then restart the consumer process. Restarting closes channels and requeues unacked messages. Do not restart the broker: it does not fix a client-side deadlock and it costs you in-memory state and a long queue recovery.

Flow control on a shared connection

Split publishing and consuming onto separate connections. Flow control is per-connection; a consumer that also publishes can have its deliveries throttled by its own publish path. This is a design fix, not a broker fix.

Nack/requeue loop

Not a black hole, but it presents identically at the ack-rate level. Configure a delivery limit (x-delivery-limit on quorum queues) and a dead-letter exchange so the failing message is routed away after N attempts instead of cycling forever. Details in the poison message guide.

Consumer timeout as a safety net

RabbitMQ’s consumer acknowledgement timeout (default 30 minutes, evaluated at 1-minute intervals) requeues messages from consumers that hold them too long. It has existed since RabbitMQ 3.8.15, but its enforcement scope has changed across versions. As of RabbitMQ 4.3, consumer timeouts are enforced only on quorum queues; classic queues and streams no longer evaluate them. In 4.3, for AMQP 0-9-1 clients that support consumer_cancel_notify, only the timed-out consumer is cancelled and the channel stays open; clients without that capability still have the channel closed, as in earlier versions. When the timeout fires, the consumer’s messages are requeued, which also means duplicate work if the consumer was merely slow. Thirty minutes is a long time for a queue that normally processes in milliseconds; if your consumers are fast, a shorter per-queue timeout (x-consumer-timeout queue argument or policy) limits how long a stuck consumer can hoard messages. Treat this as a backstop, not a fix: a consumer that regularly trips the timeout is a bug to fix, and a timeout that fires during legitimately slow processing causes redelivery and duplicates.

Prevention

  • Alert on the composite, not the parts. consumers > 0 AND ack rate ~ 0 AND messages_ready rising for 5+ minutes is the black hole signature. None of the three alone is actionable; together they are unambiguous.
  • Monitor messages_unacknowledged separately from messages_ready. Unacked is the memory-alarm precursor. A consumer holding thousands of unacked messages with no acks for 5 minutes should alert before the memory watermark is reached.
  • Track consumer_utilisation on queues with backlog. Presence of consumers means nothing; utilisation tells you whether they are effective.
  • Set heartbeats explicitly on every client and keep the interval below any load balancer idle timeout on the path.
  • Add timeouts to every downstream call in consumer code. A consumer that can block forever will eventually block forever.
  • Separate publish and consume connections in apps that do both.
  • Configure delivery limits and a DLX on every queue with consumers, so one bad message cannot impersonate a dead consumer fleet.
  • Include consumer-side metrics (processing latency, error rate) in your dashboards. RabbitMQ cannot see inside your consumers; the broker-side signals in this article tell you that consumers are stuck, not why.

How Netdata helps

Netdata’s RabbitMQ collector pulls the management API signals that define this pattern, so you can correlate them on one dashboard instead of querying five endpoints during an incident:

  • Per-queue messages_ready vs messages_unacknowledged side by side, so unacked accumulation is visible instead of hidden inside total depth.
  • Per-queue publish, deliver, ack, and redeliver rates, so the deliver-zero-vs-ack-zero split and nack loops are distinguishable at a glance.
  • Per-queue consumer count and consumer utilisation, so “consumers attached but ineffective” shows up as a metric, not a guess.
  • Node memory usage against mem_limit plus alarm state, so you can see the black hole walking toward the memory wall and how much runway is left.
  • Connection and channel churn, which catches the crash-reconnect variant where consumer count looks stable but consumers are cycling.
  • Anomaly detection on ack rate, which flags the drop to zero even when absolute queue depth is still unremarkable.