The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-poison-message-blocking-queue

Operations Guides

RabbitMQ poison message loop: a stuck queue head and endless redelivery

A queue with active consumers is making no progress. Queue depth is flat, messages are being delivered, but almost nothing is being acknowledged. The oldest message in the queue keeps getting older. This is a poison message loop: one message at the head of the queue crashes or is rejected by every consumer that picks it up, gets requeued, and is delivered again, forever.

The loop is expensive in three ways. First, the poison message itself never completes. Second, every valid message behind it is blocked, because the requeued message returns to the head. Third, each cycle burns consumer CPU, downstream calls, and log volume for zero business output. If prefetch is greater than 1, it gets worse: the poison message drags a batch of valid messages into the requeue cycle with it.

What this means

In normal operation, a consumer receives a message, processes it, and acks it. The ack removes the message from the broker. In a poison loop, the consumer either crashes while processing the message or explicitly rejects it with basic.nack or basic.reject and requeue=true. Requeue puts the message back at (or near) the head of the queue. The broker immediately redelivers it, with the redelivered flag set. The consumer fails again. Repeat indefinitely.

From the broker’s perspective, nothing is “broken”: delivery is happening, the consumer is connected, and no alarms fire. The composite signal is what gives it away:

  • Queue depth stable: messages are not accumulating, but they are not draining either.
  • Head message age climbing: the oldest message keeps getting older even though deliveries are occurring.
  • Redeliver rate elevated: the same message is being sent to consumers over and over.
  • Ack rate near zero: deliveries are not converting into completed work.
  • Unacked count stable at 1 or a small number: the in-flight set never grows because the same small set cycles.
flowchart TD
  A[Consumer fetches head message] --> B{Processing succeeds?}
  B -- ack --> C[Message removed, next message delivered]
  B -- crash or nack requeue=true --> D[Message returned to queue head]
  D --> A
  D -.-> E[Valid messages wait behind the poison message]

The prefetch interaction deserves emphasis. With prefetch greater than 1, the consumer holds the poison message plus several valid messages as unacknowledged. If the consumer crashes or its channel closes, the broker requeues the whole unacked batch, and the valid messages are trapped in the loop with the poisoned one. One bad message stalls not just itself but everything in the prefetch window behind it.

Common causes

CauseWhat it looks likeFirst thing to check
Malformed or schema-incompatible payloadConsumer throws a deserialization or parse error on every deliveryConsumer application logs for a repeating exception on the same message
Consumer bug triggered by specific contentCrash or unhandled exception only for one message shapeConsumer error rate and stack traces correlated with this queue
Message exceeds consumer limits (size, complexity)Consumer OOMs, times out, or exceeds a processing deadline on one messageConsumer memory/timeout logs; message size on the queue
Downstream dependency failureEvery message fails, not just one; looks like a poison loop but all messages cycleWhether redeliver is elevated on multiple queues sharing the same dependency
No dead-letter exchange configuredFailed messages have nowhere to go except back to the headQueue arguments and policies for a dead-letter-exchange setting
No delivery limit (classic queues, or quorum queues with the limit disabled)Redelivery count unboundedQueue type and effective x-delivery-limit

Distinguish “one bad message” from “everything fails” early. If the downstream database is down, every message looks poisoned, and the fix is the database, not the queue. The tell: poison is usually one message looping while others queue up behind it; a dependency failure shows a high redeliver rate across the whole queue with a failing downstream.

Quick checks

Substitute your vhost (the default / is URL-encoded as %2F) and queue name. The curl and list_* commands are read-only; the peek at the end requeues what it reads, so use it carefully (see note below).

# Per-queue depth, unacked count, consumers, and head message age
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | {name, messages_ready, messages_unacknowledged, consumers, head_message_timestamp, state}'

# Per-queue message rates: look for redeliver >> ack
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | {name, message_stats}'

# Same thing via rabbitmqctl
rabbitmqctl list_queues name messages_ready messages_unacknowledged consumers

# Detailed consumer info (prefetch, ack mode, channel)
rabbitmqctl list_consumers

# Peek at the head message without removing it (requeueing get)
curl -s -u guest:guest -X POST \
  http://localhost:15672/api/queues/%2F/my_queue/get \
  -H 'content-type: application/json' \
  -d '{"count":1,"ackmode":"ack_requeue_true","encoding":"auto"}'

Note on the peek: ack_requeue_true requeues the message after reading it, which sets the redelivered flag and, on quorum queues, may count toward the delivery limit depending on your broker version: releases before 4.3 counted every acquisition against the limit, while since 4.3 only genuine failures (basic.reject, client crash, AMQP 1.0 modified with delivery-failed=true) increment the delivery count. On a queue already in a redelivery loop this is usually harmless, but do not script it in a polling loop.

What to look for:

  • Depth flat, head_message_timestamp old and getting older: the head is stuck. This field depends on publishers setting the timestamp property; it can be absent or 0, so treat a missing value as “unknown”, not “healthy”.
  • redeliver rate non-trivial while ack rate is near zero on one specific queue: redelivery loop confirmed. (message_stats is only present for queues that have seen message activity since stats collection started.)
  • Unacked count equal to consumer count x prefetch, or pinned at a small constant: consumers are holding but not completing work.
  • The peeked message payload: malformed JSON, unexpected encoding, a schema version the consumer does not understand, or an abnormally large body are the usual findings.

How to diagnose it

  1. Confirm the composite signal. For the suspect queue, verify all of: stable messages_ready, rising head age, redeliver rate well above zero, ack rate near zero, and consumers > 0. If depth is growing instead of flat, you are looking at a slow-consumer backlog, not a poison loop.
  2. Rule out a downstream dependency failure. Check whether other queues consumed by the same application are also failing, and whether the consumer’s downstream (database, API) is healthy. If every message fails, fix the dependency first; the messages are not poisoned.
  3. Identify the offending message. Use the requeueing get call above, or enable the firehose tracer (rabbitmq-plugins enable rabbitmq_tracing, then rabbitmqctl trace_on -p <vhost>) to capture deliveries on the affected queue. Look for the same message ID or payload appearing repeatedly with redelivered: true. Disable tracing afterwards (rabbitmqctl trace_off); leaving it on writes every traced message to disk.
  4. Inspect the consumer failure mode. In consumer logs, find the exception thrown for that message. Determine whether the consumer is nacking with requeue=true explicitly, or crashing and letting the channel/connection close requeue the message implicitly. Both produce the same loop; the fix differs.
  5. Check the queue’s defenses. Determine the queue type (GET /api/queues shows type in the arguments, or rabbitmqctl list_queues name type), then check whether a delivery limit and a dead-letter exchange are in effect via queue arguments and policies (rabbitmqctl list_policies). Classic queues have no native poison message handling; quorum queues track redeliveries and can enforce a limit.
  6. Check prefetch. rabbitmqctl list_consumers shows prefetch per consumer. High prefetch plus a poison head means a whole batch of valid messages is trapped in the loop with it.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-queue redeliver rateDirect measure of the loopElevated and sustained on one queue
Per-queue ack rate vs deliver rateDeliveries that never completeDeliver non-zero, ack near zero
head_message_timestamp (head age)Latency of the oldest waiting messageAge climbs while depth is flat
messages_ready trendDistinguishes poison (flat) from backlog (growing)Flat depth with active consumers and no progress
messages_unacknowledgedMessages trapped in consumer buffersPinned at prefetch x consumers, or at a small constant
consumer_utilisationWhether consumers can accept workLow with a backlog present
Dead-letter queue depthWhere poison messages should be landingGrowing DLQ confirms failures are being captured; empty DLQ with high redeliver means no DLX path

Fixes

Stop the loop right now

Remove or quarantine the poison message manually. The fastest relief is getting the message out of the head position. Options, in order of preference:

  • Use the management API get with "ackmode":"ack_requeue_false" to pull the message and discard it. Copy the payload to a file first if you need it for debugging. This destroys that one message; that is the point.
  • Purge the queue only if every message in it is expendable. rabbitmqctl purge_queue destroys all messages, not just the poisoned one. Treat this as a last resort.

Do not restart the broker. A restart requeues everything and the poison message will simply resume its loop, now mixed in with everything else.

Fix the consumer’s failure handling

Nack with requeue=false and route to a DLX. The consumer should catch processing failures it cannot retry (parse errors, schema mismatches, validation failures) and reject with requeue=false. If a dead-letter exchange is configured on the queue, the message is dead-lettered instead of dropped, preserving it for inspection and replay. Reserve requeue=true for transient failures where a retry genuinely has a chance to succeed, and even then prefer bounded retries over infinite requeue.

Handle crash-based requeue. If the consumer process is dying mid-processing, the broker requeues all its unacked messages when the channel closes. That path bypasses your nack logic entirely, which is exactly why a broker-side delivery limit matters.

Enforce a delivery limit at the broker

Quorum queues: x-delivery-limit. Quorum queues track unsuccessful delivery attempts and, once a message exceeds the limit, drop it or dead-letter it if a DLX is configured. Set it as a queue argument (x-delivery-limit) or via a policy (delivery-limit). The default and the exact counting behavior have changed across releases: RabbitMQ 3.13 had no default limit (opt-in via x-delivery-limit), 4.0 introduced a default of 20, and since 4.3 the delivery count increments only for genuine failures (basic.reject, client crash, AMQP 1.0 modified with delivery-failed=true), so routine requeues such as basic.nack no longer count against the limit. Check the docs for your version before tuning.

A few operator notes:

  • The policy key is delivery-limit; the queue argument is x-delivery-limit. Using the wrong name in the wrong place silently does nothing.
  • Classic queues do not support poison message handling in any current version. If the loop is on a classic queue, your only broker-side defense is a DLX plus requeue=false discipline in the consumer, or migrating the queue to quorum.
  • Give the dead-letter queue its own delivery limit and a consumer. A DLQ with no limit and a failing consumer can loop the same message between the DLQ and the original queue.
  • Watch out for dead-letter routing cycles: if the DLX can route a message back to its original queue, the message can circulate. The broker detects a cycle and drops the message only if the cycle contains no rejections; a rejection in the cycle defeats that protection.

Tune prefetch. Keep prefetch modest so one poison message traps a small batch, not thousands. There is no universal value; size it from message processing time and consumer count, and revisit it whenever you see unacked pinned at the prefetch ceiling.

Prevention

  • Quorum queues with a delivery limit and a DLX on every workload queue. This makes the broker, not the consumer, responsible for terminating poison loops. Consumers crashing, throwing, or nacking incorrectly all funnel into the same bounded behavior.
  • Consumer-side poison handling as a library pattern. Wrap message processing in a handler that classifies exceptions: transient (retry with backoff), permanent (reject with requeue=false). Do not leave this to per-service reimplementation.
  • Monitor the composite, not one signal. Alert on the combination: redeliver rate elevated AND ack rate near zero AND head age climbing, sustained over several minutes. Any single one of these has benign explanations; together they are diagnostic. See the signal taxonomy in RabbitMQ monitoring checklist: the signals every production broker needs.
  • Treat the DLQ as a first-class queue. Alert on DLQ depth growth, give it a consumer or a replay procedure, and cap its own delivery limit. An unwatched DLQ becomes the incident later.
  • Schema discipline. Most poison messages are schema drift: a producer shipped a payload shape the consumer cannot parse. Contract testing between producers and consumers catches these before deploy.

How Netdata helps

  • Per-queue redeliver, deliver, and ack rates side by side: the deliver-vs-ack divergence plus elevated redeliver is the poison signature, and seeing all three on one queue in one view is the fastest confirmation.
  • Queue depth and unacked trends: flat ready depth with a pinned unacked count distinguishes a poison loop from a consumer backlog, which changes the entire response.
  • Head message age where available: tracking head_message_timestamp over time turns “the queue seems stuck” into “the head has not moved in 40 minutes”.
  • Consumer count and utilisation correlation: confirms consumers exist and are attached, ruling out the look-alike case of zero consumers with growing depth.
  • Cross-signal anomaly detection: a redeliver-rate anomaly on one queue while global broker rates look normal is exactly the kind of localized pattern that per-queue anomaly flags surface before depth-based alerts ever would.