The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-consumer-timeout

Operations Guides

RabbitMQ consumer timeout: delivery acknowledgement timed out and the channel is closed

You found this in the broker log or your client logs:

delivery acknowledgement on channel 1 timed out

In 3.13+ the fuller broker-side message names the consumer, queue, vhost, delivery tag, and the timeout value used, and the channel is closed with a PRECONDITION_FAILED (406) channel exception. The client library then typically reopens the channel and resubscribes, until it happens again.

This is RabbitMQ’s consumer_timeout mechanism firing. When a consumer holds a delivery without acknowledging it for longer than the configured timeout, the broker concludes the consumer is stuck, force-closes the channel, and requeues the unacknowledged messages. It is a protective feature, not a bug: an unacknowledged message is pinned in broker memory for potential redelivery, and a consumer that never acks is a slow-motion memory leak.

The timeout is almost never the root cause. It is the terminal symptom of a consumer that takes too long per message: a slow database call, a deadlocked handler, a downstream API that stopped responding, or a timeout value set without knowing the actual processing-time distribution. This guide covers how the mechanism works, how to tell “genuinely stuck” from “legitimately slow”, and how to fix it without just pushing the number up.

For the broader model of how unacked messages, prefetch, and consumer health interact, see How RabbitMQ actually works in production.

What this means

The mechanism, in order:

  1. A consumer subscribes with manual acknowledgements and receives deliveries, bounded by its prefetch count.
  2. The broker starts a per-delivery timer when it sends the message.
  3. The timeout is evaluated periodically, not continuously. Values below one minute are not supported, and values below five minutes are not recommended, so the effective timeout is approximate: it can fire somewhat later than configured.
  4. If any delivery remains unacknowledged past the timeout, the broker logs the timeout, closes the channel with PRECONDITION_FAILED, and requeues all unacked deliveries on that channel.
  5. Most client libraries react to the channel error by reopening the channel and resubscribing. The requeued messages are redelivered. If processing is still slow, the cycle repeats indefinitely.

The default is 30 minutes (1,800,000 ms). The version history matters because the default changed inside patch releases:

  • 3.8.15: mechanism introduced with a default of 15 minutes. This was a breaking change in a patch release and took down applications with long-running consumers that had upgraded routinely.
  • 3.8.17: default raised to 30 minutes after the fallout.
  • 3.12: per-queue configuration added, via the x-consumer-timeout queue argument and the consumer-timeout policy key.
  • 4.3: consumer timeout evaluation moved into quorum queues. Classic queues and streams no longer evaluate it, and for clients that advertise the consumer_cancel_notify capability the broker cancels only the timed-out consumer with a basic.cancel instead of closing the whole channel. If you upgrade to 4.3 and relied on this protection for classic queues, it silently stops applying.

Where the effective value comes from, highest precedence first:

  1. Consumer argument x-consumer-timeout on basic.consume
  2. Queue argument x-consumer-timeout at queue declaration
  3. Queue policy key consumer-timeout
  4. Global consumer_timeout in rabbitmq.conf

When both the queue argument and a policy are set, the lower of the two wins. Check all four levels before assuming you know which timeout is firing; the broker log line includes the timeout value used, which is the fastest way to identify the effective configuration.

Common causes

CauseWhat it looks likeFirst thing to check
Consumer processing genuinely too slowTimeouts cluster around the configured value; ack rate near zero between timeouts; unacked count sits at prefetchConsumer-side processing latency per message
Stuck or deadlocked consumerUnacked count frozen at exactly prefetch size, ack rate exactly zero, no progress everConsumer thread dumps, downstream dependency health
Poison messageSame queue times out repeatedly, redeliver rate elevated, head message age growsConsumer logs for a repeated exception on one payload
Downstream dependency failureTimeouts start suddenly across many consumers at onceDatabase, API, or service the consumer calls
Timeout set too low for the workloadLong batch jobs (reports, video, bulk imports) exceed 30 minActual p99 processing time vs effective timeout
Prefetch far too highOne consumer hoards hundreds of deliveries it cannot finish in timerabbitmqctl list_consumers prefetch column

The single most common pattern: unacked count sitting at exactly the prefetch size with an ack rate of zero. That is a consumer holding its entire prefetch buffer and processing none of it. See RabbitMQ unacknowledged messages growing for that failure pattern in depth.

Quick checks

All read-only and safe to run during an incident.

# Find timeout events and the timeout value actually used
grep -i "timed out waiting for\|delivery acknowledgement" /var/log/rabbitmq/rabbit@$(hostname).log | tail -20

# Which consumers exist, with their prefetch and ack mode
rabbitmqctl list_consumers queue_name consumer_tag prefetch_count ack_required active

# Per-queue: who has unacked messages piling up
rabbitmqctl list_queues name messages_ready messages_unacknowledged consumers

# Policies that may set consumer-timeout
rabbitmqctl list_policies

# Global configured value
rabbitmqctl environment | grep consumer_timeout
# Per-queue message stats: ack rate vs deliver rate vs redeliver rate
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | {name, messages_ready, messages_unacknowledged, consumers, message_stats}'

What to look for:

  • The log line names the queue and the timeout used. Start there. The queue tells you which consumer application to investigate; the timeout tells you which of the four configuration levels is in effect.
  • messages_unacknowledged at or near prefetch x consumers, ack rate zero: consumers are stuck, not slow.
  • Elevated redeliver on the affected queue: messages cycling through delivery, timeout, requeue, redelivery. Each cycle is wasted work and can mask a poison message.
  • A queue policy with consumer-timeout: per-queue overrides are a frequent surprise after upgrades, because the team that set the policy is often not the team reading the log at 3 a.m.

How to diagnose it

flowchart TD
  A[Channel closed:
delivery ack timed out] --> B{Ack rate before
the timeout?} B -->|Zero, unacked = prefetch| C[Stuck consumer:
deadlock, zombie process,
downstream down] B -->|Nonzero but too slow| D[Slow consumer:
processing time exceeds timeout] B -->|Timeouts repeat on one queue| E{Redeliver rate high?} E -->|Yes| F[Poison message loop:
same payload fails repeatedly] E -->|No| G[Workload exceeds timeout:
batch jobs, large payloads] C --> H[Fix consumer or dependency,
then reduce prefetch] D --> H F --> I[Nack with requeue=false + DLX,
x-delivery-limit on quorum queues] G --> J[Raise timeout deliberately
at queue or consumer level]
  1. Identify the queue and effective timeout from the broker log line. This pins down which application is affected and which configuration level fired.
  2. Check the ack rate for that queue over the window before the timeout. Zero acks with deliveries happening means stuck. Low but nonzero acks means slow.
  3. Check the unacked pattern. Unacked frozen at exactly prefetch size is the signature of a consumer holding its whole buffer and doing nothing, with the timeout as its visible endpoint.
  4. Look at redeliver rate on the queue. Elevated redeliver plus repeated timeouts on the same queue points at a poison message: one payload that crashes or hangs the handler every time it is delivered.
  5. Correlate with downstream dependencies. A sudden onset across many consumers at once almost always means the database, API, or service the consumers call degraded. Consumer-side health is invisible to the broker; check it on the application side.
  6. Measure actual processing time. Get the p50/p95/p99 per-message latency from the consumer application. If p99 is anywhere near the timeout, occasional timeouts are guaranteed under load, GC pauses, or dependency jitter.
  7. Check prefetch. prefetch x processing_time is the time budget a consumer needs to drain its buffer. Prefetch of 500 with 10-second processing is 83 minutes of buffered work against a 30-minute timeout: a timeout is mathematically certain after any hiccup.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
messages_unacknowledged per queueUnacked messages are pinned in broker memory and cannot be paged outFrozen at prefetch x consumers, or growing steadily
Ack rate per queueGround truth for “consumers are finishing work”Zero while deliver rate is nonzero
Redeliver rate per queueDetects timeout/requeue cycles and poison messagesSustained elevation on one queue
Consumer utilisation per queueWhether attached consumers can actually receive workBelow 0.5 with messages_ready > 0
Head message ageLatency of the oldest waiting messageGrowing while depth stays flat (poison message)
Channel churn rateTimeout events close channels; churn is the fingerprintSpikes correlated with timeout log lines
Consumer-side processing latencyThe real input to timeout sizingp99 approaching the configured timeout

Protocol-level channel errors like PRECONDITION_FAILED are visible only in the logs, not in the Management API or Prometheus metrics. Log collection for the timeout message is what ties the metric symptoms to the mechanism.

Fixes

Fix the stuck consumer

If acks are zero, raising the timeout only delays the next closure. Find the stall: thread-dump the consumer, check its downstream dependencies, look for lock contention or an exhausted database connection pool. A consumer whose process died but whose TCP socket is still open (a zombie connection) will also hold deliveries until the timeout fires; closing the zombie connection via the management API requeues its messages immediately instead of waiting.

Fix the poison message

A payload that deterministically hangs or crashes the handler will time out on every delivery attempt. The fix is in the consumer: catch processing failures, nack with requeue=false, and route to a dead-letter exchange. On quorum queues, set x-delivery-limit so a repeatedly failing message dead-letters after N attempts instead of cycling forever.

Reduce prefetch

If consumers are honest but slow, prefetch is usually the amplifier. Each delivery in the prefetch buffer starts its own timer on receipt, and the consumer can only work one (or a few) at a time. Size prefetch so the whole buffer drains well within the timeout: prefetch x worst_case_processing_time should be a small fraction of the timeout, not a multiple of it.

Raise the timeout deliberately

Legitimate for long-running work: report generation, media processing, bulk imports. Set it as close to the workload as possible, not globally:

# Per-queue policy (quorum queues in 4.x; queue-scoped override in 3.12+)
rabbitmqctl set_policy queue_consumer_timeout "^batch_imports\." \
  '{"consumer-timeout":7200000}' --apply-to quorum_queues

Tradeoffs to be explicit about:

  • A higher timeout means a genuinely stuck consumer holds pinned memory longer. The timeout exists because unacked messages are the number-one precursor to memory alarms.
  • The mechanism detects stuck consumers; it does not enforce processing SLAs. Do not tune it down to chase latency.
  • Since RabbitMQ 4.3 moved timeout handling into quorum queues, setting consumer_timeout to undefined in advanced.config no longer disables the timeout and breaks consumer registration (the quorum-queue path requires an integer timeout). If you must effectively disable it, set a very high value instead.
  • In 4.3+, this only protects quorum queues. Classic queues no longer evaluate the timeout at all.

What not to do: restart the broker. The stuck deliveries are requeued on channel closure anyway; a restart loses in-memory state, forces every consumer to reconnect, and changes nothing about why processing is slow.

Prevention

  • Measure before you configure. Know the p99 processing time per queue before touching consumer_timeout. Size the timeout as a multiple of p99, not a round number.
  • Keep prefetch honest. The prefetch buffer must be drainable in a fraction of the timeout. This single constraint prevents most timeout incidents on healthy consumers.
  • Dead-letter by default. Every queue whose consumers can fail on a payload should have a DLX, and quorum queues should carry x-delivery-limit.
  • Alert on the leading indicators, not the log line. Unacked count at prefetch with zero ack rate, and per-queue redeliver rate, both fire minutes to hours before the first channel closure.
  • Watch consumer-side latency. The broker cannot see why a consumer is slow. Instrument processing time in the consumer application and alert when p99 trends toward the timeout.
  • Plan the 4.3 migration. If you rely on consumer timeouts on classic queues, that protection disappears on upgrade to 4.3. Migrate queues that need it to quorum queues first.

How Netdata helps

  • Per-queue messages_ready and messages_unacknowledged charts make the “unacked frozen at prefetch” signature visible at a glance, before the timeout ever fires.
  • Message rate charts (publish, deliver, ack, redeliver) let you confirm the diagnosis in one view: deliver nonzero, ack zero, redeliver climbing is the timeout loop in progress.
  • Channel churn correlation ties channel closures back to consumer timeout events in the broker log, distinguishing them from client bugs or protocol errors.
  • Consumer utilisation per queue separates “no consumers attached” from “consumers attached but ineffective”, which changes the fix.
  • Alerting on sustained unacked growth and elevated redeliver rate gives you the early warning the timeout mechanism itself cannot provide, since the log line only appears after the damage is done.
  • Correlating broker memory against unacked totals shows the cost of a raised timeout: how much pinned memory a stuck consumer accumulates while it waits to be detected.