The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-prefetch-count-tuning

Operations Guides

RabbitMQ prefetch count: unacked hoarding versus idle consumers

The basic.qos prefetch count is the most consequential consumer-side setting in RabbitMQ, and the one most teams never touch. The default is unlimited. That default is why a single slow consumer can pin thousands of messages in RAM while every other consumer on the same queue sits idle, and why “we added more consumers but nothing got faster” is such a common report.

Prefetch has two opposite failure modes, diagnosed with different signals. Set too high, one consumer hoards unacknowledged messages that are pinned in broker memory and invisible to other consumers. Set too low, consumers spend most of their time waiting for the next delivery instead of processing. This guide covers the mechanism, both failure modes, and how to pick a defensible number.

What prefetch actually controls

Prefetch is a per-consumer credit limit. When a consumer subscribes with basic.consume, the broker pushes up to prefetch_count messages to that consumer without waiting for acknowledgements. Every ack frees one slot, and the broker immediately delivers one more message to fill it.

Three properties of this mechanism drive everything else:

  • Unacked messages are pinned in memory. The broker must hold every unacked message available for instant redelivery if the consumer dies. Unlike ready messages, unacked messages cannot be paged to disk. Memory usage grows linearly with unacked count x message size, a direct precursor to the memory alarm, which blocks all publishers cluster-wide.
  • A hoarded message is a stalled message. While a message sits in one consumer’s prefetch buffer, no other consumer can process it. If that consumer is slow, stuck, or deadlocked, every message it holds is frozen until it acks, nacks, times out, or its channel closes.
  • Channel close triggers a mass requeue. When a consumer’s channel or connection dies, all of its unacked messages return to the queue in one burst. A consumer holding 10,000 unacked messages creates a 10,000-message redelivery spike when it disconnects.

Prefetch has no effect on consumers using auto-ack mode (there is nothing to wait for) or on basic.get pull consumers (each get is an explicit request). If your workload uses either, this knob does not apply.

The two failure modes

flowchart TD
  A[Prefetch count] --> B[Too high or unlimited]
  A --> C[Too low]
  B --> D[One consumer hoards thousands of unacked messages]
  D --> E[RAM pinned, other consumers idle]
  E --> F["Tell: unacked / (ready + unacked) > 0.8"]
  C --> G[Consumer acks, then waits for next delivery]
  G --> H[Throughput capped by ack round-trip]
  H --> I["Tell: consumer_utilisation < 1.0 with backlog"]

Failure mode 1: unacked hoarding

With unlimited or very high prefetch, the broker delivers messages to whichever consumer has capacity as fast as TCP allows. The first consumer to connect, or the one on the fastest network path, drains the queue into its own buffer. The queue looks almost empty in the management UI, but the work is not done: it is sitting, unacknowledged, inside one consumer.

The tell is the ratio messages_unacknowledged / (messages_ready + messages_unacknowledged). Above roughly 0.8, the vast majority of the queue’s work is in flight rather than awaiting delivery. Combined with consumer_utilisation below 1.0 and a low ack rate relative to the deliver rate, you have a hoarding consumer.

A second sanity check: unacked count on a queue should not exceed prefetch_count x active consumer count. If it does, you likely have a dead consumer whose channel never closed cleanly.

The failure escalates in three ways:

  1. Memory pressure. Unacked messages hold RAM on the broker. Sustained hoarding is one of the most common precursors to the memory alarm, at which point every publisher in the cluster is blocked.
  2. Head-of-line blocking. If the hoarding consumer hits a poison message or a slow downstream dependency, the thousands of messages it holds are stuck behind that one problem, even though other consumers are idle and healthy.
  3. Consumer timeout. RabbitMQ’s delivery acknowledgement timeout (consumer_timeout, default 30 minutes) closes the channel of a consumer that holds a delivery too long. A consumer that hoards a large batch and processes it sequentially will eventually trip this, mass-requeueing everything it held. See RabbitMQ consumer timeout: delivery acknowledgement timed out and the channel is closed.

Failure mode 2: idle consumers

With prefetch set very low, especially to 1, the consumer processes a message, sends the ack, and then waits. The broker only sends the next message after the ack arrives, so every message costs at least one network round trip of dead time.

The arithmetic is brutal on high-latency paths. If the round trip between broker and consumer is 125 ms and processing takes 5 ms, a prefetch of 1 means the consumer is busy 5 ms out of every 130 ms: idle about 96% of the time. Adding consumers does not fully compensate, because each one has the same serialization.

The tell here is consumer_utilisation persistently below 1.0 on a queue that has a backlog (messages_ready > 0), combined with a deliver rate well below what the queue could sustain. The queue has work, consumers are attached, and throughput is capped by the ack round trip. See RabbitMQ consumer_utilisation low: consumers attached but not keeping up for the broader diagnostic.

consumer_utilisation on an empty queue is not meaningful. A queue with nothing to deliver can show any utilisation value; only interpret it alongside backlog.

Sizing prefetch: the working method

There is no universal number, but there is a bounded method. You are balancing two constraints.

Throughput floor: cover one round trip. To keep a consumer busy continuously, its prefetch buffer must hold at least as many messages as it can process during one broker-to-consumer round trip plus ack time:

prefetch_min ~= (round_trip_time + processing_time) / processing_time

With a 125 ms round trip and 5 ms processing, that is about 26. With a 1 ms round trip on a local network and 50 ms processing, prefetch of 2 or 3 saturates the consumer. Consumers in the same datacenter as the broker need small values; consumers across regions or the public internet need larger ones.

Memory ceiling: cap by message size. The worst-case memory a single consumer can pin on the broker is prefetch_count x max message size. With 1 MB messages and prefetch 10,000, that is 10 GB held hostage by one consumer. Divide your acceptable per-consumer memory budget by your maximum message size to get a hard ceiling:

prefetch_max ~= memory_budget_per_consumer / max_message_size

Pick a value between the floor and the ceiling. Practical notes:

  • RabbitMQ maintainers have stated that prefetch values above a few hundred make virtually no difference to consumer throughput. Past that point you are buying risk, not speed.
  • Prefetch of 1 is correct when fairness matters more than throughput. If messages have highly variable processing times and you want strict work distribution, prefetch 1 (or a small single-digit value) is the right choice. Accept the throughput cost as the price of even load.
  • If processing time varies wildly within one queue, a large prefetch concentrates the slow tail on whichever consumer grabbed it. Prefer smaller prefetch and more consumers over large prefetch and few consumers in that case.
  • Measure the actual round trip. Do not guess from region names; cross-AZ traffic inside one cloud region can already be meaningful relative to fast processing.

Enforcing a sane default

Because the broker default is unlimited, every client that forgets to set prefetch gets the dangerous behavior. RabbitMQ offers a broker-side default via the default_consumer_prefetch setting, configured in advanced.config as {rabbit, [{default_consumer_prefetch, {false, 250}}]}. This applies a per-consumer limit of 250 to consumers that do not set their own; clients that set prefetch explicitly still override it. It is a useful safety net on shared clusters where you do not control every client library configuration.

Version and queue-type differences that change the answer

  • Global QoS is deprecated. The global flag on basic.qos applies one shared prefetch limit across all consumers on a channel. It was deprecated in RabbitMQ 4.0, and since RabbitMQ 4.3.0 it is denied_by_default: basic.qos with global=true is rejected unless the deprecated feature is re-enabled via deprecated_features.permit.global_qos. Quorum queues never supported global QoS: consuming from a quorum queue on a channel with global QoS returns a channel error. Client libraries that call global QoS during connection autorecovery can fail to recover after a disconnect. Use per-consumer QoS (global=false) exclusively.
  • Quorum queues cap unacked messages per consumer. Even with prefetch set to 0 (unlimited), a quorum queue will not deliver more than 2,000 unacked messages to a single consumer — a cap on consumer prefetch intended to protect against runaway Raft log growth. Operators migrating from classic queues who relied on unlimited prefetch for batch processing will hit this ceiling; the maintainers’ recommendation for that pattern is streams.
  • Classic mirrored queues are gone. They were removed in RabbitMQ 4.0, so any prefetch behavior you observed on mirrored queues no longer applies after migration to quorum queues.
  • Consumer timeout gives hoarders a hard clock. The delivery acknowledgement timeout (consumer_timeout, default 30 minutes) was introduced in RabbitMQ 3.9 and closes the channel of a consumer that holds a delivery too long; since RabbitMQ 4.3 it is enforced for quorum queues only. On 4.3+, clients that advertise the consumer_cancel_notify capability receive a basic.cancel for the timed-out consumer instead of a full channel close. A large prefetch combined with slow per-message processing turns this timeout from a safety net into a recurring channel-churn event.

Signals to watch in production

SignalWhy it mattersWarning sign
messages_unacknowledged per queueDirect measure of in-flight work pinned in RAMGrowing without bound; or steady at prefetch x consumers with a low ack rate
unacked / (ready + unacked) ratioDistinguishes hoarding from backlogSustained above 0.8 means nearly all work is in flight
consumer_utilisation per queueWhether attached consumers can accept deliveriesBelow 1.0 with messages_ready > 0 means consumers are the bottleneck
Ack rate vs deliver rateConfirms processing completesDeliver high, ack near zero: consumers receive but never finish
Deliver rate vs publish rateWhether the queue drainsDeliver capped well below what backlog should allow: prefetch too low
Redeliver rateMass requeues from dead or timed-out consumersSpikes correlated with consumer disconnects or timeout channel closures
mem_used / mem_limitUnacked hoarding ends hereRatio climbing in step with unacked growth

Correlate these, do not read them in isolation. A queue with 50,000 ready messages and utilisation 1.0 is a capacity problem (add consumers). A queue with 200 ready, 40,000 unacked, and utilisation 0.3 is a hoarding problem (fix prefetch and the stuck consumer). Same total depth, opposite fixes.

How Netdata helps

  • Tracks messages_ready and messages_unacknowledged as separate per-queue series, so the hoarding ratio is visible at a glance instead of requiring mental arithmetic on the management UI.
  • Collects per-queue message rates (publish, deliver, ack, redeliver), which is what you need to distinguish “consumers stuck” from “prefetch too low” from “queue simply overloaded”.
  • Correlates queue-level unacked growth with node memory usage on one dashboard, so you can see hoarding turning into memory pressure before the memory alarm fires.
  • Exposes consumer_utilisation per queue alongside depth, making the idle-consumer pattern (backlog plus low utilisation) directly observable.
  • Per-second granularity catches the requeue burst when a hoarding consumer’s channel closes, which slower polling intervals smooth into invisibility.