The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-publish-deliver-imbalance

Operations Guides

RabbitMQ publish outpacing deliver: reading the rate imbalance before the backlog

When publish rate runs ahead of deliver rate on a RabbitMQ broker, nothing is broken yet. That is exactly why this signal matters. Queue depth is a lagging indicator: by the time messages_ready looks alarming, the imbalance has usually been running for minutes. The rate comparison tells you the same thing earlier, and it tells you which side of the pipe changed.

The comparison has one trap that produces both false alarms and missed incidents: the correct counter to compare against publish is deliver_get, not deliver. Pull-mode consumers using basic.get only show up in get, and deliver_get is the sum of deliver plus get. Compare publish against deliver alone on a workload with pull consumers and you will see a permanent fake imbalance.

This article covers how to read the publish versus deliver_get rates correctly, what the different imbalance shapes mean, and which signals to correlate before you act. For the broader broker mental model and failure pattern catalogue, see How RabbitMQ actually works in production.

What this means

RabbitMQ exposes message rates as cumulative counters: publish, deliver, get, deliver_get, ack, redeliver, and others, available cluster-wide via GET /api/overview, per vhost via GET /api/vhosts, and per queue via GET /api/queues/{vhost}/{name} in each object’s message_stats block. Your monitoring system differentiates the counters into rates; the management API also ships pre-computed *_details.rate fields from a sliding window.

Three rate relationships carry almost all of the diagnostic value:

  • publish vs deliver_get: the queue growth equation. If publish exceeds deliver_get sustained, depth will grow. This is the leading read; queue depth is the confirmation.
  • deliver vs ack: the consumer health equation. If deliver is positive but ack is zero, consumers are receiving messages and not completing them. That is a different failure than consumers being absent.
  • publish dropping to zero: publisher-side failure or publisher-side blocking. A rate collapse is not a capacity problem; it is publishers dead, blocked by a resource alarm, or throttled by flow control.

Batch publishers are spiky by design. A job that pushes 50,000 messages in ten seconds and then goes quiet will show publish far above deliver_get on any short sample. Use 1-minute averages for the comparison, not instantaneous samples, and judge imbalance over 5 to 10 minutes, not seconds.

flowchart LR
  P[publish rate] --> Q[queue process]
  Q --> D[deliver_get rate]
  D --> A[ack rate]
  P -- "publish > deliver_get" --> G[depth grows]
  G --> M[memory rises]
  M --> PG[paging to disk]
  PG --> F[flow control on publishers]
  M --> AL[memory alarm: all publishers blocked]
  D -- "deliver > 0, ack = 0" --> U[unacked builds: RAM pinned]

Common causes

CauseWhat it looks likeFirst thing to check
Consumers absent or crashedpublish > 0, deliver_get near 0, consumers = 0 on the queuePer-queue consumer count
Consumers present but too slowpublish > deliver_get sustained, deliver_get > 0, depth growingconsumer_utilisation and ack rate on the queue
Consumers connected but not ackingdeliver > 0, ack = 0, messages_unacknowledged pinned at prefetchUnacked count vs prefetch times consumer count
Pull-consumer miscountpublish looks above deliver, but deliver_get matches publishCompare against deliver_get, not deliver
Traffic spike beyond capacitypublish 2-3x baseline, deliver_get at its ceiling, both healthyRolling baseline comparison, 1-minute averages
Publishers blocked or failedpublish drops to zero or near zeromem_alarm, disk_free_alarm, connection states
Silent routing losspublish > 0, deliver_get = 0, depth flat and emptyExchange bindings and return_unroutable
Fanout multiplierdeliver_get higher than publishExpected: one publish routed to N queues counts N deliveries

Quick checks

All read-only. Substitute your management user and vhost (the default vhost / is URL-encoded as %2f).

# Cluster-wide message counters and rates
curl -s -u guest:guest http://localhost:15672/api/overview | jq '.message_stats'

# Global queue totals: ready vs unacknowledged
curl -s -u guest:guest http://localhost:15672/api/overview | jq '.queue_totals'

# Per-queue depth and consumer state in one pass
rabbitmqctl list_queues name messages_ready messages_unacknowledged consumers consumer_utilisation

# Per-queue rates: message_stats for a specific queue
curl -s -u guest:guest 'http://localhost:15672/api/queues/%2f/my_queue' | \
  jq '.message_stats | {publish: .publish_details.rate, deliver_get: .deliver_get_details.rate, ack: .ack_details.rate}'

# Consumer effectiveness on the same queue
curl -s -u guest:guest 'http://localhost:15672/api/queues/%2f/my_queue' | \
  jq '{consumer_utilisation, consumers, head_message_timestamp}'

# If publish collapsed: check resource alarms first
curl -s -u guest:guest http://localhost:15672/api/nodes | \
  jq '.[] | {name, mem_alarm, disk_free_alarm}'

# If publish collapsed: count connection states
curl -s -u guest:guest http://localhost:15672/api/connections | \
  jq 'group_by(.state) | map({state: .[0].state, count: length})'

Two notes on collection cost. Iterating /api/connections to count states is expensive on deployments with many connections; sample it, do not poll it tightly. And the management API itself has overhead: polling more frequently than every 5-10 seconds can measurably load the broker.

How to diagnose it

  1. Confirm the imbalance is real. Pull message_stats from /api/overview twice, 60 seconds apart, and compare 1-minute average rates. A single spiky sample from a batch publisher is not an imbalance. Sustained publish above deliver_get for 5+ minutes is.

  2. Localize it. Cluster-wide rates can hide a single hot queue. List per-queue depths and consumers with rabbitmqctl list_queues, then pull message_stats per queue from the HTTP API and sort by publish rate. Per-queue rates can be zero while global rates look healthy because load is concentrated elsewhere.

  3. Check which side moved. If publish jumped above baseline, you have an ingress event: a flood, a retry storm, a new producer. If publish is at baseline and deliver_get fell, you have a consumer-side problem. The fix paths are completely different.

  4. Characterize the consumer side. On the affected queue, read consumers, consumer_utilisation, and the deliver versus ack relationship. Zero consumers means nobody is subscribed. Utilisation below 0.5 with a growing backlog means consumers are attached but not effective. Deliver positive with ack zero means the consumer black hole: connected, receiving, never completing. See consumers connected but not acknowledging.

  5. Check depth as confirmation, not discovery. messages_ready should now be growing at roughly (publish - deliver_get) per second. If publish exceeds deliver_get but depth is flat and empty, suspect routing loss: check exchange bindings and return_unroutable. Messages published without mandatory=true to an exchange with no matching bindings are silently discarded and appear in no counter.

  6. Project the runway. For classic queues, the imbalance rate times average message size approximates memory growth while messages sit in RAM. Compare against mem_used / mem_limit: paging begins at the paging ratio (default 0.5 of the watermark) and the cluster-wide publisher block fires at ratio 1.0. That gives you a time-to-alarm estimate instead of a surprise.

  7. If publish dropped to zero, switch playbooks. Zero publish with connections still open means publishers are blocked, not slow. Check mem_alarm and disk_free_alarm first (memory resource limit alarm, disk free limit alarm), then connection states. Connections in blocked or blocking mean a resource alarm; connections in flow mean per-connection credit throttling, which is surgical and transient rather than cluster-wide. The distinction matters: flow control vs resource alarms.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
publish rate (1-min avg)Ingress ground truthDrop >80% from rolling baseline for 5+ min: publisher failure or block
deliver_get rate (1-min avg)Total egress, push plus pullSustained gap vs publish: depth will grow
publish / deliver_get ratioThe imbalance itself>2x sustained for 10+ min
deliver vs ack rateConsumer completiondeliver > 0 with ack = 0: stuck consumers
messages_ready trendConfirmation of the rate mathPositive growth for 10+ min
messages_unacknowledgedRAM pinned by in-flight messagesPinned at prefetch x consumers, no acks
consumer_utilisationConsumer effectiveness, not presence< 0.5 with messages_ready > 0
redeliver rateConsumer failure loopsElevated: nacks, timeouts, poison messages
head_message_timestampLatency depth cannot showOldest message aging past queue SLA
return_unroutableRouting loss (mandatory=true only)Any sustained rate > 0

Fixes

Consumers absent or crashed. Restore the consumer fleet first. Depth will drain on its own once deliver_get exceeds publish. Do not restart the broker: it does not fix a consumer problem and it forces queue recovery on top of an active backlog.

Consumers too slow. If consumer_utilisation is low, the bottleneck is consumer-side processing or prefetch tuning, not consumer count; adding consumers that all wait on the same slow downstream changes nothing. If utilisation is high and depth still grows, you genuinely need more consumer capacity. See consumer_utilisation low.

Consumers not acking. This is the black hole pattern. Check consumer application health and downstream dependencies, and consider the consumer timeout behavior introduced in 3.8 (15-minute default, raised to 30 minutes in 3.8.17), which closes channels whose deliveries go unacknowledged past the timeout (default 30 minutes). Unacked messages are pinned in RAM and cannot be paged to disk, so this pattern reaches the memory alarm faster than ordinary backlog.

Publish-side flood. If the ingress spike is legitimate (batch backfill, replay), the options are consumer scale-up, or letting the backlog drain if your message-age SLA tolerates it. Quorum and lazy-queue depth decouples backlog from RAM more than classic queues do, which buys time but not infinite time.

Publish rate zero. Do not tune consumers. Find the block: memory alarm, disk alarm, or widespread flow state. Rate imbalance analysis resumes after publishers can publish.

Silent routing loss. Fix the binding, and set mandatory=true on publishers so unroutable messages surface in return_unroutable instead of vanishing. This is the one imbalance shape where depth never confirms the problem, because the messages never arrive.

Prevention

  • Alert on the ratio, not just depth. Sustained publish / deliver_get above 2x for 10+ minutes catches capacity mismatches while depth is still boring. Depth alerts alone fire late.
  • Alert on publish collapse. A drop of more than 80% from rolling baseline for 5+ minutes catches publisher failure and alarm-driven blocks before users report missing work.
  • Baseline per queue. Per-queue rates can be zero while global rates are healthy. The queues that matter need their own publish and deliver_get baselines.
  • Pair depth with age. Depth tells you volume; head message age tells you latency. The rate imbalance predicts both.
  • Watch unacked separately. messages_unacknowledged is the fast path to the memory wall and the clearest stuck-consumer signal.
  • Respect the batch pattern. If your publishers are batch jobs, alert on 1-minute averages and longer sustained windows, or you will page yourself every time a job runs.

How Netdata helps

  • Netdata collects RabbitMQ message_stats counters and differentiates them into publish, deliver_get, and ack rates, so the imbalance is visible as a rate comparison rather than something you compute by hand.
  • Per-queue charts for ready, unacknowledged, and message rates let you localize a cluster-wide imbalance to the one queue actually driving it.
  • Correlating the rate gap against memory usage and mem_used / mem_limit on the same dashboard turns “depth will grow” into a concrete time-to-alarm estimate.
  • Unacked messages and consumer counts alongside ack rate make the black hole pattern (deliveries without acknowledgements) obvious in one view.
  • Alarm flags and connection states on the same timeline explain publish-rate collapses immediately: blocked connections plus an active alarm is a resource problem, not a publisher bug.