The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-run-queue-scheduler-saturation

Operations Guides

RabbitMQ Erlang run queue high: scheduler saturation and slow heartbeats

The RabbitMQ node is alive, queues are accepting messages, but everything feels slow: publishes lag, acks lag, the management API responds sluggishly, and the cluster log starts showing “missed heartbeats” from peers or clients. The node’s run_queue metric sits persistently above the CPU core count. OS-level CPU may even look unremarkable.

This is Erlang scheduler saturation. The BEAM VM has more runnable processes than its scheduler threads can execute, so work queues up and every operation on the node, including heartbeat processing, waits its turn. The secondary effect is worse than the primary one: delayed heartbeats get misread as dead peers, and the cluster starts reporting network partitions that do not exist.

This guide covers how to confirm scheduler saturation, separate it from ordinary CPU load, find the driver, and fix it without making the heartbeat problem worse.

What this means

RabbitMQ runs on the Erlang VM (BEAM). The VM starts a set of scheduler threads, typically one per CPU core, and every connection, channel, and queue is an Erlang process that those schedulers execute cooperatively. When a process is runnable but no scheduler is free, it waits in the run queue.

The run_queue metric (exposed in the Management API per node and via erlang:statistics(run_queue)) counts processes waiting to run, aggregated across all schedulers. Read it against the total core count, not against 1:

  • run_queue at or near 0: schedulers keep up. Spikes during startup, queue purges, or policy application are normal and transient.
  • run_queue sustained above 1 per core: work is queuing. Latency rises for everything.
  • run_queue sustained above total cores: the node is saturated. This is the condition that delays heartbeats.

The heartbeat link is the dangerous part. RabbitMQ connections rely on heartbeat frames to detect dead peers: the default timeout is 60 seconds, frames are sent roughly every timeout/2, and about two missed frames mark the peer unreachable. Cluster nodes additionally use Erlang’s net_ticktime (default 60 seconds) to judge node health. When the schedulers are backlogged, the node may be perfectly healthy but too slow to send or process these frames on time. GC pauses make it worse: a long pause under memory pressure can exceed net_ticktime and produce a transient false partition that appears and heals on its own.

One more trap: OS CPU and Erlang scheduler utilization diverge. Schedulers may be pinned to a subset of cores, the machine may run other workloads, or the scheduler count may not match the actual cores available (a classic container misconfiguration). You can see 40% OS CPU while Erlang schedulers are at 95%. Trust the run queue and scheduler utilization, not top.

flowchart TD
  A[CPU driver: TLS, routing, churn, GC] --> B[Schedulers busy]
  B --> C[run_queue sustained above cores]
  C --> D[All ops slow: publish, ack, mgmt API]
  C --> E[Heartbeat frames delayed]
  E --> F[Missed heartbeat errors in logs]
  E --> G[net_ticktime exceeded: false partition]
  D --> H[Flow control on connections]
  G --> I[Quorum queues lose members, clients reconnect]

Common causes

CauseWhat it looks likeFirst thing to check
High throughput with expensive routingrun_queue tracks publish rate; topic exchanges with many bindingsMessage rates vs run_queue trend; exchange types and binding counts
TLS termination at the brokerHigh CPU with modest message rates; worse during reconnect burstsAre clients on 5671? Connection churn rate
Connection or channel churnScheduler time burned on handshakes and process setup, not messageschurn_rates in /api/overview; connection count oscillation
Fan-out to many queuesOne publish fans out to hundreds of queue processesTopology: bindings per exchange
GC pressurerun_queue high with rising memory; periodic stallsGC counters; binary heap size
Scheduler/core mismatchHigh run_queue but OS CPU lowScheduler count vs actual cores (containers, CPU limits)
CPU contention from neighborsNode shares a host with other heavy workloads; latency spikes without load growthHost-level CPU steal and other processes
Management/stats overheadPeriodic CPU spikes roughly every 5 seconds even when idleManagement plugin stats collection interval

Quick checks

All read-only and safe on a live node.

# 1. Run queue from the Management API (aggregate across schedulers)
curl -s -u guest:guest http://localhost:15672/api/nodes | jq '.[] | {name, run_queue}'

# 2. Same value directly from the VM
rabbitmqctl eval 'erlang:statistics(run_queue).'

# 3. Core count, for the comparison
nproc

# 4. Per-scheduler utilization (this call samples for 5 seconds and blocks)
rabbitmqctl eval 'scheduler:utilization(5000).'

# 5. Where scheduler time goes: emulator, gc, port, sleep
rabbitmq-diagnostics runtime_thread_stats

# 6. Churn: handshakes and channel setup eat scheduler time
curl -s -u guest:guest http://localhost:15672/api/overview | jq '.churn_rates'

# 7. Heartbeat and partition evidence in the log
grep -i "missed heartbeats\|partition" /var/log/rabbitmq/rabbit@$(hostname).log | tail -30

# 8. Cluster view: are partitions actually being reported?
rabbitmq-diagnostics cluster_status

Two notes on interpretation. First, scheduler:utilization(5000) blocks the caller for the sample window; that is expected, not a hang. The scheduler module (with scheduler:utilization/1) has shipped in Erlang/OTP since OTP 21, so it is available on every supported RabbitMQ (3.13 and 4.x mandate OTP 26 or newer). Second, if the management API itself is slow to answer check 1, that is consistent with saturation, not a separate API problem.

How to diagnose it

  1. Confirm saturation, not a spike. Sample run_queue every few seconds for a couple of minutes. Startup, queue purges, and post-restart sync legitimately spike it. You care about sustained elevation above the core count.

  2. Compare scheduler utilization to OS CPU. Run scheduler:utilization(5000) and watch OS CPU in parallel. Schedulers hot with OS CPU low points to a scheduler count mismatch or pinning problem. Both hot points to genuine CPU demand.

  3. Break down where scheduler time goes. rabbitmq-diagnostics runtime_thread_stats splits time between emulator (executing code), gc, port (I/O), and sleep. High gc share points at memory pressure. High emulator with high publish rates points at routing and serialization. High port points at I/O-heavy work.

  4. Identify the driver. Correlate the saturation window with message rates (message_stats in /api/overview), churn rates, and connection counts. Saturation that tracks publish rate is throughput-driven. Saturation with flat message rates but high churn is handshake-driven. Saturation with neither points at GC, stats collection, or a hot internal process.

  5. Check GC and memory pressure. Take two samples of rabbitmqctl eval 'erlang:statistics(garbage_collection).' a few seconds apart and compare the deltas. Check rabbitmq-diagnostics memory_breakdown for a bloated binary heap. As a diagnostic only, rabbitmq-diagnostics force_gc forces GC across all processes: if run_queue and memory drop right after, lazy GC was contributing. force_gc itself is a CPU burst; do not run it repeatedly on a saturated node.

  6. Quantify the heartbeat damage. Grep the log for “missed heartbeats” and check cluster_status for partitions. If partitions appear and self-heal, correlate their timing with run_queue peaks and GC activity before treating them as network events.

  7. Rule out noisy neighbors. If the node shares a host or VM with other workloads, check host-level CPU. The Erlang runtime assumes it does not share CPU; time-slicing against other tenants inflates latency far beyond what the load suggests.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
run_queue (per node)Direct CPU saturation measure, aggregate across schedulersSustained above total core count
Scheduler utilizationErlang-specific busy time; catches saturation OS CPU missesAverage above 80% for 5+ minutes; above 95% risks unresponsiveness
OS CPU per coreDivergence from scheduler utilization reveals misconfigurationOS low, schedulers high, or one core pinned at 100%
churn_rates.connection_created / channel_createdHandshakes and process setup are scheduler-intensiveSustained well above the rolling baseline
Message rates (publish, deliver_get, ack)Ties saturation to real workload vs overheadSaturation growing faster than throughput
GC countersGC steals scheduler time and pauses processesGC share of scheduler time climbing; long pauses
partitions array / net_ticktime eventsThe cliff-edge secondary effect of saturationAny non-empty partitions array, especially transient self-healing ones
mem_used / mem_limitMemory pressure drives GC, which drives scheduler loadRatio above 0.7 during saturation events
Connections in flow stateSaturation backpressures publishers through credit flowMany connections in flow sustained

Fixes

Reduce CPU per message

If saturation tracks publish rate, the routing path is too expensive. Topic exchanges with thousands of bindings make every publish match routing keys against patterns. Large fan-outs turn one publish into work for hundreds of queue processes. Simplify routing topology, consolidate bindings, and split hot exchanges. Check message sizes too: serialization cost scales with payload.

Move TLS termination off the broker

TLS at the broker is one of the most common saturation drivers, and reconnect bursts multiply it because every handshake is CPU-heavy. Terminating TLS at a load balancer or proxy in front of the broker removes that cost, at the tradeoff of plaintext between the proxy and broker, which may be unacceptable in your threat model. If TLS must stay on the broker, reducing connection churn (below) matters even more.

Fix connection and channel churn

Connection setup involves the TLS handshake, AMQP negotiation, and Erlang process creation; channels are each a process too. Clients that open a connection per operation, reconnect without backoff, or use a channel-per-message pattern will saturate schedulers while doing almost no messaging. Enforce connection pooling, reuse channels, and add reconnect backoff. See RabbitMQ connection storm: reconnect loops, FD pressure, and CPU spent on handshakes and RabbitMQ channel leak and churn: the channel-per-message anti-pattern.

Fix scheduler and core mismatches

If scheduler utilization is high but OS CPU is low, check how many schedulers the VM started versus how many cores the node can actually use. In containers, a node started with the host core count under a 2-core CPU quota will thrash. Verify the container CPU limit and that the runtime sees it.

Relieve GC pressure

Saturation with a high gc share in runtime_thread_stats and a rising memory ratio means the VM is spending scheduler time collecting. Find the memory driver first: queue depth, unacked messages, or a bloated binary heap (rabbitmqctl eval 'erlang:memory(binary).'). Draining the memory pressure removes the GC load with it. Do not treat force_gc as a remediation loop; it is a diagnostic.

Rebalance and scale

Quorum queue leaders do all the publish and dispatch work for their queues, so a node holding most leaders works much harder than its peers. If one node saturates while others idle, check leader distribution and rebalance. If the workload genuinely exceeds one node’s cores, add capacity or move queues; no tuning makes a saturated scheduler faster.

Do not fix heartbeats by tightening timeouts

Lowering the heartbeat timeout on a saturated node makes false detections more frequent, not fewer. Values in the low single digits produce false positives even on a healthy node under burst load. Keep the default 60 seconds or a moderate value, and fix the saturation instead. The same applies to net_ticktime: raising it is sometimes justified on lossy links, but here it is a mask, not a fix.

Prevention

  • Alert on the ratio, not the absolute. Ticket (do not page) when run_queue exceeds total cores sustained for several minutes, and when average scheduler utilization stays above 80%. Both are leading indicators well before heartbeats start failing.
  • Trend churn separately from connection count. A stable connection count can hide heavy open/close churn that is quietly burning scheduler time.
  • Keep CPU headroom. Plan capacity so peak load leaves schedulers below roughly 60% average. Past about 80%, small load increases produce disproportionate latency.
  • Do not colocate. Run RabbitMQ nodes on hosts, or with CPU guarantees, that do not share cores with other heavy workloads.
  • Correlate partitions with run_queue before declaring network incidents. A transient partition that coincides with a run_queue spike or GC pause is a symptom, not a network fault. Require partitions to persist before paging.
  • Watch the write path after fixes. After reducing churn or topology cost, confirm run_queue drops and connections leave flow state; see RabbitMQ connection in flow state: credit-based backpressure explained for what that state means.

How Netdata helps

  • Saturation over time: charting the node’s run_queue over minutes and hours shows sustained saturation instead of point samples during an incident.
  • Scheduler vs OS CPU correlation: charting per-core OS CPU from Netdata next to the node’s scheduler utilization and run_queue makes the divergence pattern (saturated VM, idle host) visible immediately, which is the fastest route to a misconfiguration fix.
  • Churn correlation: Netdata’s RabbitMQ connection and channel churn charts on the same view as CPU let you see whether scheduler time is going to handshakes or to messaging.
  • Throughput overlay: publish, deliver, and ack rates against CPU and run_queue separate workload-driven saturation from overhead-driven saturation.
  • Timeline correlation: partition events and resource alarms on the same timeline as run_queue spikes let you confirm false partitions instead of chasing phantom network faults.