The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-node-quorum-critical

Operations Guides

RabbitMQ node is quorum critical: checking before a rolling restart

You are about to restart a RabbitMQ node: a rolling upgrade, an OS patch, an instance resize. Before you stop it, one question decides whether this is routine maintenance or a queue outage: if this node goes down right now, does any quorum queue or stream lose its online majority?

RabbitMQ ships a purpose-built check for exactly this. rabbitmq-diagnostics check_if_node_is_quorum_critical returns unhealthy when stopping the target node would drop a quorum queue below the number of online members it needs to accept writes. The same logic is exposed over HTTP at GET /api/health/checks/node-is-quorum-critical, which makes it usable from automation, load balancer health gates, and CI-driven upgrade pipelines.

This guide covers what the check evaluates, how to run it from the CLI and the API, how to sequence a rolling restart around it, and the edge cases where it can mislead you.

What the quorum-critical check does

Quorum queues replicate each message across a group of members using Raft consensus. A queue with 3 members needs 2 online members (a majority) to accept writes and elect a leader. A queue with 5 members needs 3. When a queue falls below majority, it shows state: minority in the management API: it may still serve already-committed messages to consumers, but it cannot accept new publishes.

The check inverts that logic from the operator’s perspective. Instead of asking “is any queue currently in minority?”, it asks “if I stop this node, would any queue end up in minority?” For every quorum queue and stream with a member on the target node, it looks at the currently online members and computes whether removing this node leaves fewer than a majority online.

The critical detail: the check cares about online members, not configured members. With 3 members configured and 1 already down, only 2 are online. Taking a second node down leaves 1 of 3 online, below majority. The queue goes unavailable. The check on the second node returns unhealthy precisely because of this.

flowchart TD
    A[Node restart requested] --> B{Run quorum-critical check}
    B -->|healthy, exit 0| C[Drain and stop node]
    C --> D[Restart node]
    D --> E{All members back online?}
    E -->|yes| F[Proceed to next node]
    E -->|no| G[Do not continue. Investigate the lagging member]
    B -->|unhealthy, exit 69| H[Do not restart. A queue would lose majority]
    H --> I[Find which queue members are offline]
    I --> J[Restore the offline member first]
    J --> B

Why this is the gate that matters during rolling restarts

Rolling restarts fail in a predictable way with quorum queues:

  1. Node 1 goes down for patching. Quorum queues with a member on node 1 now run at 2 of 3 online. Still accepting writes. Leader elections happen for queues led by node 1. All normal.
  2. The operator, or the automation, does not wait for node 1 to fully return and rejoin its queue groups. Node 2 goes down.
  3. Every queue with members on both node 1 and node 2 is now at 1 of 3 online. Minority. Publishing to those queues stops. Consumers keep draining committed messages, but producers block or fail.

The trap is that a quorum queue with one member down looks like it is working fine, because commits still succeed. Fault tolerance is gone, and one more failure makes the queue unavailable. The check surfaces that hidden fragility before you act on it.

Two behaviors during restarts are worth expecting rather than fighting:

  • Leader elections are normal. Rolling restarts cause Raft leader elections as members leave and return. Do not alarm on leadership changes during the maintenance window.
  • Followers need resync time. After a node restarts, its quorum queue followers resynchronize with the leader. Depending on queue depth this can take minutes to hours. The node being “up” does not mean its queue members are caught up. Wait for full membership before moving to the next node.

How to run the check

CLI

# Quorum-critical gate before restarting this node
rabbitmq-diagnostics check_if_node_is_quorum_critical

The command is also available as rabbitmq-queues check_if_node_is_quorum_critical.

Exit behavior:

  • Exit 0: stopping this node leaves every quorum queue and stream with an online majority. Safe to proceed.
  • Exit 69: stopping this node would leave at least one quorum queue or stream without an online majority. Do not proceed.
  • Other non-zero codes: the check itself failed (for example, it could not reach the node). Treat this as “unknown”, not as “safe”.

The exit code distinction matters for automation. The Kubernetes Cluster Operator’s preStop hook loops on this command and blocks pod termination only on exit code 69; other failures are not treated as quorum-critical. Mirror that behavior in your own scripts: gate on 69 specifically, and fail closed (do not restart) on anything you cannot classify.

HTTP API

# Quorum-critical gate over the management API
curl -s -o /dev/null -w "%{http_code}\n" \
  -u "$USER:$PASS" \
  http://localhost:15672/api/health/checks/node-is-quorum-critical

A 200 response means the node is not quorum critical. A failure status means stopping it would break a majority somewhere. This endpoint is the right choice for load balancer health gates and external orchestration that should not shell into the node.

There is also a management CLI equivalent, rabbitmqadmin health_check node_is_quorum_critical.

Scope note

The check covers both quorum queues and streams in current releases. Classic mirrored queues are a separate concern (check_if_node_is_mirror_sync_critical), and mirroring was removed entirely in RabbitMQ 4.0, so on current versions the quorum check is the only gate of this kind you need.

Interpreting the result

ResultMeaningAction
Healthy (exit 0 / HTTP 200)Every quorum queue and stream keeps an online majority without this nodeProceed with drain and restart
Unhealthy (exit 69 / HTTP failure)At least one queue or stream drops below majority if this node stopsStop. Find the offline members and restore them first
Check errors (other exit code)Check could not completeTreat as unknown; investigate before restarting

When the check says unhealthy, find out why before doing anything else:

# Which quorum queues exist, where are their members, and who is online
rabbitmqctl list_queues name type state members online

Look for queues where the online member count is lower than the members count. That gap is your exposure. Common reasons a member is offline:

  • Another node is down, intentionally or not. Check cluster status and node uptime across the cluster.
  • A node recently restarted and its followers are still resyncing. Wait for resync to complete.
  • A queue was declared with members on nodes that no longer exist or were renamed.

The fix is almost always “restore the missing member first, then re-run the check”. Do not work around the gate by stopping the node anyway: the check is telling you a specific queue will stop accepting writes the moment you do.

Sequencing a rolling restart safely

For a 3-node cluster, the safe sequence is:

  1. Baseline. Confirm all nodes are running, no partitions are detected, and every quorum queue shows its full member count online.
  2. Gate node 1. Run the quorum-critical check on node 1. Only proceed on a healthy result.
  3. Drain and stop node 1. Put it into maintenance mode if your workflow uses it, stop the node, do the work, start it again.
  4. Wait for full rejoin. Do not move on when the node process is up. Wait until its quorum queue members are online and caught up, leader elections have settled, and resync has finished. rabbitmq-upgrade await_online_quorum_plus_one exists for exactly this wait in automated flows: it blocks until there are enough nodes online to maintain quorum plus one.
  5. Re-verify the cluster. Confirm list_queues name type state members online shows full membership everywhere and no queue is in minority.
  6. Gate node 2. Run the check on node 2. Repeat.

The discipline is the same in Kubernetes, where the Cluster Operator already implements most of it: the preStop hook blocks pod termination while the node is quorum critical, and the operator surfaces a quorumStatus field on the RabbitmqCluster resource so you can see the state externally. If you run RabbitMQ outside the operator, you own this sequencing yourself, and the check is the primitive to build it on.

One cadence note: environments with very high queue churn (tens of queue declarations and deletions per second) are harder to restart cleanly, because queue groups are forming and dissolving while members are offline. If your workload churns queues aggressively, schedule restarts in the quietest window you have and re-run the check immediately before each stop, not at the start of the whole procedure.

Known false positives and blind spots

The check is the right tool, but it is not perfect. Know its failure modes before you build a hard gate on it:

  • False positive after recent replica changes (fixed in 3.13.4). The check could report a node as quorum critical when some quorum queue replicas had been very recently added or very recently restarted. If you are on an older 3.13 release and the check disagrees with what list_queues name type state members online shows, upgrading past 3.13.4 resolves it.
  • False positive during maintenance mode (fixed in 3.8.10). A node drained via maintenance mode could leave a stale process entry that made the check report quorum critical incorrectly. Only relevant on very old versions.
  • False negative with a dead follower. In a 3-node cluster with a 3-member queue where one follower process has crashed (shown as “noproc”), the check historically did not flag the leader as quorum critical, even though stopping the leader would leave no quorum. A fix was merged (PR #12727) and has shipped in stable releases: RabbitMQ 4.0.4 and 4.1.0. Until you can verify the fix is in your running version, do not rely on the check alone: verify member online status directly with list_queues name type state members online before each stop.
  • The check is point-in-time. A healthy result at 09:00 says nothing about 09:05 if another node fails in between. Re-run the check immediately before each stop, and never run two stops concurrently “because both checked out earlier”.

Signals to watch during the restart

SignalWhy it mattersWarning sign
Quorum-critical check resultThe gate itselfExit 69 or HTTP failure before any stop
Queue state per queueminority means a queue already lost majorityAny queue in minority outside a planned stop
Members vs online membersShows degraded redundancy before it becomes an outageOnline count below member count after the node should have rejoined
Node uptimeConfirms which node restarted and whenUptime reset on a node you did not touch
Partition statusA partitioned node looks running but cannot contribute to quorumNon-empty partitions array
Publish and ack ratesConfirms traffic resumed after each node returnsPublish rate not recovering after a node rejoins
Cluster link trafficQuorum replication flows over inter-node linksZero traffic between peers, corroborated by partition state

How Netdata helps

  • Netdata collects per-node availability, uptime, and partition status, so you can see exactly when each node left and returned during the rolling restart, and catch a node that came back but partitioned.
  • Per-queue state and message metrics let you spot a queue stuck in minority or not recovering after its member node rejoined, which is the failure mode the quorum-critical check is designed to prevent.
  • Publish, deliver, and ack rate charts give you the “did traffic actually resume” confirmation after each node returns, before you gate the next one.
  • Correlating uptime resets with queue state changes on one timeline makes it obvious whether a minority event lines up with your maintenance window or with an unplanned failure that should stop the rollout.