The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-quorum-queue-minority

Operations Guides

RabbitMQ quorum queue in minority: lost majority and read-only queues

You are looking at a queue with state: minority in the management API, or publishers are timing out on a queue that was healthy an hour ago. A quorum queue enters minority when fewer than a majority of its Raft members are online. Without a majority, the queue cannot elect a leader and cannot commit new writes.

This is not a queue-level bug. It is the direct, by-design consequence of a network partition or the loss of multiple cluster nodes. The queue is protecting your data by refusing writes it cannot replicate safely.

For a queue with active traffic (messages > 0 or consumers > 0), treat this as urgent. Publishers to that queue are stuck, and depending on client behavior they will block, time out, or fail over. This guide walks through confirming the scope, deciding between member recovery and shrinking the member set, and preventing recurrence.

What this means

A quorum queue is replicated across N members using Raft consensus (typically 3 or 5 members). Every write must be committed by a majority: 2 of 3 members, 3 of 5. When the online member count drops below that majority, the queue cannot elect a leader and cannot commit anything new.

In the management API (GET /api/queues), the queue’s state field flips from running to minority. A queue in minority cannot accept writes but may still serve reads of previously committed messages. Treat it as effectively unavailable for new work: publishers will stall, and any client expecting confirms will hang or fail.

The two triggers, in order of likelihood:

  1. Network partition. The cluster split and this queue’s majority ended up on the other side. Other quorum queues with different member placement may still be running, which makes the failure look arbitrary.
  2. Multi-node loss. Two or more nodes hosting this queue’s members are down (crash, hardware failure, botched rolling restart, eviction).
stateDiagram-v2
  running --> minority: majority of members offline (partition or node loss)
  minority --> running: members recover, quorum restored
  minority --> running: dead members removed, member set shrunk
  minority --> unavailable: members permanently lost, quorum unrecoverable
  unavailable --> running: force delete and recreate queue (data loss)

Common causes

CauseWhat it looks likeFirst thing to check
Network partition (split brain)partitions array non-empty on one or more nodes; cluster link traffic dropped to zero between peers; some queues minority, others runningrabbitmq-diagnostics cluster_status
Multi-node failureTwo or more nodes hosting this queue’s members show running: false or do not respond to pingrabbitmq-diagnostics -q ping on each node
Rolling restart took too many members downNodes with very low uptime; members of the same queue restarted back-to-back without waiting for resyncNode uptime via GET /api/nodes plus your maintenance timeline
Permanent node loss (disk, hardware)One or more members offline for an extended period; node will not come backrabbitmq-queues quorum_status <queue> to see which members are offline
Transient false partitionMinority state appeared briefly and self-healed; coincides with GC pauses or CPU saturationErlang run queue on the nodes (GC pauses exceeding net_ticktime, default 60s, cause false node-down detection)

Quick checks

All of these are read-only and safe to run during an incident. Substitute your actual credentials for guest:guest in the API examples.

# Find every queue currently in minority state
curl -s -u guest:guest http://localhost:15672/api/queues | \
  jq '.[] | select(.state=="minority") | {name, vhost, node, state, type, messages: .messages_ready, consumers}'

# Check cluster-wide partition status
rabbitmq-diagnostics cluster_status
# or via API
curl -s -u guest:guest http://localhost:15672/api/nodes | jq '.[] | {name, running, partitions}'

# Per-queue Raft detail: leader, followers, online members
rabbitmq-queues quorum_status <queue_name>

# Member list and which are online
rabbitmqctl list_queues name type online members

# Is each expected node actually alive?
rabbitmq-diagnostics -q ping

# Before any further node shutdown: would it break quorum anywhere else?
rabbitmq-diagnostics check_if_node_is_quorum_critical

Two things to note in the output. First, the activity fields: a queue in minority with messages_ready > 0 or consumers > 0 is actively impacting traffic and should be treated as a page. An idle queue in minority is still a correctness problem but rarely urgent. Second, which specific members are offline. That answer drives the entire recovery decision.

How to diagnose it

  1. Confirm the scope. Run the minority-state query above. One queue in minority points to unlucky member placement. Many queues in minority points to a cluster-level event (partition or multiple node loss). Also check GET /api/vhosts for partially available vhosts if several queues are affected.

  2. Check for a partition first. A non-empty partitions array, corroborated by cluster link traffic dropping to zero between previously communicating peers, confirms a split brain. Verify with independent network checks between nodes (Erlang distribution runs on port 25672 by default). Do not confuse this with a false partition: GC pauses or a saturated Erlang run queue exceeding net_ticktime (default 60s) can trigger transient partition detection that self-heals. Require the condition to persist before acting.

  3. Check node availability. For each node that hosts a member of the affected queue, confirm running status. A node can be running: true but partitioned, so node health and partition health are separate checks.

  4. Classify the offline members: transient or permanent. A node that crashed and will restart cleanly is transient. A node with a dead disk or a decommissioned VM is permanent. This distinction decides the fix: recovery versus shrinking the member set.

  5. Assess downstream impact. Publish rates to the affected queue will be zero or near zero. Look for publisher-side symptoms: connections piling up, timeouts, or retry storms. Consumers may still drain previously committed messages, but plan for the queue to be effectively out of service for new work.

  6. Freeze further node changes. Until quorum is restored, do not stop, restart, or decommission any node that hosts a member of an affected queue. Run rabbitmq-diagnostics check_if_node_is_quorum_critical before touching anything. One more member offline can turn a recoverable minority into permanent loss.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Queue state per queueDirect detection of minority, down, crashedAny state other than running/idle on an active queue
partitions array per nodeMinority almost always follows a partitionNon-empty array sustained across collection intervals
Node running statusDistinguishes dead members from partitioned onesrunning: false on a node hosting queue members
Cluster link peer trafficConfirms real isolation vs. idle clusterBytes dropping to zero between previously connected peers, corroborated by partition detection
Erlang run queueExplains false partitions from heartbeat timeoutsSustained run queue above the scheduler count
Publish rate to the queueMeasures actual impactPublish rate collapses to zero while connection count stays stable
messages_ready / messages_unacknowledged on the queueTracks what is stranded during the outage and what drains after recoveryDepth growing on adjacent healthy queues while the minority queue is frozen

Fixes

Recover the failed members (preferred)

If the offline members are transiently down (crashed node, fixed network partition, restarted VM), bring them back and do nothing to the queue itself. Raft restores quorum automatically once a majority is online again.

After recovery, followers must resynchronize with the leader. On deep queues this resync can take minutes to hours, during which the queue serves traffic while replication catches up. Expect elevated inter-node traffic and disk I/O during this window; it is normal recovery behavior, not a second incident.

If the cause was a partition, resolve the network issue first and let the cluster heal before restarting anything. Premature restarts during an unstable network extend the outage.

Shrink the member set (members permanently lost)

Membership changes are queue state changes: rabbitmq-queues shrink <node> and rabbitmq-queues delete_member <queue> <node> only work while a quorum of the queue’s members is available. If a minority of members is permanently lost (for example 2 of 5) but the remaining majority is online, remove the dead members node by node so the survivors form the new member set; a 5-member queue that permanently lost 2 members can be shrunk to 3 while those 3 are online.

If a majority of members is permanently lost (for example 3 of 5), no quorum exists to approve membership changes, so shrinking is impossible; the queue is permanently unavailable, and the only path is the force-delete-and-recreate procedure in the next section, which loses the queue’s messages.

Shrinking is a deliberate, queue-by-queue operation. Do not attempt it while a partition is still unresolved; you can end up with two sides that each believe they are the queue.

Force delete and recreate (last resort)

If a majority of members is permanently and unrecoverably lost (for example, 2 of 3 nodes destroyed with no backups), the queue cannot be recovered. RabbitMQ’s own documentation states the queue is permanently unavailable and must be force deleted and recreated.

This loses all messages in the queue. Treat it as a data-loss event: confirm with the owning team that the messages are expendable or recoverable upstream, then delete the queue, redeclare it, and let clients reconnect. Normal deletion may fail on a queue with no online replicas; check your version’s documentation for the force-delete procedure.

What not to do

  • Do not restart healthy nodes hosting surviving members. You are one member away from permanent loss.
  • Do not redeploy the queue while a partition is unresolved.
  • Do not wait passively if the queue is active. Minority on an idle queue can wait for business hours; minority on a hot queue is an active outage.

Prevention

  • Use odd member counts and spread them. Three members tolerates 1 failure; five tolerates 2. Place members across failure domains (racks, availability zones) so a single event cannot take a majority.
  • Gate all node maintenance on quorum safety. Run rabbitmq-diagnostics check_if_node_is_quorum_critical (or GET /api/health/checks/node-is-quorum-critical) before every stop, restart, or decommission. During rolling restarts, wait for each node to return and for followers to resync before proceeding.
  • Alert on state: minority with an activity condition. Page when the queue has messages or consumers; ticket otherwise. A bare minority alert on every idle test queue trains people to ignore the real one.
  • Alert on partitions with a sustained condition. Transient false partitions from GC pauses self-heal; real ones do not. Corroborate with cluster link traffic before paging.
  • Rebalance leaders after restarts. Rolling restarts leave leaders concentrated on surviving nodes. rabbitmq-queues rebalance all (a command shipped in current releases; verify it on your version) redistributes them; avoid running it at peak traffic since leader transfer causes brief per-queue unavailability.
  • Keep the run queue healthy. Sustained scheduler saturation causes heartbeat timeouts, which cause false partitions, which cause minority states that were entirely avoidable.

How Netdata helps

  • Per-queue state tracking surfaces the transition from running to minority the moment it happens, with the queue’s message and consumer counts attached so you can judge urgency without a manual query.
  • Partition and node availability signals together let you separate a true split brain (partitions array plus zero cluster link traffic) from a dead node (running: false) and from a false positive (elevated run queue, self-healing).
  • Publish rate collapse alongside stable connection count confirms the queue is refusing writes while clients stay attached, which distinguishes minority from a publisher-side failure.
  • Cluster link traffic rates corroborate isolation between specific peers, telling you which side of the partition holds the majority.
  • Erlang run queue monitoring catches the CPU/GC saturation that causes heartbeat timeouts and false partition detection before it cascades into a real minority event.

Correlating queue state, partition status, and node availability on one timeline turns “some queue is broken” into “nodes 2 and 3 partitioned at 03:14, this queue’s majority was on the other side” in one look.