The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-quorum-loss-no-writes

Operations Guides

ZooKeeper quorum loss: no leader elected and every write is failing

Every write to your ZooKeeper ensemble is timing out. Clients report ConnectionLoss and SessionExpired. Downstream systems that depend on ZK for coordination, such as Kafka controller elections or HBase region assignment, are cascading into failure. On the surviving ZK nodes, ruok still returns imok. The process is alive; the ensemble is not.

Quorum loss is ZooKeeper’s worst-case availability scenario. When fewer than floor(N/2)+1 voting members can communicate, no leader can be elected and every write fails. Surviving nodes sit in LOOKING state, unable to make progress through ZAB.

First response priority: count reachable nodes, check for a network partition, and restore enough members to form quorum. Do not blindly restart survivors. A rolling restart of an already-degraded ensemble can destroy the only in-memory copy of recent transactions.

What this means

Quorum is floor(N/2)+1 voting members. For a 3-node ensemble, quorum is 2; for 5 nodes, 3; for 7 nodes, 4. Observers do not count toward quorum. When the number of reachable, in-sync voting members drops below the threshold, FastLeaderElection cannot complete (FastLeaderElection has been the only election algorithm since ZK 3.6.0). Surviving members stay in LOOKING and ZAB cannot make progress.

User-visible symptoms during confirmed quorum loss:

  • Every write request returns ConnectionLoss or times out.
  • Reads may or may not work, depending on readonlymode.enabled (default false).
  • No node reports leader via mntr or srvr.
  • zk_sum_leader_unavailable_time (3.7+) is non-zero and growing.
  • zk_looking_count may keep incrementing as members retry election.

Suggested page rule: page when no leader exists for more than 60 seconds AND ensemble_size > 1 AND majority of nodes have uptime greater than 600 seconds (this rules out a cold full-cluster start). Split-brain, more than one node reporting leader simultaneously, pages unconditionally because it means data divergence.

Critical nuance: ruok is shallow. It confirms the JVM is alive and the client port is open. A node in LOOKING with no quorum still returns imok. Do not rely on ruok for quorum health. Pair it with isro and mntr.

isro has its own trap. The command returns ro only when readonlymode.enabled=true and the node is actually serving stale reads. With the default readonlymode.enabled=false, surviving nodes do not enter read-only mode at all during quorum loss; isro returns rw even though the node is not serving client requests. Treat isro as a reliable quorum-loss signal only when you have explicitly enabled read-only mode.

flowchart TD
  A[Healthy ensemble
1 leader, N-1 followers] --> B{floor N/2 +1
reachable?} B -- No --> C[Surviving nodes
enter LOOKING] C --> D[No leader elected
all writes fail] B -- Yes --> E[New leader elected
within seconds] D --> F{readonlymode
enabled?} F -- true --> G[isro returns ro
stale reads served] F -- false --> H[Node refuses
client requests]

Common causes

CauseWhat it looks likeFirst thing to check
Network partition between membersSubset of nodes can reach each other; cross-subset TCP failsnc -zv peer 3888 from each node to each peer
Multiple node failuresProcess gone or JVM hung on majority of membersps -f QuorumPeerMain, jstack on each
Firewall or security group changeAll members up, none can reach election portCloud security group audit, iptables -L
Correlated storage failureMulti-node crash tied to shared storage layerdf and iostat on shared volume
DNS resolution failureMembers up but unable to resolve peer hostnamesgetent hosts peer-hostname from each node
Election port blocked asymmetricallyQuorum port open, election port one-directionalBidirectional port probe between every pair
Client-induced thundering herdMass session expiry cascading into leader instabilityzk_stale_sessions_expired, zk_num_alive_connections

Quick checks

Run these on each surviving member. All are read-only.

# Check server state on every node - is any 'leader' present?
echo mntr | nc localhost 2181 | grep zk_server_state

# Read-only mode status. 'ro' only when readonlymode.enabled=true
echo isro | nc localhost 2181

# Shallow liveness - returns 'imok' even during quorum loss
echo ruok | nc localhost 2181

# Server mode from srvr (alternative path if mntr is not whitelisted)
echo srvr | nc localhost 2181 | grep Mode

# Leader-only: how many followers are connected and synced?
echo mntr | nc localhost 2181 | grep -E 'zk_(learners|synced_followers|pending_syncs)'

# Compare zxids across members - divergent zxids signal partition or lag
echo mntr | nc localhost 2181 | grep zk_zxid

# Uptime - gates cold-start false positives
echo mntr | nc localhost 2181 | grep zk_uptime

# Leader election history - increments on each LOOKING entry
echo mntr | nc localhost 2181 | grep zk_looking_count

# Cumulative write unavailability
echo mntr | nc localhost 2181 | grep -E 'zk_.*leader_unavailable_time'

If mntr returns nothing, your 4lw.commands.whitelist probably does not include it. Since ZK 3.5.3, only srvr is whitelisted by default. Add mntr and isro to the whitelist in zoo.cfg, otherwise monitoring silently reports zeros.

Network reachability between members:

# Probe both inter-server ports to each peer
nc -zv peer-hostname 2888
nc -zv peer-hostname 3888

Both ports must be reachable in both directions. A firewall that blocks only the election port will not be noticed until the next leader election, at which point it is catastrophic.

How to diagnose it

  1. Confirm quorum loss. Run mntr against every member. Count nodes that respond. If fewer than floor(N/2)+1 respond, or if every responding node reports no leader for more than 60 seconds, quorum loss is confirmed.
  2. Classify unreachable nodes. For each non-responding member, determine whether the JVM is alive (ps -f QuorumPeerMain), the client port is closed (nc -zv host 2181), or only the election port is unreachable. The classification drives the recovery action.
  3. Check for a network partition. From each surviving member, probe both inter-server ports on every peer. Cloud environments are particularly prone to silent security-group changes that block one direction or one port.
  4. Check uptime to rule out cold start. If majority uptime is under 600 seconds, the ensemble may still be converging after a coordinated restart. Wait two minutes before declaring quorum loss.
  5. Check for split-brain. If more than one node reports leader, treat as a partitioned-brain condition. Page unconditionally and treat as a data integrity risk.
  6. Check leader election churn. zk_looking_count incrementing repeatedly means members are stuck in election loops. Common causes: TCP proxy timing out on election traffic, asymmetric connectivity, or members unable to persist votes due to disk problems.
  7. Check correlated failures. Did all members of one availability zone, rack, or storage tier fail simultaneously? Correlated failures point to infrastructure, not application, root cause.
  8. Check leader-side replication metrics once a leader exists. zk_learners and zk_synced_followers only appear on the leader. If they are absent, no leader has been elected yet.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_server_stateDirect indicator of leader existenceNo node reports leader, or multiple do
zk_looking_countCounts election eventsSustained increments outside maintenance
zk_sum_leader_unavailable_time (3.7+)Cumulative write unavailabilityAny non-zero delta means writes are failing now
zk_uptimeCold-start suppressionReset on majority of nodes signals restart cascade
zk_synced_followers (leader only)Replication healthAt floor(N/2) means one failure from quorum loss
zk_learners (leader only)Connected learner countBelow expected means peer connectivity loss
isroRead-only statero means quorum lost but serving stale reads
zk_proposal_count / zk_commit_countWrite pipeline flowproposal_count advancing while commit_count stalls
zk_outstanding_requestsPipeline backlogGrowing on all nodes signals system-wide stall

Leader-only metrics (zk_learners, zk_synced_followers, zk_pending_syncs) are invisible on followers. Monitoring that queries only one node will never see them. Either identify and query the current leader, or query all members and filter for the leader-reported values.

Fixes

Recovery follows one principle: restore enough voting members to form quorum. How you do that depends on the cause.

Network partition

If a partition splits the ensemble, restore connectivity. Once floor(N/2)+1 members can reach each other on both inter-server ports, FastLeaderElection completes within seconds and writes resume. Verify with mntr showing exactly one leader and isro returning rw on the leader.

Cloud action: audit security groups, network ACLs, route tables. On-prem: check switch fabric, firewall rules, and recent change control. The most common cause in cloud environments is a silent security-group change that blocks inter-server traffic.

Multiple node failures

Restart failed members one at a time, in the order they failed if known. Wait for each to rejoin and sync before restarting the next. Do not restart all members simultaneously; you risk losing the in-memory state on the only node with the most recent committed transactions.

If a member has a corrupt transaction log or snapshot, it cannot rejoin cleanly. Inspect the logs for snapshot errors or transaction-log IOException. You may need to rebuild the member from a healthy snapshot plus transaction log replay.

Dynamic reconfiguration will not rescue you

A frequent operator instinct is to shrink the ensemble via dynamic reconfiguration to exclude failed members. This does not work during quorum loss. Dynamic reconfiguration requires a quorum of the old configuration to make progress. You need quorum to change quorum. If three of five members are permanently gone, dynamic reconfig cannot reduce the ensemble to three. You must either restore enough members or rebuild from a known-good snapshot and transaction log with a reconfigured zoo.cfg.

Read-only mode trap

If readonlymode.enabled=true, surviving nodes serve stale reads while writes fail. This is dangerous because dependent services may appear healthy while losing writes. Kafka leader election, HBase region assignment, and any system using ephemeral nodes for liveness will fail silently. Confirm write recovery explicitly before declaring the incident resolved.

# Verify writes work after recovery - run from a client or zkCli.sh
create /quorum-recovery-test "ok"
get /quorum-recovery-test
delete /quorum-recovery-test

Do not blindly restart survivors

A surviving member in LOOKING state still holds its in-memory data tree and the transaction log on disk. If it is the most-recently-updated survivor, its state is your recovery source. Restarting it discards in-memory state and forces recovery from the on-disk snapshot and logs, which may be older than what was in memory. Always exhaust connectivity and configuration causes before restarting any survivor.

Prevention

  • Monitor zk_server_state across the whole ensemble. Page when no leader exists for more than 60 seconds outside cold-start windows.
  • Alert on zk_looking_count increments outside maintenance. Every unplanned election is an incident worth investigating.
  • Track zk_sum_leader_unavailable_time (3.7+) deltas. Any non-zero delta is a direct measurement of write unavailability.
  • Watch zk_synced_followers on the leader. Alert at floor(N/2), which means one more failure breaks quorum.
  • Probe both inter-server ports in monitoring, not just the client port. Election port reachability is invisible until you need it.
  • Run chaos exercises. Kill one member in production regularly. Validate that monitoring catches it and that the ensemble re-elects within seconds.
  • Use odd-sized ensembles. Even-sized ensembles waste a member without adding fault tolerance, and a half-split partition produces two minorities.
  • Pin the ZK version and track CVEs. Recent releases fix serious quorum-security issues. CVE-2023-44981 (SASL quorum peer auth bypass, fixed in 3.9.1/3.8.3/3.7.2) and CVE-2024-23944 (persistent watcher ACL bypass, fixed in 3.9.2/3.8.4) both affect cluster integrity.
  • Document the recovery procedure before you need it. Quorum loss at 3 a.m. is the wrong time to learn that dynamic reconfig cannot help.

How Netdata helps

  • Per-second zk_server_state collection across every ensemble member surfaces leader absence within seconds, faster than typical 30 to 60 second scrape intervals.
  • zk_sum_leader_unavailable_time deltas directly measure write unavailability, so you alert on impact rather than proxy signals.
  • Leader-only metrics (zk_followers, zk_synced_followers, zk_pending_syncs) are collected from whichever node is currently leader, eliminating the gap left by static monitoring configs that query a fixed node.
  • Correlation with host-level network, disk, and CPU metrics on the same timeline distinguishes quorum loss caused by a network partition from quorum loss caused by disk stalls or GC pauses.
  • ML anomaly detection on zk_looking_count, zk_uptime, and connection counts flags election churn and restart cascades before they become full outages.
  • Cold-start suppression via zk_uptime gating prevents false pages during planned rolling restarts while still paging on real quorum loss once majority uptime crosses the threshold.