The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-leader-election-storm

Operations Guides

ZooKeeper leader election storm: an ensemble that keeps re-electing

zk_looking_count is incrementing on multiple nodes. The ZooKeeper log fills with repeated LEADING and FOLLOWING transitions. Every few seconds or minutes, the ensemble elects a new leader, and that leader quickly loses quorum. Writes are intermittent, and downstream systems that depend on ZooKeeper for coordination (Kafka controllers, HBase region assignment, distributed locks) experience cascading failures.

This is a leader election storm. Unlike a single failover where remaining nodes elect a stable replacement, here every potential leader hits the same wall. The root cause is shared across ensemble members: fsync stalls on overloaded storage, GC pauses exceeding the quorum timeout, or intermittent network failures between server pairs. Because every candidate experiences the same problem, no leader holds the role long enough for the cluster to recover.

More than one unplanned election per hour is instability. Each election is an availability event. During the LOOKING phase, no writes are processed, client sessions approach expiry, and downstream coordination systems react to the gap.

What this means

In a healthy ensemble, elections happen during planned maintenance or single-node failure. The remaining nodes elect a new leader, the cluster converges, and writes resume within seconds, inside the initLimit x tickTime window.

An election storm is a different pattern. The ensemble cycles through elections without stabilizing. The mechanism is a feedback loop: a shared resource problem causes the current leader to become unreachable. Followers exceed the syncLimit x tickTime window (default: 5 x 2000ms = 10 seconds), declare the leader dead, and start a new election. The new leader is elected, but it runs on the same hardware with the same overloaded disk or the same JVM heap pressure. Within seconds, it too misses heartbeats. The cycle repeats.

flowchart TD
    A["Shared root cause"] --> B["Leader misses heartbeats"]
    B --> C["Followers declare leader dead"]
    C --> D["Followers enter LOOKING"]
    D --> E["New leader elected"]
    E --> F["Same problem recurs"]
    F --> B

Key diagnostic insight: because the root cause is shared, every node is affected. If only one node were the problem, the ensemble would elect a different leader and stabilize. The storm pattern means you need to look for a common factor across all members: the same storage tier, the same heap configuration, the same network path.

Common causes

CauseWhat it looks likeFirst thing to check
Fsync stalls on txnlog diskzk_max_fsynctime elevated on leader, write latency spikes precede each election, OS iowait highecho mntr | nc localhost 2181 | grep zk_.*fsynctime
GC pauses exceeding syncLimit x tickTimezk_p99_jvm_pause_time_ms (3.7+) spikes align with election events, latency spikes are rhythmicecho mntr | nc localhost 2181 | grep zk_.*jvm_pause
Intermittent network between pairsElections trigger without disk or GC correlation, zk_p99_quorum_ack_latency spikes, specific followers repeatedly dropecho mntr | nc localhost 2181 | grep zk_.*quorum_ack_latency
maxTimeToWaitForEpoch too low (3.6+)Leader abandons leadership on stray LOOKING notifications without heartbeat missesCheck JVM property zookeeper.leader.maxTimeToWaitForEpoch
Version-specific election bug (3.5.7 and earlier)3-node ensemble re-elects after leader shutdown, connectOne() to dead node delays vote delivery by 5 secondsCheck ZooKeeper version against ZOOKEEPER-2164 fix (3.5.8+, 3.6.1+, 3.7.0+)

Note: the percentile metrics shown above require ZooKeeper 3.6+ (jvm_pause_time and leader_unavailable_time only from 3.7+), and mntr must be listed in 4lw.commands.whitelist (default since 3.5.3: only srvr).

Quick checks

Run these on each ensemble member. They are read-only and safe for production.

# Current role - any node reporting other than leader/follower is in trouble
echo mntr | nc localhost 2181 | grep zk_server_state

# Election count since process start
echo mntr | nc localhost 2181 | grep zk_looking_count

# Cumulative leader unavailability (sum = total ms, cnt = episode count)
echo mntr | nc localhost 2181 | grep zk_.*leader_unavailable_time

# Fsync latency percentiles - smoking gun for disk problems
echo mntr | nc localhost 2181 | grep zk_.*fsynctime

# JVM pause time percentiles - smoking gun for GC problems
echo mntr | nc localhost 2181 | grep zk_.*jvm_pause

# Quorum ACK latency (leader only) - network and follower processing
echo mntr | nc localhost 2181 | grep zk_.*quorum_ack_latency

# Synced followers and pending syncs (leader only)
echo mntr | nc localhost 2181 | grep -E "synced_followers|pending_syncs"

# Recent fsync warnings from the log
grep "fsync-ing the write ahead log" /var/log/zookeeper/zookeeper.log | tail -20

# Election transitions in the log
grep -E "LEADING|FOLLOWING|LOOKING" /var/log/zookeeper/zookeeper.log | tail -30

# OS-level disk latency on the transaction log device
iostat -x 1 5

How to diagnose it

  1. Confirm it is a storm, not a single election. Check zk_looking_count on all members. If only one node incremented, it may have restarted. If multiple nodes show increments within the same window, the ensemble is cycling.

  2. Correlate election timestamps with fsync latency. Pull zk_max_fsynctime (3.6+; mntr exposes fsynctime only as avg/min/max/cnt/sum) alongside election events from the log. If fsync spikes immediately precede each LOOKING transition, the disk is the trigger. The fsync warning log line fires when fsync exceeds fsync.warningthresholdms (default 1000ms). Any appearance of that message during a storm is a direct signal.

  3. Correlate election timestamps with JVM pause time. Pull zk_p99_jvm_pause_time_ms (3.7+, requires jvm.pause.monitor=true) alongside election events. GC pauses approaching syncLimit x tickTime (10 seconds with defaults) will cause followers to declare the leader dead. Check the GC log for Full GC events aligned with each election.

  4. Check for network-specific patterns. If neither fsync nor GC correlates, examine zk_p99_quorum_ack_latency for spikes on specific follower connections. ZooKeeper uses two inter-server ports: one for follower-to-leader communication and one for leader election. If only one is blocked, the ensemble may elect a leader but cannot re-elect after a failure. Check for asymmetric connectivity between specific pairs.

  5. Verify ZooKeeper version. If running 3.5.7 or earlier on a 3-node ensemble, ZOOKEEPER-2164 is a known cause. When the leader shuts down, the remaining two nodes both vote, but the winner does not receive votes within 5 seconds because the connection attempt to the dead node does not time out in time. By the time votes arrive, the follower has given up. This is fixed in 3.5.8, 3.6.1, and 3.7.0 (ZOOKEEPER-2164 fixVersions per JIRA).

  6. Check for sidecar proxy interference. If running ZooKeeper behind a service mesh (for example, Istio sidecar proxies), connection timeouts introduced by the sidecar can trigger repeated elections. This is tracked as ZOOKEEPER-3923.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_looking_countDirect count of election events per nodeMore than 1 increment per hour outside maintenance
zk_server_stateCurrent role in ensembleCycling between states or stuck in LOOKING
zk_sum_leader_unavailable_time (3.7+)Cumulative ms without a leaderAny non-zero delta means writes were impossible
zk_cnt_leader_unavailable_time (3.7+)Number of unavailability episodesGrowing count confirms repeated events
zk_max_fsynctime / zk_avg_fsynctimeDisk write path health on each memberAbove 10ms, or any growing trend
zk_p99_jvm_pause_time_ms (3.7+)GC pause impactp99 approaching syncLimit x tickTime (10s default)
zk_p99_quorum_ack_latencyNetwork and follower ACK speedp99 above 50ms, or spikes on specific pairs
zk_outstanding_requestsRequest pipeline backlogSustained non-zero with active traffic
zk_synced_followers (leader)Replication healthBelow ensemble_size - 1

Fixes

Fsync stalls

Confirm with zk_max_fsynctime and iostat -x on the transaction log device. If dataLogDir is not configured, ZooKeeper writes transaction logs to dataDir, sharing the disk with snapshots. Snapshot I/O competes with fsync, causing latency spikes during snapshot creation. Separating dataLogDir onto a dedicated device is the single most impactful configuration change for ZooKeeper write path stability.

If dataLogDir is already on a separate disk, check whether that disk is shared with other I/O-heavy workloads. In cloud environments, EBS gp2/gp3 burst credit exhaustion causes sudden fsync latency cliffs. The transition from burst to baseline is abrupt. Check provisioned IOPS and burst credit balance.

If all members share the same storage tier and that tier is the bottleneck, the fix must address the storage layer across all members. Moving one member to better storage does not help because any member can become leader.

GC pauses

Confirm with zk_p99_jvm_pause_time_ms (3.7+) and the GC log. If Full GC pauses exceed 1 second, the JVM is under heap pressure. Check heap usage against data tree size: zk_znode_count and zk_approximate_data_size tell you how much live data the heap carries. If znode count is growing without bound, application frameworks are leaking nodes.

Short-term mitigation: increase heap size. But if the data tree is growing, this only delays the next crisis. Long-term fix: identify and clean up the leaking subtree, then size the heap with at least 30% headroom above the live data set.

GC algorithm choice matters. G1GC has been the default JVM collector since JDK 9 (ZooKeeper’s launch scripts do not set a collector). Older versions may use CMS or Parallel GC, which produce longer stop-the-world pauses on large heaps. ZGC (available from JDK 15+) dramatically reduces pause times and is worth evaluating for ensembles with large data trees.

Also check Transparent Huge Pages on the host. THP can cause GC pauses to be significantly longer because GC must page in objects to scan them. Verify with cat /sys/kernel/mm/transparent_hugepage/enabled and disable if set to always.

Network issues between pairs

If fsync and GC both look clean, the problem is likely network. ZooKeeper requires both inter-server ports to be reachable for healthy operation: the follower-to-leader port and the leader election port. A firewall blocking only the election port will not be noticed until the next election, at which point the ensemble cannot converge.

Check for asymmetric connectivity: can every node reach every other node on both ports? In cloud environments, security group changes are a common cause. If the ensemble spans datacenters, network latency between sites increases quorum_ack_latency baseline. Cross-datacenter deployments may need maxTimeToWaitForEpoch tuning.

If running with a service mesh sidecar, verify that the sidecar’s connection timeout does not interfere with ZooKeeper’s quorum communication timeout (quorumCnxnTimeoutMs, default -1, which means it uses syncLimit x tickTime).

maxTimeToWaitForEpoch tuning (3.6+)

Introduced in 3.6.0, maxTimeToWaitForEpoch controls how long a leader waits for epoch packets from a majority after receiving a LOOKING notification from a voter. If the leader does not receive epoch packets within this window, it goes back to LOOKING and triggers a new election. If this value is too low for your network latency profile, the leader may prematurely abandon leadership on stray notifications.

The property defaults to -1, which disables the wait entirely. If elections correlate with stray LOOKING notifications rather than heartbeat misses, check this setting. The admin guide notes this can be tuned to reduce quorum unavailability in cross-datacenter environments.

Version-specific bugs

If running 3.5.7 or earlier, upgrade. ZOOKEEPER-2164 causes 3-node ensembles to repeatedly re-elect after a leader shutdown because the connection attempt to the dead node does not time out fast enough. This is fixed in 3.5.8, 3.6.1, and 3.7.0 (ZOOKEEPER-2164 fixVersions per JIRA). Versions 3.5.7 and earlier are reported as unreliable for 3-node leader election.

If running 3.9.x, check for ZOOKEEPER-4925 (diff sync can introduce a hole in a stale follower’s committed log, fixed in 3.9.4 and 3.10.0) and ensure you are on the latest patch release.

Prevention

  • Separate transaction log storage. Put dataLogDir on a dedicated low-latency device with no competing I/O. This is the highest-leverage configuration change for write path stability.
  • Enable GC logging. Without GC logs, you cannot correlate GC pauses with election events. Enable with -Xlog:gc*:file=/var/log/zookeeper/gc.log:time,uptime,level,tags:filecount=5,filesize=100m (JDK 9+ unified logging).
  • Monitor znode growth. All ensemble members carry the same data tree. They will OOM at the same time. Track zk_znode_count and zk_approximate_data_size against heap capacity.
  • Size heap with headroom. Maintain at least 30% free heap above the live data set. The degradation curve is cliff-edge above 85% sustained usage.
  • Disable Transparent Huge Pages. THP amplifies GC pause duration on ZooKeeper hosts.
  • Configure autopurge. Set autopurge.purgeInterval and autopurge.snapRetainCount to prevent unbounded transaction log and snapshot accumulation.
  • Test failover regularly. Controlled leader kills validate that the ensemble can re-elect within the initLimit x tickTime window and that monitoring detects the event.
  • Verify both inter-server ports. Network issues that block only the election port are invisible until the next election.

How Netdata helps

  • Per-second collection of zk_looking_count across all ensemble members surfaces election events the moment they happen, rather than after a scrape interval delay.
  • Correlating zk_p99_fsynctime and zk_p99_jvm_pause_time_ms against election timestamps distinguishes disk-caused storms from GC-caused storms in seconds.
  • ML anomaly detection on zk_server_state transitions catches rapid role cycling that a static threshold might miss during the first minutes of a storm.
  • Leader-only metrics (zk_synced_followers, zk_pending_syncs, zk_p99_quorum_ack_latency) are collected from whichever node currently holds the leader role, so replication health remains visible even during transitions.
  • zk_sum_leader_unavailable_time deltas quantify the actual write availability impact, converting an election count into a severity signal that reflects business impact.