The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-synced-followers-degraded

Operations Guides

ZooKeeper synced_followers below ensemble size: degraded fault tolerance

zk_synced_followers reports how many followers are currently synced with the leader. In a healthy ensemble it equals ensemble_size - 1 (voting members only; observers are excluded). When it drops, a follower is disconnected or lagging, and your fault tolerance margin has shrunk.

This metric is emitted only by the leader. Followers, observers, and standalone nodes do not report it. If your collector scrapes a fixed node or only followers, you have a blind spot: identify the leader dynamically, or scrape every node and keep only the values from the node reporting leader in zk_server_state.

Severity depends entirely on ensemble size and where the new value sits relative to the quorum floor. The critical edge is floor(ensemble_size / 2). At that point the leader plus its synced followers exactly form quorum, and one more failure breaks it. In a 3-node ensemble that edge is reached the moment you lose a single follower, so any drop is already a page.

What this means

zk_synced_followers counts followers fully synchronized with the leader and participating in the proposal-ack-commit pipeline. The expected value is ensemble_size - 1 (voting members only; observers are not counted). Anything below means a follower is not connected, or is connected but still catching up and not yet acknowledging proposals.

zk_learners counts connected learners (followers plus observers) whether or not they are synced. Comparing the two tells you which case you are in:

  • zk_learners below the expected count (all followers plus observers): a learner is disconnected. The leader has no TCP session to it.
  • zk_learners at the expected count but zk_synced_followers below the number of followers: a follower is connected but lagging. It has not finished catching up and is not part of the ack quorum.

Page thresholds:

  • zk_synced_followers == floor(ensemble_size / 2) and ensemble_size > 1 and zk_uptime > 600s: page. The leader plus the synced followers exactly form quorum. One more follower loss means no quorum, no leader election success, no writes.
  • One follower missing in a larger ensemble (a 5-node cluster at synced_followers = 3): ticket. Redundancy is degraded but the cluster can still survive another loss.
  • Brief drop during a rolling restart: info, expected.

The following diagram shows how a single dropped follower maps to severity across common ensemble sizes.

flowchart TD
  A[Follower disconnects or lags] --> B[zk_synced_followers drops]
  B --> C{Ensemble size and loss}
  C -->|"3 nodes, 1 lost"| D["synced = 1 = floor(3/2)
PAGE: zero fault tolerance"] C -->|"5 nodes, 1 lost"| E["synced = 3
TICKET: degraded, still tolerant"] C -->|"5 nodes, 2 lost"| F["synced = 2 = floor(5/2)
PAGE: at quorum edge"] C -->|"7 nodes, 1-2 lost"| G["TICKET: still tolerant"] C -->|"7 nodes, 3 lost"| H["synced = 3 = floor(7/2)
PAGE: at quorum edge"]

A 3-node cluster deserves special attention. Quorum is floor(3/2)+1 = 2, and the leader is always one of those two. When synced_followers drops from 1 to 0, you are sitting exactly at quorum. There is no “degraded but okay” window in a 3-node ensemble the way there is in a 5- or 7-node ensemble. If you run 3 nodes, every missing follower is a page.

Common causes

CauseWhat it looks likeFirst thing to check
Follower process downzk_learners and zk_synced_followers both drop together; the missing node’s ruok fails or zk_uptime resetruok and srvr on the suspect node
Follower GC pauseSingle follower’s zxid stalls, then catches up; zk_jvm_pause_time_ms on that node is elevatedGC log and zk_jvm_pause_time_ms on the follower
Follower disk stallFollower connected but not synced; its zk_fsynctime elevated; zk_pending_syncs on leader non-zerozk_fsynctime and host iowait on the follower
Network partition isolating one followerFollower looks healthy to itself but leader has no session; intermittent if flappingInter-node connectivity, election port reachability
Follower in SNAP sync after restartFollower connected, zxid far behind leader, leader network output to that node sustained highLeader and follower logs for snapshot transfer messages
Non-voting follower inflating the count (3.6+)zk_synced_followers looks fine but zk_synced_non_voting_followers > 0; real voting synced count is lowerzk_synced_non_voting_followers and dynamic reconfig state

Quick checks

These are read-only and safe to run during an incident. The leader-only metrics return nothing on a follower, so run them against every node and filter.

# Identify the leader and its replication metrics, across all ensemble members
for host in zk1 zk2 zk3; do
  echo "=== $host ==="
  echo mntr | nc -w 2 $host 2181 | grep -E 'zk_server_state|zk_learners|zk_synced_followers|zk_pending_syncs'
done
# Confirm which node is leader (followers fields are empty on non-leaders)
echo srvr | nc localhost 2181 | grep Mode
# Compare last-processed zxid across nodes. They should match or be within a few transactions.
for host in zk1 zk2 zk3; do
  printf "%s " "$host"; echo srvr | nc -w 2 $host 2181 | grep '^Zxid'
done
# Check pending syncs on the leader. Sustained non-zero means followers cannot keep up.
echo mntr | nc localhost 2181 | grep zk_pending_syncs
# If running 3.6+, check for non-voting followers that inflate synced_followers.
echo mntr | nc localhost 2181 | grep -E 'zk_synced_non_voting_followers|zk_learners'
# Verify the four-letter-word whitelist is not silently blocking mntr (3.5.3+).
echo ruok | nc localhost 2181
echo mntr | nc localhost 2181 | head -1

If mntr returns nothing or returns mntr is not executed because it is not in the whitelist, your monitoring is likely blind to all of these metrics. Check 4lw.commands.whitelist in zoo.cfg.

How to diagnose it

  1. Confirm you are looking at the leader. Filter all nodes by zk_server_state and keep only the leader’s values for zk_learners, zk_synced_followers, and zk_pending_syncs. A collector pulling from a fixed node may be scraping a follower and seeing nothing.

  2. Distinguish disconnected from lagging. Compare zk_learners and zk_synced_followers.

    • If zk_learners is also low: a learner is not connected at all. Move to network/process checks.
    • If zk_learners is at the expected count but zk_synced_followers is low: a follower is connected but lagging. Move to disk and sync checks.
  3. Find which follower is missing. Compare zxid across all nodes. The lagging or disconnected follower will have a lower zxid, or be unreachable. For a follower mid-SNAP sync, the zxid will be far behind and the leader’s logs will show snapshot transfer activity.

  4. Check the missing follower locally. On that node, look at zk_jvm_pause_time_ms, zk_fsynctime, and host-level iowait. A follower that cannot fsync proposals fast enough will fall behind and drop out of the synced set.

  5. Check inter-node connectivity. ZooKeeper uses two ports between members: one for follower-to-leader traffic and one for leader election. A firewall blocking only the election port will not show up until the next election, at which point the ensemble may fail to re-elect. Verify both ports are reachable between every pair.

  6. Account for non-voting followers (3.6+). zk_synced_followers counts all forwarding followers, including non-voting ones (for example, members removed via dynamic reconfig but still connected). If zk_synced_non_voting_followers is non-zero, subtract it to get the true voting synced count. Otherwise the metric can mask a quorum risk.

  7. Gate on uptime. If zk_uptime on the leader is under 600 seconds, suppress non-critical alerts. Leader-only metrics only appear after election completes, and a cold start can produce transient low values.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_synced_followers (leader)Direct read on fault tolerance marginBelow ensemble_size - 1, or at floor(N/2) for page
zk_learners (leader)Distinguishes disconnected from laggingLower than expected means a learner is gone entirely
zk_pending_syncs (leader)Replication lag in flightSustained non-zero means followers cannot keep up with write rate
zk_synced_non_voting_followers (leader, 3.6+)Corrects the inflated synced countNon-zero means subtract to get true voting synced
zk_follower_sync_timeHow long catch-up takesApproaching syncLimit * tickTime means follower about to be ejected
zk_quorum_ack_latency (leader)Network plus follower processing delayp99 climbing means followers are slow to ack
zk_jvm_pause_time_ms on followersGC stalls freeze the followerp99 approaching syncLimit * tickTime threatens quorum
Last-processed zxid per node (from srvr; mntr does not expose it)Real-time replication positionDivergence between leader and a follower shows lag directly
zk_looking_count (3.6+)Election frequencyIncrementing alongside follower loss suggests instability cascading
zk_sum_leader_unavailable_time (3.7+)Cumulative write unavailabilityGrowing means the cluster already had leaderless periods

Fixes

Treat the fix as specific to the cause. Do not restart services as a first move; restarting the leader loses leader-only metrics context and can trigger an election that worsens things.

Follower process down or unresponsive

Check ruok and srvr on the suspect node. If the process is gone, look at zk_uptime history and the JVM GC log for an OOM or long pause. Restarting that single follower is fine, but identify why it died first. If it OOM-killed, the data tree is likely too large for the heap; address that or it will die again.

Follower GC pause

If zk_jvm_pause_time_ms p99 on the follower is elevated, GC is freezing it. Java 9+ defaults to G1GC (the ZooKeeper startup scripts do not set a collector; ZooKeeper 3.6+ still supports JDK 8, where the JVM default is Parallel GC); older deployments may run CMS. Check heap sizing against zk_znode_count and zk_approximate_data_size. If the follower has the same data tree as the leader (it should), the same heap pressure that eventually hurts the leader is already hurting the follower.

Follower disk stall

If the follower’s zk_fsynctime is elevated, its transaction log disk cannot keep up. Confirm with host-level iostat -x on that node. Common causes: dataLogDir not separated from dataDir, shared storage, cloud storage burst credit exhaustion, or a colocated workload. The durable fix is dedicated low-latency storage for the transaction log. As a stopgap, identify and remove the competing I/O.

Network partition isolating one follower

If the follower looks healthy to itself but the leader has no session to it, suspect the network. Check both inter-node ports between the leader and that follower. A server-to-server auth problem can present the same way, so also check the quorum auth counters: zk_ensemble_auth_fail (mntr, 3.6+) reports ensemble (server-to-server) authentication failures.

Follower in SNAP sync

A follower that was offline long enough that the leader’s transaction log no longer covers its catch-up needs will receive a full snapshot transfer. This is expensive for both sides and can briefly degrade the leader. Verify the follower’s zxid is advancing. If it is not advancing at all, the sync has stalled and you need to investigate why (network, disk on either side). Do not interrupt a progressing SNAP sync.

Non-voting followers inflating the count (3.6+)

If zk_synced_non_voting_followers is non-zero, you have members still connected but no longer voting (common after dynamic reconfig). Your alerting math must subtract these. If you want them fully gone, complete the reconfig removal so they disconnect. Do not leave stale non-voting members in place; they make the quorum math misleading.

Prevention

  • Query the leader, not a random node. Leader-only metrics (zk_learners, zk_synced_followers, zk_pending_syncs) are invisible on followers. Your collector must either identify the leader dynamically or scrape all nodes and filter by zk_server_state.
  • Alert on the quorum edge, not just on “below expected”. A 5-node cluster at synced_followers = 3 is degraded; at synced_followers = 2 it is one failure from outage. Use floor(ensemble_size / 2) as the page threshold.
  • Prefer odd-sized ensembles. A 4-node cluster needs 3 for quorum, the same as a 5-node cluster, but tolerates only 1 failure. You pay for 4 nodes and get the fault tolerance of 3.
  • Gate alerts on uptime. Suppress during cold starts (zk_uptime < 600s) and brief rolling-restart transitions.
  • Watch zk_synced_non_voting_followers on 3.6+. If your alerting uses zk_synced_followers directly, non-voting members can mask a real quorum risk.
  • Separate transaction log storage on every node. Follower disk stalls are a leading cause of followers dropping out of the synced set.
  • Verify the 4lw whitelist includes mntr. Since 3.5.3, a missing whitelist silently returns nothing, and many monitors interpret empty responses as healthy zeros.

A note on version drift: in 3.6.0, the mntr gauge zk_followers was renamed to zk_learners because the new gauge counts every connected learner (followers plus observers), which is what the old name actually measured (verified in LeaderZooKeeperServer.registerMetrics: 3.5.x exposed followers; 3.6.0 replaced it with learners). zk_synced_followers was not renamed and retains its semantics. Dashboards and alerts written against pre-3.6 metric names will silently return no data after upgrade.

How Netdata helps

  • Netdata collects zk_server_state per node, so you can build alerts that automatically use only the leader’s values for zk_synced_followers, zk_learners, and zk_pending_syncs without hardcoding which host is the leader.
  • Per-second resolution on zk_synced_followers and zk_pending_syncs catches transient drops that a 60-second scraper misses entirely, which matters for followers that flap in and out of the synced set.
  • Correlating zk_synced_followers with zk_jvm_pause_time_ms and zk_fsynctime on each follower pinpoints whether a drop was caused by GC, disk, or network, without switching tools.
  • Anomaly detection on zk_learners minus zk_synced_followers surfaces the “connected but lagging” case before it becomes a quorum edge.
  • Tracking zk_synced_non_voting_followers alongside zk_synced_followers keeps the quorum math honest on 3.6+ ensembles that use dynamic reconfig.
  • Alerting on the floor(ensemble_size / 2) edge, gated by zk_uptime, avoids both cold-start false positives and the far worse failure of not paging when you are one loss from outage.