The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-avg-latency-hiding-write-stalls

Operations Guides

ZooKeeper avg_latency hides write stalls: why the headline number lies

The dashboard says zk_avg_latency is 1.2 ms. Clients are timing out on writes. Both can be true. On a read-heavy ZooKeeper ensemble, the headline latency number can look healthy while the write path is stalled.

Two properties cause this. First, zk_avg_latency, zk_min_latency, and zk_max_latency aggregate reads and writes into one number. Reads are served from local memory and complete in microseconds. Writes require a quorum round-trip plus a transaction log fsync before acknowledgment. When reads dominate the request mix, a severe write stall is diluted by thousands of cheap reads and disappears into the average.

Second, all three metrics are server-cumulative since the last srst reset or process start. They are not sliding windows. A three-hour-old GC pause keeps zk_max_latency pinned at that duration indefinitely, and the average drifts toward the workload average over the entire uptime. Dashboards built on these values are unreliable as recent-behavior signals.

What this means

The aggregation problem is arithmetic. Suppose an ensemble handles 950 reads per interval at 0.5 ms each and 50 writes. If write latency climbs to 500 ms during a disk stall, the weighted average across all requests is (950 * 0.5 + 50 * 500) / 1000, or about 25 ms. A 25 ms average looks like mild congestion. It is not. Every write is taking half a second, and clients with short session timeouts or tight lock-acquire budgets are failing. The same arithmetic that makes the average look benign also breaks threshold alerts: a “page when avg_latency > 100 ms” rule will not fire even when every write is broken.

The cumulative trap compounds this. zk_max_latency only goes up until you issue srst or restart the process. A 15-second GC pause from last Tuesday pins the max at 15,000 ms indefinitely. Teams learn to ignore it, and when a real stall arrives the alert is already suppressed as noise.

Together these produce a classic failure pattern: a write-heavy downstream system (Kafka controller election, HBase region assignment, a distributed lock service) starts failing, the on-call engineer checks the ZooKeeper dashboard, sees a flat average and a max that has “always been high,” and concludes ZooKeeper is fine. It is not.

Common causes

CauseWhat it looks likeFirst thing to check
Transaction log disk stallFsync p99 elevated, write latency p99 tracking it, zk_outstanding_requests climbing on the leaderFsync percentiles via the metrics endpoint, plus iostat -x on the txnlog device
Shared or co-located storagePeriodic fsync spikes, often aligned with snapshot creation when dataLogDir equals dataDirWhether dataLogDir is set in zoo.cfg and points to a dedicated device
Cloud storage throttlingSudden fsync cliff after a period of normal latency, correlates with burst credit exhaustionCloud provider IOPS and burst metrics for the txnlog volume
JVM GC pausesJVM pause p99 elevated, both read and write latency spike together, zk_outstanding_requests builds then drainsJVM pause percentiles via the metrics endpoint, plus GC logs
Quorum ACK delayQuorum ACK p99 elevated on the leader, writes slow even when leader disk is fineQuorum ACK latency on the leader, follower fsync times
Cumulative-metric dashboard illusionzk_avg_latency looks flat and low, zk_max_latency pinned high for days, clients report write timeoutsWhether your monitoring ever issues srst, and whether you collect per-type write latency at all

Quick checks

The mntr four-letter-word command exposes the aggregated latency metrics (zk_avg_latency, zk_min_latency, zk_max_latency), basic pipeline counters (zk_outstanding_requests, zk_server_state), and, since 3.6, per-type latency percentiles, fsync times, quorum ACK latency, JVM pause metrics, and throttled operations. These 3.6+ metrics are also available in Prometheus format from the Prometheus metrics provider’s HTTP endpoint (default port 7000).

# Confirm the version - per-type latency metrics require 3.6+
echo srvr | nc localhost 2181 | grep -E "Zookeeper version|Mode"

# Aggregated latency (cumulative since last srst) - from mntr
echo mntr | nc localhost 2181 | grep -E "zk_(avg|min|max)_latency"

# Per-type read/write latencies (3.6+) - from mntr
echo mntr | nc localhost 2181 | grep -E "zk_.*(update|read)latency"

# Fsync percentiles (3.6+) - the usual root cause of write stalls
echo mntr | nc localhost 2181 | grep -E "zk_.*fsynctime"

# Request pipeline backlog - from mntr
echo mntr | nc localhost 2181 | grep "zk_outstanding_requests"

# Quorum ACK latency (leader) and JVM pause time (3.6+) - from mntr
echo mntr | nc localhost 2181 | grep -E "zk_.*(quorum_ack|jvm_pause)"

# Throttled operations (3.6+) - means clients are being dropped at the global limit
echo mntr | nc localhost 2181 | grep "zk_throttled_ops"

# Verify the four-letter-word whitelist includes mntr (3.5.3+)
echo mntr | nc localhost 2181 | head -1
# An empty response means mntr is not whitelisted and monitoring is silently broken

How to diagnose it

  1. Confirm the ZooKeeper version exposes per-type latency metrics. Anything before 3.6.0 only provides the aggregated zk_avg/min/max_latency, so you cannot separate reads from writes at all. If you are on an older version, the only options are computing deltas between scrapes and periodically issuing srst.

  2. Pull per-type write and read latency percentiles from mntr. If write latency is several orders of magnitude above read latency, the write path is the problem and the aggregated average is hiding it.

  3. Pull fsync percentiles. On a healthy dedicated SSD, p99 fsync should be under 2 ms. If it is in the tens or hundreds of milliseconds, disk I/O is the root cause.

  4. Check zk_outstanding_requests on the leader via mntr. It should be zero in steady state. A sustained non-zero value with active traffic means the pipeline cannot keep up. If it is approaching globalOutstandingLimit (default 1000), throttling will kick in and clients will see timeouts. The zk_throttled_ops counter tracks operations already dropped at this limit.

  5. Check quorum ACK latency percentiles on the leader. If this is elevated while leader fsync is fine, the bottleneck is follower-side: follower GC, follower disk, or network between leader and followers.

  6. Check JVM pause time percentiles. If GC pauses are the cause, both read and write latency will spike together rather than writes alone.

  7. If you are stuck on pre-3.6 metrics and cannot upgrade, you can issue srst at a fixed interval and compute deltas. Warning: srst resets all server statistics simultaneously. Every monitoring system scraping that node will see a counter reset at the same moment. If multiple tools depend on the cumulative counters, coordinate the reset cadence or only one of them will see meaningful deltas.

flowchart TD
    A["zk_avg_latency looks fine"] --> B{"Collecting per-type write latency?"}
    B -- "No (pre-3.6 or not scraping)" --> C["Hidden: cannot tell reads from writes"]
    B -- "Yes" --> D{"Write latency p99 high?"}
    D -- "No" --> E["Read path or client-side issue"]
    D -- "Yes" --> F{"Fsync p99 high?"}
    F -- "Yes" --> G["Disk stall: shared storage, throttling, or hardware"]
    F -- "No" --> H{"Quorum ACK p99 high?"}
    H -- "Yes" --> I["Follower GC, follower disk, or network"]
    H -- "No" --> J["Check JVM pauses and pipeline queue"]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Write latency percentilesWrite-path latency separated from reads. The metric that actually reflects write stalls.p99 climbing while read latency stays flat
Read latency percentilesRead-path latency. Reads are memory lookups, so elevation usually means GC or deep hierarchies.p99 above single-digit milliseconds
Fsync percentilesTime to fsync the transaction log. The single most common write-stall root cause.p99 above 10 ms on SSD, or any sustained upward trend
Quorum ACK latency percentilesTime from PROPOSE to quorum ACK on the leader. Captures follower and network delays.p99 above 50 ms, or approaching syncLimit * tickTime
zk_outstanding_requestsRequest pipeline backlog. Leading indicator: it fills before latency spikes.Sustained non-zero, or approaching globalOutstandingLimit
zk_throttled_opsCounter of operations throttled at the global limit. Means clients are already being dropped.Any non-zero rate
JVM pause time percentilesGC pause impact. When elevated, both reads and writes spike together.p99 approaching a meaningful fraction of minSessionTimeout
zk_max_latency (handled with care)Cumulative extreme since last srst. Useful only as a delta.Pinned at a high value for days means nobody is resetting or computing deltas

Fixes

Separate reads from writes in monitoring

The first fix is instrumentation, not infrastructure. If you are on 3.6 or later, collect write latency and read latency percentiles separately via mntr and alert on them independently. Alert on write latency p99 crossing a baseline-relative threshold, not on zk_avg_latency.

If you are on a version before 3.6, you have two options. The first is to issue srst at the end of each scrape and compute deltas, which gives you per-interval min, avg, and max. The second is to upgrade. Given that 3.6 has been out for years and per-type percentile metrics are the only reliable way to see write stalls, upgrading is usually the right call.

Tradeoff: srst resets the cumulative counters that other tools may depend on. If you have multiple monitoring systems scraping the same node, coordinate the reset cadence or only one of them will see meaningful deltas.

Put the transaction log on dedicated storage

If fsync p99 is the smoking gun, the most impactful single change is setting dataLogDir to a dedicated low-latency device that holds nothing else. The default is for the transaction log and snapshots to share dataDir, which means snapshot writes (large, bulk, sequential) compete with fsync (small, latency-critical). This shows up as periodic fsync spikes aligned with snapshot creation.

On cloud instances, also check whether the txnlog volume is a burstable type. Burst credit exhaustion produces a sudden fsync cliff that looks like hardware failure but is really throttling. Move to provisioned IOPS or a higher tier.

Tradeoff: dedicated storage costs money and adds an operational variable. It is almost always worth it for ZooKeeper, because fsync latency dominates write latency and ensemble stability.

Investigate quorum ACK latency

If leader fsync is fine but quorum ACK p99 is elevated, the bottleneck is on the followers or the network. Pull fsync percentiles from each follower. A single slow follower (failing disk, long GC) can inflate quorum ACK latency because the leader must wait for a majority.

Check whether zk_synced_followers on the leader is below ensemble_size - 1. A follower that is connected but lagging will drag quorum ACK time without failing health checks.

Tradeoff: removing a slow follower from the ensemble to protect quorum ACK latency reduces fault tolerance. In a three-node ensemble, removing one follower leaves no redundancy.

Address JVM GC pauses

If JVM pause p99 is elevated and both read and write latency spike together, GC is the cause. Common fixes: size the heap to the data tree (monitor zk_znode_count and zk_approximate_data_size), switch to G1GC if you are still on CMS or Parallel, and consider ZGC on JDK 15+ for low-pause collection.

Tradeoff: a larger heap means longer full GC pauses when they do occur. The goal is enough headroom that full GC is rare, not so large that a full GC is catastrophic.

Stop trusting zk_max_latency as a live signal

If your dashboards show zk_max_latency pinned at a high value for days, either reset it periodically with srst or compute deltas in your monitoring system. Treat the raw cumulative value as a “worst ever since start” marker, not a current-condition signal.

Prevention

  • Collect per-type latency from day one. On 3.6+, scrape write latency and read latency with percentiles. The aggregated average is structurally unable to surface write stalls on read-heavy clusters.
  • Alert on write latency p99, not zk_avg_latency. Use a baseline-relative threshold (for example, 3x the rolling p99) rather than an absolute number, since write latency baselines vary by workload.
  • Track fsync p99 as a first-class metric. It is the leading indicator for the most common write-stall root cause and is almost never collected by default.
  • Set dataLogDir to a dedicated device. Every production ensemble should isolate the transaction log from snapshot I/O to avoid periodic fsync spikes during snapshot creation.
  • Reset or delta zk_max_latency. Decide on a reset cadence (per scrape, per minute, per hour) and stick to it, or have your monitoring system compute deltas. A pinned zk_max_latency is a monitoring smell.
  • Gate cold starts. Suppress non-critical latency alerts when zk_uptime is below 300 seconds, because latency metrics are noisy immediately after restart.

How Netdata helps

  • Netdata collects the full mntr output per second, including zk_avg_latency, zk_min_latency, zk_max_latency, zk_outstanding_requests, and leader-only metrics like zk_synced_followers and zk_pending_syncs.
  • Per-second collection means a write stall that lasts only a few seconds still shows up as a distinct spike, rather than being averaged away in a longer scrape window.
  • The anomaly advisor correlates write latency with fsync time, outstanding requests, quorum ACK latency, and JVM pause time, so when latency deviates from baseline the related root-cause metrics are highlighted in the same view.
  • Because Netdata derives per-interval rates from cumulative counters, the zk_max_latency cumulative trap is handled by computing deltas rather than charting the raw monotonically-increasing value.
  • Leader-only metrics are collected from every node and filtered by reported zk_server_state, so replication health is visible without manually identifying the leader first.