The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-how-it-works-in-production

Operations Guides

How ZooKeeper actually works in production: a mental model for operators

ZooKeeper is a distributed coordination service built on a replicated state machine. It maintains a hierarchical namespace of data nodes (znodes) entirely in memory and replicates every mutation across an ensemble of servers using the ZAB (ZooKeeper Atomic Broadcast) protocol.

This is the mental model the rest of the ZooKeeper runbooks assume. It is the set of abstractions an on-call engineer needs to reason about why writes stall, sessions expire, a “healthy” node can return stale reads, and why the leader matters more than any other node in the cluster.

Two facts anchor the whole model. First, the entire znode tree lives in the JVM heap of every server. Second, every write is fsync’d to the transaction log before it is acknowledged. Reads touch local memory. Writes touch local memory plus quorum plus disk. That asymmetry explains almost every ZooKeeper production incident.

What it is and why it matters

From the outside, ZooKeeper is a small hierarchical key-value store with strong consistency guarantees on writes. From the inside it is a replicated state machine: every server applies the same ordered sequence of transactions and converges on the same tree state. The ordering primitive is the zxid, a monotonically increasing 64-bit transaction ID assigned by the leader.

This determines the cost of every operation:

  • A read is a local memory lookup. No quorum, no disk, no network round-trip beyond the client connection. Reads are cheap, fast, and scale horizontally by adding followers or observers.
  • A write is a quorum operation. It must be assigned a zxid by the leader, persisted to the transaction log on disk with an fsync, broadcast to followers, acknowledged by a quorum, then committed. Writes are serialized through the leader and gated by disk I/O.

When something breaks, the symptom points back to one of those two paths: slow writes to the transaction log disk, stale reads to a follower lagging behind the leader, and mass session expirations to a process freeze (GC pause) that interrupted heartbeat traffic.

How it works

The request pipeline is the same on every server, but the work it does depends on role.

flowchart LR
  Client[Client request] --> Split{Read or write}
  Split -->|read| Tree[Local znode tree]
  Tree --> ReadResp[Respond from memory]
  Split -->|write| Leader[Leader]
  Leader --> Zxid[Assign zxid, fsync txnlog]
  Zxid --> Quorum[Quorum of followers ACK]
  Quorum --> Commit[Commit and apply]
  Commit --> WriteResp[Respond to client]

The in-memory znode tree

Every server holds the full data tree in JVM heap: every znode, its data payload, its ACL reference, its children list, and its stat structure. Session state, watch tables, queued requests, and serialization buffers consume heap on top of the tree itself.

This is why zk_znode_count and zk_approximate_data_size are leading indicators of heap exhaustion. A tree that grows slowly over months looks fine until heap pressure crosses roughly 85-90 percent, at which point garbage collection frequency spikes, stop-the-world pauses grow from milliseconds to seconds, and sessions begin expiring during those pauses. All servers carry the same tree, so they all OOM at roughly the same time.

The write path

Writes are processed exclusively by the leader. Followers that receive a write from a client forward it. For every mutation, the leader does the following:

  1. Assigns a monotonically increasing zxid.
  2. Appends the transaction to its write-ahead transaction log (txnlog) and fsyncs it to disk.
  3. Broadcasts a PROPOSE message containing the transaction to all followers via ZAB.
  4. Waits for a quorum of followers to ACK. Each follower ACKs only after it has also fsynced the transaction to its own txnlog.
  5. Commits the transaction once quorum is reached, applies it to its in-memory tree, and sends a COMMIT to followers.
  6. Followers apply the commit to their own in-memory trees.

The leader’s fsync is on the critical path of every write. Each follower’s fsync is on the critical path of every ACK. If either is slow, the whole write pipeline stalls. Dedicated low-latency storage for the txnlog is the single most important deployment decision for ZooKeeper. Shared or slow storage is the most common cause of write stalls in production.

The read path

Reads are served from the local in-memory copy. They never touch disk at runtime, never contact the leader, and never require quorum agreement. This is what makes ZooKeeper fast for read-heavy workloads, and it is also what makes reads potentially stale.

A follower that has fallen slightly behind the leader will happily serve reads from its older in-memory state. ZooKeeper offers sequential consistency for reads, not strict linearizability. Clients that need read-after-write consistency must issue a sync() call before reading, which forces the follower to catch up to the leader’s latest committed state.

Sessions, ephemeral nodes, and watches

A session is the unit of client identity. Ephemeral nodes, watches, and ACL enforcement all bind to sessions. Sessions have a negotiated timeout enforced by the leader’s session tracker. If a client fails to send a heartbeat within the timeout window, the session expires.

Session expiry is one of the most impactful events in ZooKeeper because it cascades:

  • Every ephemeral node created by the session is deleted.
  • Every watch registered by the session is cleared.
  • Every system that watches those ephemeral nodes (Kafka broker registration, HBase RegionServer registration, distributed locks) receives a notification.

A mass session expiry event is ZooKeeper’s thundering-herd failure mode: ephemerals vanish, watches fire, clients reconnect simultaneously, and the leader (new or old) inherits the load spike.

Watches are one-shot triggers in classic ZooKeeper. When a watched znode changes, the server queues a notification to every watcher and then the watch is removed. A heavily-watched znode (a service discovery path with thousands of consumers) produces a fan-out notification storm when it changes. Persistent and recursive watches introduced in later versions change the volume characteristics, but not the underlying fan-out risk.

Persistent and recursive watches were introduced in ZooKeeper 3.6.0 (ZOOKEEPER-1416).

Persistence and recovery

Every mutation is appended to the write-ahead txnlog and fsynced before acknowledgment. This is the durability primitive. Periodically the server serializes the entire in-memory data tree to a snapshot file. Snapshots are fuzzy: they are taken without pausing the request pipeline and may include concurrent updates. They are consistent when replayed against the txnlog from the snapshot’s starting zxid forward.

On startup, a server loads the latest snapshot and replays the txnlog forward. For a large data tree this can take minutes, during which the node appears unresponsive. Recovery time scales linearly with snapshot size plus the volume of txnlog to replay.

Key defaults: forceSync defaults to true; setting it to false skips the fsync and is documented as unsafe. fsync.warningthresholdms defaults to 1000ms and logs a warning when an fsync exceeds it. snapCount defaults to 100,000 transactions between snapshots, with randomization to avoid synchronized snapshot storms across the ensemble. autopurge.purgeInterval defaults to 0, which disables autopurge; many deployments forget to enable it and slowly accumulate txnlog and snapshot files until disk fills. jute.maxbuffer defaults to roughly 1MB and caps the size of any single znode payload.

Where it shows up in production

The mental model maps directly onto recurring operational patterns:

  • The leader is a serialization point. All writes go through it. Its disk, its CPU, and its network are the bottleneck for write throughput. A leader under write load can saturate while followers sit idle.
  • Heap is the entire state. Slow znode growth is a silent killer. Track zk_znode_count and zk_approximate_data_size against heap allocation, not just current heap utilization.
  • Disk fsync is the write latency floor. A write cannot complete until both the leader and a quorum of followers have fsynced. Anything that interferes with that fsync (shared storage, cloud burst credit exhaustion, colocated workloads, snapshot I/O on the same disk) directly elevates write latency.
  • GC pauses freeze everything. A stop-the-world GC pause stops heartbeats, request processing, and quorum ACKs simultaneously. On the leader, a long enough pause triggers re-election. On any node, a long enough pause expires client sessions.
  • Reads from followers can be stale. Applications that need strict read-after-write consistency must use sync(), or accept sequential consistency.

The default per-IP connection limit is another recurring trap. maxClientCnxns defaults to 60 per source IP, not per server. In containerized environments where many pods share a host IP behind NAT, new client connections are silently rejected.

zk_connection_rejected is a real counter (metric connection_rejected) surfaced by mntr on ZooKeeper 3.6.0+ when the default metrics provider is used; it increments when maxClientCnxns rejects a connection. It does not exist on 3.5.x.

Tradeoffs and when it matters

PropertyConsequence
Reads are local memoryFast, scalable, potentially stale on followers
Writes are quorum + fsyncStrongly consistent, serialized through leader, disk-bound
Entire tree in heapHeap exhaustion is the silent killer; all nodes OOM together
Sessions bind ephemeral stateMass expiry cascades to every dependent system
Watches are fan-outA popular znode changing triggers a thundering herd
maxClientCnxns default 60 per source IPSilently rejects new connections in NAT’d environments

ZooKeeper is a coordination service, not a database. Storing large data in znodes is an anti-pattern that bloats the tree, slows snapshots, and lengthens recovery. Use it for what it is designed for: small configuration, leader election, distributed locks, service discovery registrations, and small state machines that benefit from linearizable writes.

For deployment topology, an ensemble of three, five, or seven voting members is standard. Quorum is floor(N/2)+1. A three-node ensemble tolerates one failure; a five-node tolerates two. Observers are non-voting members that receive the commit stream for read scaling without participating in quorum, useful when read load is geographically distributed but write latency must stay low.

Signals to watch in production

Metric availability is version-dependent with the default metrics provider: on 3.5.x mntr exposes only the aggregate zk_avg_latency, zk_min_latency, zk_max_latency (plus the legacy counters); the per-path summaries zk_fsynctime, zk_readlatency, zk_updatelatency and the counters zk_digest_mismatches_count, zk_connection_rejected, zk_unrecoverable_error_count, zk_stale_sessions_expired require ZooKeeper 3.6.0+, and zk_jvm_pause_time_ms requires 3.6.2+. With the Prometheus metrics provider (3.6+), all of these are available over the /metrics endpoint instead.

SignalWhy it mattersWarning sign
zk_server_stateWhich path each node is onNo leader, or more than one leader
zk_fsynctime (p99)The write latency floorSustained p99 above 10ms, or any upward trend
zk_updatelatency (p99)Full write round-tripTracks fsync; divergence means queue or quorum problem
zk_readlatency (p99)Local memory path healthSustained elevation implies GC or deep tree
zk_outstanding_requestsPipeline backlogSustained non-zero means saturation
zk_jvm_pause_time_ms (p99)Process freeze durationApproaching tickTime threatens quorum
zk_synced_followers (leader only)Replication healthBelow N - 1 means degraded fault tolerance
zk_znode_countTree size trendLinear or unbounded growth
zk_stale_sessions_expiredSession cascade signalAny non-zero rate outside maintenance
zk_digest_mismatches_countData integrityAny increment is critical

zk_followers, zk_synced_followers, and zk_pending_syncs are leader-only metrics. Monitoring must identify and query the current leader, or query all nodes and filter for the leader-reported values. A scrape that only hits followers will never see replication health, which is one of the most common monitoring gaps in ZooKeeper deployments.

How Netdata helps

zk_throttled_ops (counter throttled_ops, 3.7.0+) and zk_unrecoverable_error_count (counter unrecoverable_error_count, 3.6.0+) are the exposed names for these signals.

Per-second collection and anomaly detection suit ZooKeeper’s failure modes: slow drift (znode growth, fsync latency creep) that becomes a cliff-edge incident (heap exhaustion, write stall, session cascade).

  • Correlate zk_fsynctime with zk_updatelatency on the same chart to confirm disk as the root cause of write stalls.
  • Overlay zk_jvm_pause_time_ms against zk_stale_sessions_expired and zk_looking_count to identify GC as the trigger for session cascades and elections.
  • Track zk_znode_count and zk_approximate_data_size as long-running trends to surface silent heap exhaustion months before OOM.
  • Watch zk_outstanding_requests and zk_throttled_ops together to catch pipeline saturation before clients see timeouts.
  • Surface zk_digest_mismatches_count and zk_unrecoverable_error_count as immediate-page signals for data integrity violations.