The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-session-count-climbing

Operations Guides

ZooKeeper session count climbing: leaks and duplicate sessions

zk_global_sessions climbing while your client fleet is stable is a slow ZooKeeper failure mode. The ensemble keeps serving reads and writes, latency looks fine, quorum is intact, but the session table keeps growing. Each entry costs heap and periodic heartbeat processing. Eventually you hit a GC death spiral, an OOM, or a maxClientCnxns-shaped outage.

The signal is simple to read but easy to misinterpret. ZooKeeper exposes two distinct populations: global sessions, which the leader echoes across the ensemble, and local sessions (only present when localSessionsEnabled=true, added in ZK 3.5, default false), which live on a single follower and upgrade to global when the client creates an ephemeral node. If you only watch zk_global_sessions you may be looking at a subset of the actual session population, and if you only watch zk_num_alive_connections you cannot tell a leak from a deployment event.

The three signals that matter are zk_global_sessions, zk_num_alive_connections, and zk_ephemerals_count.

What this means

A global session in ZooKeeper is the unit of client identity that the leader has agreed to track across the ensemble. Ephemeral nodes, watches, and ACL enforcement all bind to a session ID. The leader maintains a session tracker with expiry buckets; each session must renew within its negotiated timeout or it expires, taking its ephemeral nodes and watches with it.

zk_global_sessions from mntr (3.6+; on 3.5.x the value is JMX-only) reports the count of these leader-tracked sessions. Each session consumes a small but non-zero amount of heap (session state, watch table entries if the client sets watches, queued request state) and requires periodic heartbeat processing on the request pipeline.

Monotonic growth in zk_global_sessions that never plateaus, without a matching change in client deployment, is the canonical tell for one of two pathologies.

  1. Session leak: a client library or application opens new sessions without retiring old ones. The leak is usually client-side (a framework that spawns new ZooKeeper handles without closing the previous one), less often server-side (a buggy reconnection path that establishes a new session before the old one expires).

  2. Duplicate sessions from reconnection churn: clients enter a reconnect loop where each cycle creates a new session alongside the previous one, either because the client never observes the SESSION_EXPIRED event that should gate new session creation, or because a framework’s reset path races the user code’s close path.

The third pattern worth knowing is benign growth from local-session promotion. With localSessionsEnabled=true, sessions start local to a follower and are upgraded to global when they create ephemeral nodes. A workload shift that increases ephemeral-node creation will inflate zk_global_sessions without any change in actual client count.

flowchart TD
    A[zk_global_sessions climbing] --> B{num_alive_connections also climbing?}
    B -- Yes, proportional --> C[Real client growth, check deployment]
    B -- No, flat or declining --> D{ephemerals_count also climbing?}
    D -- Yes, proportional --> E[Legitimate ephemeral use or local-session upgrade]
    D -- No --> F[Session leak or duplicate-session churn]
    F --> G{stale_sessions_expired climbing?}
    G -- Yes --> H[Reconnection loop, investigate client libraries]
    G -- No --> I[Pure leak, clients opening new handles without closing]

Common causes

CauseWhat it looks likeFirst thing to check
Client framework leak (Curator and similar)zk_global_sessions rises steadily; cons shows duplicate session IDs from the same client IP; zk_stale_sessions_expired flat or moving slowlyClient library version; grep client logs for repeated Expired events followed by new session establishment
Reconnection loop creating duplicateszk_global_sessions and zk_stale_sessions_expired both climb; connection churn visible in zk_connection_drop_countClient GC logs; network stability between clients and ensemble
Local-to-global session promotionzk_global_sessions climbs when localSessionsEnabled=true; total connection count is flatlocalSessionsEnabled and localSessionsUpgradingEnabled in zoo.cfg; ephemeral creation rate
maxClientCnxns saturation masquerading as leakzk_connection_rejected increments alongside zk_num_alive_connections plateauingmaxClientCnxns per source IP; container-to-host IP fan-in ratio
Client-side GC pauses causing session churnzk_stale_sessions_expired spikes rhythmically; client-side JVM pauses correlateClient JVM GC logs; client-side heap sizing

Quick checks

All read-only and safe in production. cons and dump are O(n) over the connection and session tables and can cause brief latency spikes on large ensembles; sample sparingly.

# Headline session count
echo mntr | nc localhost 2181 | grep zk_global_sessions

# Pair with live connection count and ephemeral count
echo mntr | nc localhost 2181 | grep -E 'zk_(num_alive_connections|ephemerals_count)'

# Session expiration and connection drop rates
echo mntr | nc localhost 2181 | grep -E 'zk_(stale_sessions_expired|connection_drop_count)'

# Per-connection detail, including session IDs (sid=0x...). Expensive, sample sparingly
echo cons | nc localhost 2181

# Outstanding sessions and ephemeral nodes. Expensive, sample sparingly
echo dump | nc localhost 2181

# Check whether local sessions are enabled (changes interpretation of zk_global_sessions).
# Path varies by distribution.
grep -E 'localSessionsEnabled|localSessionsUpgradingEnabled' /etc/zookeeper/conf/zoo.cfg

# JVM pause times. A leading cause of session expiration cascades.
# zk_jvm_pause_time_ms (avg/min/max/cnt/sum + p50-p999) is in mntr from 3.7+,
# and only while jvm.pause.monitor=true; on 3.5.x/3.6.x use GC logs.
echo mntr | nc localhost 2181 | grep -i jvm_pause

# Per-IP connection limit and rejected-connection counter
echo mntr | nc localhost 2181 | grep connection_rejected
grep -E 'maxClientCnxns' /etc/zookeeper/conf/zoo.cfg

If mntr returns nothing, your four-letter words may not be whitelisted. Since ZK 3.5.3 you must set 4lw.commands.whitelist (or the Java system property zookeeper.4lw.commands.whitelist) to include at least mntr. The AdminServer on port 8080 is the long-term replacement for four-letter words.

How to diagnose it

  1. Establish the baseline. Capture zk_global_sessions, zk_num_alive_connections, and zk_ephemerals_count at one-minute intervals for at least 30 minutes. The single most important question is whether the growth is monotonic (never plateaus, never reverses) or stepwise (jumps, then flat). Pure monotonic growth without any plateau is the leak signature.

  2. Compute the session-to-connection ratio. In steady state, zk_global_sessions / zk_num_alive_connections should be roughly constant. If the ratio is climbing, you have more sessions than connections can explain, which points at duplicate sessions or sessions that outlive their connections. If the ratio is constant and both are climbing together, look for a real client growth event first.

  3. Check whether local sessions are in play. If localSessionsEnabled=true, zk_global_sessions only counts sessions promoted to global (typically by creating an ephemeral node). A workload shift that increases ephemeral creation will inflate zk_global_sessions without any leak. Confirm by checking whether zk_ephemerals_count is climbing proportionally.

  4. Pull per-connection detail and look for duplicate session IDs from the same source IP. Use cons to list every connection with its session ID. Multiple live sessions originating from the same client IP and process indicate a duplicate-session bug, not a generic leak.

  5. Correlate with session expirations. If zk_stale_sessions_expired is also climbing, the leak is dynamic: sessions are expiring but new ones are being created faster. If zk_stale_sessions_expired is flat while zk_global_sessions climbs, sessions are accumulating without expiring, which is a harder leak where the server believes the sessions are alive even though the clients have moved on.

  6. Check JVM pause times on both server and client. Server-side GC pauses cause session expiration cascades that masquerade as churn. Client-side GC pauses cause clients to miss heartbeats and reconnect, often creating transient duplicate sessions. The server-side signal is the JVM pause metric from mntr; the client-side signal must be read from the client’s own GC logs.

  7. Inspect the client library. The single most common cause of duplicate ZooKeeper sessions in production is a client framework bug, not a ZooKeeper bug. Apache Curator has a known class of issues where the framework’s reset path can race the user’s close path on session expiration, leaving an orphaned session connected alongside the new one. Check the framework version against the upstream issue tracker before changing anything on the server. CURATOR-722 (“Zookeeper connection leak after session expiration”) was still open with no fix version as of this review.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_global_sessionsPrimary leak signalMonotonic growth without plateau; growth without matching client deployment
zk_num_alive_connectionsValidates whether session growth is real client growthSessions climbing while connections are flat indicates duplicate sessions
zk_ephemerals_countTracks the dynamic part of the data tree; correlates with legitimate session useClimbing in proportion to sessions is benign; climbing faster indicates per-session ephemeral leaks
zk_stale_sessions_expiredRate of sessions that died from missed heartbeatsNon-zero sustained rate indicates churn that can mask a leak
zk_connection_drop_countConnection-level instabilityBurst pattern correlates with reconnection storms
JVM pause (server)Long GC pauses cause cascading session expirationsp99 climbing toward minSessionTimeout (default 2 x tickTime = 4000ms)
zk_connection_rejectedSilent rejection from maxClientCnxns per source IPAny non-zero rate in production
zk_open_file_descriptor_countSessions and connections consume FDs; leak eventually exhausts themopen/max ratio above 80%

Fixes

Client framework session leak

The right fix is in the client, not the server. For Apache Curator specifically, upgrade to the latest patch release and audit the application’s CuratorListener implementation. The mandatory rule from the ZooKeeper FAQ is: only create a new session when you are notified of SESSION_EXPIRED, and SESSION_EXPIRED automatically closes the existing ZooKeeper handle. Application code that opens a new handle on any connection-state change (not just expiration) will leak sessions.

If you cannot immediately push a client fix, you can buy time by raising the client-side session timeout so fewer expirations occur per unit time. This slows the leak but does not stop it. A scheduled bounce of the leaking client process, keeping the session count below your heap headroom threshold, is a stopgap, not a fix.

Do not bounce the ZooKeeper ensemble to clear leaked sessions. The clients will reconnect and recreate the leak within minutes, and you will have paid an availability cost for nothing.

Duplicate sessions from reconnection churn

Treat this as a client bug with server-side symptoms. The diagnostic question is: why is the client creating a new session before the old one is observed as expired? Common root causes.

  • Client-side JVM GC pauses longer than the negotiated session timeout. Fix the client heap, switch to G1GC (JVM default since JDK 9) or ZGC (production-ready since JDK 15), and consider raising the session timeout if the client legitimately needs long GC pauses. ZooKeeper’s own launch scripts do not set a collector: G1GC is the JDK default since JDK 9, so the GC choice is a JVM property, not a ZooKeeper one.
  • Network instability between the client and ensemble that causes repeated TCP resets. Stabilize the network path or move the client closer to the ensemble.
  • A framework’s reconnect logic that does not wait for the SESSION_EXPIRED callback before establishing a new session.

Local-to-global session promotion misread as a leak

If localSessionsEnabled=true and your workload is creating more ephemeral nodes than before, zk_global_sessions will climb because local sessions are being promoted. This is benign and expected. Confirm by checking that zk_ephemerals_count is climbing proportionally, that zk_num_alive_connections is stable, and that no workload change increased ephemeral-node creation unexpectedly.

No fix is required. If the promotion volume is operationally problematic, reconsider whether the ephemeral-creation workload belongs on ZooKeeper at all.

maxClientCnxns saturation

If zk_connection_rejected is non-zero, clients are being silently denied service because their source IP has hit the per-IP connection limit. This is not strictly a session leak, but it produces symptoms (clients cannot connect, applications spawn new connection attempts that also fail) that look like one. Raise maxClientCnxns in zoo.cfg (default 60 per source IP in all 3.5/3.6/3.7/3.8 releases; set to 0 to remove the limit) and rolling-restart. In containerized environments where many pods share a host IP, the default 60 is easily exceeded.

Prevention

  • Track the session-to-connection ratio over time. It should be stable. Alert on sustained drift, not on absolute session count.
  • Cap and monitor client-side connection pools. Each ZooKeeper handle is a session. Frameworks that let application code spawn handles freely will eventually leak.
  • Pin client framework versions and watch their issue trackers. Curator and similar frameworks have shipped session-leak bugs more than once.
  • Enable GC logging on ZooKeeper clients, not just the server. Client-side GC pauses are a leading root cause of session churn.
  • Run a periodic chaos exercise that disconnects a fraction of clients and watches the ensemble’s session table return to baseline. If it does not return, you have a leak; find it before it finds you at 3 a.m.
  • Set maxClientCnxns deliberately based on your container-to-host-IP fan-in ratio. The default 60 is too low for most microservice deployments that fan many pods through one host IP.

How Netdata helps

  • Per-second collection of zk_global_sessions, zk_num_alive_connections, and zk_ephemerals_count makes the monotonic-growth signature visible within minutes rather than after a daily aggregation rolls over.
  • ML anomaly detection flags session-to-connection ratio drift even when both metrics are individually within normal absolute ranges.
  • Correlating zk_global_sessions with JVM pause, zk_stale_sessions_expired, and zk_connection_drop_count in a single view distinguishes a server-side GC cascade from a client-side reconnect loop without manual cross-referencing.
  • Composite dashboards surface leader-only metrics (zk_followers, zk_synced_followers) alongside connection metrics, so quorum degradation can be ruled out as a confounding factor.
  • The same per-second resolution applies to host-level signals (FD usage, disk latency, network retransmits) that often explain why a client is reconnecting in the first place.