The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-server-stuck-looking

Operations Guides

ZooKeeper server stuck in LOOKING: a node that never rejoins the quorum

A single ZooKeeper node sits in LOOKING long after the ensemble has settled, or flaps between LOOKING and FOLLOWING. The rest of the ensemble holds a stable leader with healthy write throughput. Restart the orphaned node and it drops back into LOOKING. Restart the leader and it briefly rejoins, then falls out again.

This is not quorum loss. Quorum loss is every node entering LOOKING at once because no majority can form. That case is covered in ZooKeeper quorum loss: no leader elected and every write is failing. This article covers the narrower symptom: one permanently-orphaned node while the rest of the ensemble serves traffic.

The causes cluster into three themes. The node cannot reach its peers on the election port (3888) or quorum port (2888) in the direction that matters. The node’s local state is incompatible with the leader (zxid ahead of the leader’s epoch, or missing epoch files). Or the node’s local disk and JVM cannot persist a vote fast enough to survive an election round. Distinguishing among these three is the whole job, and the rest of this article walks through how to do it without making the situation worse.

What this means

A ZooKeeper ensemble member participates in two distinct TCP planes. The quorum port (2888 by default, the first port in server.X=host:2888:3888) carries follower-leader traffic: PROPOSE, ACK, COMMIT. The leader election port (3888, the second port) carries FastLeaderElection notifications during leader election. Both must be reachable, in both directions, between every pair of ensemble members. A firewall that blocks 3888 in one direction is invisible until the next election. At that point the affected node cannot send or receive votes and gets stuck in LOOKING indefinitely.

When a node enters LOOKING, it opens TCP connections to every peer on 3888 and exchanges vote notifications. To leave LOOKING it must receive a valid notification from a quorum of peers establishing a leader. If 3888 is reachable from peers but not from this node, or vice versa, the vote never completes. The same applies to the subsequent sync phase on 2888. A node that elected a leader but cannot reach it on 2888 falls back into LOOKING and the cycle repeats.

This produces the characteristic “flapping” signature. zk_looking_count on the affected node increments steadily while zk_looking_count on every other node is flat. The rest of the ensemble holds a stable leader in BROADCAST. From the leader’s point of view, zk_synced_followers is one short of ensemble_size - 1.

flowchart TD
    A[One node stuck in LOOKING] --> B{Other nodes have
stable leader?} B -- No --> C[Quorum loss:
different playbook] B -- Yes --> D{Can this node reach
peers on 3888?} D -- No --> E[Election port blocked
or DNS to loopback] D -- Yes --> F{zxid comparable
to leader?} F -- Far behind --> G[Long catch-up
or SNAP sync] F -- Ahead of leader --> H[Unrecoverable:
rebuild from leader] F -- Comparable --> I{fsync or GC
healthy on this node?} I -- No --> J[Disk or JVM stalls
preventing vote persistence] I -- Yes --> K[Known bug pattern
ZOOKEEPER-2938]

Common causes

CauseWhat it looks likeFirst thing to check
Election port (3888) unreachable from this nodeNode completes startup but cannot send or receive vote notifications; log shows repeated LOOKING entries without resolutionBidirectional TCP reachability on 3888 between this node and every peer
DNS resolving own hostname to loopback (containers)Node binds 3888 to 127.0.0.1; peers cannot reach it; common in Kubernetes and Dockergetent hosts <this-node-name> and the server.X= line in zoo.cfg
zxid ahead of leader’s epochLog shows “Got zxid 0x… expected 0x…” with ClosedChannelException; node cannot syncCompare zk_zxid across all ensemble members
Long catch-up after extended downtimeNode elected a leader but stuck in synchronization; large snapshot transfer in progresszk_zxid of affected node slowly advancing toward leader’s
Local fsync or GC stalls prevent vote persistencezk_fsynctime p99 spiking or zk_jvm_pause_time_ms p99 approaching tickTime; node falls out of election roundszk_fsynctime and zk_jvm_pause_time_ms on the affected node
“Have smaller server identifier” drop patternLeader drops reconnecting follower’s connection because of lower server ID; unrecoverable without interventionZooKeeper log for the literal string

Quick checks

Run these read-only commands before changing anything.

# This node's state. Should report follower (or leader). LOOKING means election in progress.
echo srvr | nc localhost 2181 | grep Mode

# Functional state. rw means serving writes; ro means read-only mode (quorum lost).
echo isro | nc localhost 2181

# Election and state counters on this node.
echo mntr | nc localhost 2181 | grep -E 'zk_server_state|zk_looking_count|zk_uptime|zk_zxid'

# Same metrics from the leader's perspective.
echo mntr | nc <leader-host> 2181 | grep -E 'zk_server_state|zk_followers|zk_synced_followers|zk_pending_syncs'

# Compare zxid across every ensemble member. They should match within a few transactions.
for h in zk1 zk2 zk3; do
  printf '%s ' "$h"
  echo mntr | nc "$h" 2181 | grep zk_zxid
done

# TCP reachability from this node to every peer on both inter-server ports.
for h in zk1 zk2 zk3; do
  for p in 2888 3888; do
    timeout 2 bash -c "exec 3<>/dev/tcp/$h/$p" 2>/dev/null && echo "$h:$p open" || echo "$h:$p blocked"
  done
done

# What the affected node is logging right now.
tail -n 200 /var/log/zookeeper/zookeeper.log | grep -E 'LOOKING|FOLLOWING|LEADING|smaller server|Got zxid|Not following'

# Cold-start suppression check. If uptime is small, give the node time to settle before diagnosing.
echo mntr | nc localhost 2181 | grep zk_uptime

Replace zk1 zk2 zk3 and /var/log/zookeeper/zookeeper.log with your actual hostnames and log path. The four-letter commands require 4lw.commands.whitelist to include srvr, mntr, and isro (ZooKeeper 3.5.3+). On 3.6+, the AdminServer on port 8080 exposes the same data over HTTP, for example /commands/server_stats and /commands/leader.

How to diagnose it

  1. Confirm only one node is affected. Check zk_server_state on every ensemble member. If two or more nodes are in LOOKING, you have a different problem; see the quorum-loss guide. This article applies only when exactly one node is stuck and the rest have a stable leader.

  2. Check whether this is a cold start. If zk_uptime on the affected node is below 300 seconds, suppress investigation for a few minutes. Large data trees take time to load from snapshot plus transaction log replay. The node can look stuck while it is actually recovering.

  3. Verify bidirectional election-port reachability. Run the TCP reachability loop in Quick checks from this node to every peer on 3888, then run the same check from one peer back to this node. Asymmetric blocking is the most common cause of a single stuck node. A firewall rule that allows outbound 3888 but blocks inbound 3888 (or vice versa) is invisible until election.

  4. Check whether DNS resolves this node’s hostname to a loopback address. In containerized deployments (Kubernetes, Docker, Strimzi), getent hosts <this-node-name> may return 127.0.0.1, which causes the node to bind 3888 to loopback and become unreachable from peers. The server.X= entry in zoo.cfg for the local node should use a routable IP, or 0.0.0.0:2888:3888 to bind to all interfaces. The 0.0.0.0 local entry only affects the local bind and still works on 3.8.x/3.9.x; it does not break quorum TLS as long as the peer entries (the addresses other members use to reach this node) remain hostnames, because ssl.quorum.hostnameVerification (new in 3.5.5, on by default) validates the connecting side’s hostname against the certificate.

  5. Compare zxid across the ensemble. A node with a zxid significantly behind the leader is in long catch-up. If the leader must send a full snapshot (SNAP sync), this can take minutes for a large data tree. A node with a zxid ahead of the leader’s epoch is in the unrecoverable “Got zxid X expected Y” state and must be rebuilt from the leader.

  6. Check local fsync and GC. Look at zk_fsynctime p99 and zk_jvm_pause_time_ms p99 on the affected node. Election rounds have timeouts governed by tickTime (default 2000ms). If fsync or GC stalls push vote persistence past that boundary, the node cannot complete an election round and re-enters LOOKING.

  7. Inspect the log for known failure signatures. The string “Have smaller server identifier, so dropping the connection” indicates the ZOOKEEPER-2938 pattern where the leader drops a reconnecting follower with a lower server ID. “Got zxid 0x… expected 0x…” with ClosedChannelException indicates the unrecoverable zxid-ahead state. “currentEpoch not found!” indicates the epoch file is missing from the data directory.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_server_state per nodeTells you which nodes are leader, follower, or LOOKINGOne node anything other than leader/follower while the rest are stable
zk_looking_count per nodeCounts entries into leader electionIncrements on one node only, flat on peers
zk_zxid per nodeReplication positionDiverges from leader, either far behind (catch-up) or ahead (unrecoverable)
zk_followers and zk_synced_followers on leaderCount of voting members in syncsynced_followers < ensemble_size - 1 for more than a minute outside maintenance
zk_fsynctime p99 on this nodeTime to fsync transaction logSpikes preceding LOOKING transitions
zk_jvm_pause_time_ms p99 on this nodeGC stop-the-world eventsp99 approaching tickTime (default 2000ms)
zk_uptime per nodeDetects crashes and restartsUnexpected reset
zk_follower_sync_time on leaderTime for followers to syncSustained elevation indicates a follower is struggling to catch up

zk_followers, zk_synced_followers, and zk_pending_syncs are leader-only metrics. A monitoring setup that scrapes only followers will never see them. Query the leader, or query all nodes and filter for the one reporting leader state.

Fixes

Election port (3888) blocked from this node

Verify firewall, security group, and iptables rules allow 3888 bidirectionally between this node and every peer. Test in both directions, not just outbound from the affected node. Do not restart ZooKeeper until reachability is verified; the restart will not help and will lose diagnostic state in the logs. Once connectivity is fixed, the next election round (initiated by the node itself or by a planned leader restart) should let it rejoin.

DNS resolving to loopback in containers

In Kubernetes, Docker, or any environment where DNS may resolve a pod’s own hostname to 127.0.0.1, set the local server entry to bind all interfaces:

server.1=0.0.0.0:2888:3888
server.2=zk2.example.com:2888:3888
server.3=zk3.example.com:2888:3888

Alternatively, use static IPs in every server.X= line. The 0.0.0.0 form binds to all interfaces on that node only; peers continue to use their routable addresses. Verify the bind with ss -lntp | grep 3888 after restart.

zxid ahead of leader (unrecoverable)

This state is unrecoverable without intervention. The node’s local transaction log has a zxid incompatible with the leader’s epoch, and ZAB will not allow it to rejoin.

Destructive operation. This procedure discards the local copy of the data tree. It is safe only because the node is already not serving traffic.

  1. Stop ZooKeeper on the affected node only.
  2. Back up the data directory (the contents of dataDir, typically version-2/).
  3. Remove the contents of the data directory on the affected node, including currentEpoch, acceptedEpoch, and all log.* and snapshot.* files.
  4. Preserve the myid file (dataDir/myid). It identifies this server’s position in the ensemble and must not be deleted.
  5. Restart ZooKeeper. The node performs a SNAP sync from the leader, receiving a full snapshot.
  6. Verify zk_server_state returns follower and zk_zxid matches the leader.

Long catch-up after extended downtime

If zk_zxid on the affected node is advancing, just slowly, this is a SNAP sync in progress. Do not interrupt it. Watch zk_follower_sync_time on the leader; once it returns to baseline, the node should report follower. If the leader is overloaded by the snapshot transfer and client traffic is affected, reduce client load on the leader temporarily.

Local fsync or GC stalls

See the playbook’s Disk Sync Deadlock and GC Death Spiral composite patterns. The short version: move the transaction log to its own dedicated disk (dataLogDir on a separate volume, not shared with snapshots or other workloads), enable GC logging (-Xlog:gc*:file=/var/log/zookeeper/gc.log:time,uptime,level,tags:filecount=5,filesize=100m), and size the heap so that the post-GC trough stays below roughly 50% of max. A node that cannot persist a vote within tickTime cannot complete an election round.

ZOOKEEPER-2938 (“Have smaller server identifier”)

This is a known bug pattern where the leader drops the reconnecting follower’s connection because the follower has a lower server ID. ZOOKEEPER-2938 remains open without an assigned fix version as of the last JIRA update, including on the 3.9.x line. The temporary workaround is to restart the leader so a fresh election round runs cleanly. The longer-term mitigation is to upgrade to the newest practical release and watch the JIRA for a fix.

Prevention

  • Alert on zk_server_state per node, not per ensemble. Aggregating state across nodes hides single-node LOOKING. Each node must be alertable independently.
  • Alert on zk_looking_count rate per node. Any sustained increment outside a planned maintenance window is a ticket. The threshold is rate, not absolute value.
  • Verify symmetric firewall rules for 2888 and 3888 between every pair of ensemble members, in both directions. Add this to provisioning checks. Cloud security group changes are a recurring cause of asymmetric blocking.
  • In containers, use static IPs or 0.0.0.0 binding for the local server.X= entry. Do not rely on DNS resolving pod hostnames to routable addresses.
  • Put the transaction log on a dedicated disk. Set dataLogDir to a separate volume from dataDir. This is the single highest-impact configuration change for write-path stability and the most common cause of fsync-induced LOOKING flapping.
  • Compare zk_zxid across the ensemble as a scheduled check. Divergence is the earliest signal of a node about to fall out.
  • Suppress non-critical alerts when zk_uptime < 300 seconds. Cold-start LOOKING is normal during snapshot and log replay and should not page.

How Netdata helps

  • Per-second zk_server_state per node surfaces single-node LOOKING quickly, before it becomes a long-running incident.
  • Per-node correlation of zk_looking_count increments with zk_fsynctime p99 and zk_jvm_pause_time_ms p99 distinguishes network causes from local disk and JVM causes without manual log scraping.
  • Side-by-side zk_zxid across ensemble members shows divergence early, including the slow drift of a node in long catch-up versus the abrupt ahead-of-leader signature.
  • Leader-only metrics (zk_followers, zk_synced_followers, zk_pending_syncs) are collected automatically and shown in the leader context, so a missing follower is visible in one view.
  • ML anomaly detection on zk_looking_count and zk_fsynctime p99 catches the slow trend that precedes a stuck election, even when absolute values still look normal.