The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-noauth-error

Operations Guides

ZooKeeper KeeperErrorCode = NoAuth: ACL denials on protected znodes

KeeperErrorCode = NoAuth for /path appears in client logs when a ZooKeeper operation is rejected because the calling session lacks the ACL permission required for that operation on that znode. ZooKeeper does not log the denial server-side by default; to record it, enable audit logging (audit.enable=true, ZooKeeper 3.6+). NoAuth is not a transient connectivity issue. The request reached a server, the server evaluated the znode’s ACL, and the session did not match.

ZooKeeper always enforces ACLs. A fresh ensemble uses OPEN_ACL_UNSAFE (world:anyone with CREATE, READ, WRITE, DELETE, ADMIN). NoAuth only appears after someone has explicitly set a restrictive ACL on a znode using one of the schemes digest, sasl, ip, auth, or x509. The error implies a protected znode exists and a client reached it without the matching credential.

The usual causes are a misconfigured or missing SASL/Digest credential, a client that never called addAuthInfo before issuing operations, a credential rotation that left one client behind, or an unauthorized access attempt. Repeated NoAuth from a single source warrants investigation as either a misconfiguration or a probe.

What this means

ZooKeeper evaluates ACLs per operation. Each znode stores an ACL: a list of (scheme:expression, permissions) entries. When a client performs an operation, the server walks the ACL and allows the operation only if some entry matches the client’s authenticated identity and grants the required permission. The permission bits are CREATE, READ, WRITE, DELETE, and ADMIN. getACL requires READ or ADMIN (fixed in 3.4.14 and 3.5.5, per ZOOKEEPER-1392; present in 3.6.0 and later). setACL requires ADMIN.

Facts that bite operators:

  • ACLs are not recursive. A parent’s ACL does not protect its children. Each znode carries its own ACL, set at creation time. A common cause of NoAuth is assuming an ACL set on a parent propagates to children.
  • The auth scheme is a wildcard for “identities this session has authenticated as.” Creating a node with an auth-scheme ACL before the creating session has called addAuthInfo fails with InvalidACL. If a node ends up with an empty or misapplied ACL, no client will match.
  • The world scheme with anyone is the permissive default. Once a node is created with a restrictive ACL, only matching identities or the super user can change it.
  • exists() now checks READ in versions 3.9.2, 3.8.4, and 3.7.3 and later (ZOOKEEPER-2590). Clients that previously probed for node existence without READ permission now receive a permission error where they previously received a silent ok.

The error surface differs by client. The Java client throws org.apache.zookeeper.KeeperException$NoAuthException: KeeperErrorCode = NoAuth for /path. The CLI (zkCli.sh) reports KeeperError = NoAuth, or, prior to ZOOKEEPER-3891 (fixed in 3.7.0), a misleading Authentication is not valid message for insufficient permissions. The server does not log the denial by default; with audit logging enabled (3.6+), the audit log records it. All are the same event.

flowchart TD
  A["NoAuth on /path"] --> B{"Client called addAuthInfo?"}
  B -- No --> C["Add addAuthInfo before
protected operations"] B -- Yes --> D{"Identity matches ACL?"} D -- No --> E["Fix digest password
or SASL principal"] D -- Yes --> F{"Upgraded to
3.9.2 / 3.8.4 / 3.7.3?"} F -- Yes --> G["exists now needs READ.
Grant it or stop probing."] F -- No --> H["Source IP expected?
If not, treat as probe."]

Common causes

CauseWhat it looks likeFirst thing to check
Client never called addAuthInfoFirst operation on a protected znode fails immediately with NoAuth; same client works on world-readable nodesInspect the connection path in the client. Confirm addAuthInfo is called before any create, getData, or setData on protected paths
Digest credential mismatchNoAuth appears after a password rotation; only some clients affected; zk_auth_failed_count stays 0 if addAuthInfo was omitted, but increments if an invalid digest was sentDiff the password on the client against the password used to set the ACL; recompute the SHA-1 hash
SASL/Kerberos mismatchNoAuth tied to a specific principal; ticket renewal failures in client logs; possibly the ZOOKEEPER-4885 pattern of a non-SASL client used after Kerberos recoveryCheck JAAS config and keytab principal; compare against the ACL expression; run klist on the client host
ACL set with auth before authenticationNode created with a missing or wrong ACL; intended owner cannot operate on itgetACL /path (as super user if needed) and inspect the actual entries
Version-driven exists() changeNoAuth or Insufficient permission on exists() probes after upgrade to 3.9.2/3.8.4/3.7.3+; previously workedConfirm server version via srvr; identify clients probing without READ
Unauthorized access attemptNoAuth from source IPs outside the client inventory; bursts from a single sourceCross-reference source IP against known clients; check connection logs

Quick checks

These are read-only. In ZooKeeper 3.5.3+, four-letter commands must be whitelisted via 4lw.commands.whitelist in zoo.cfg. If mntr returns nothing, whitelisting is the first thing to fix.

# The server does not log ACL denials by default (client-side NoAuth is the signal).
# With audit logging enabled (3.6+, audit.enable=true), read the audit log instead:
grep -hE "result=failure" /var/log/zookeeper/zookeeper_audit.log | tail -30

# Failed authentication handshakes (SASL/digest). Distinct from ACL denials.
echo mntr | nc localhost 2181 | grep zk_auth_failed_count

# Confirm the server is healthy enough to be the source of truth
echo ruok | nc localhost 2181
echo isro  | nc localhost 2181

# Server version. Determines whether the ZOOKEEPER-2590 exists() change applies.
echo srvr | nc localhost 2181 | grep -E "Zookeeper version|Mode"

# Connected clients and source IPs. Expensive on busy servers; use sparingly.
echo cons | nc localhost 2181 | head -40

The log path differs by distribution. /var/log/zookeeper/zookeeper.log is the common default; check your service definition if that file is absent.

How to diagnose it

  1. Confirm the error is NoAuth. Confirm the client log shows KeeperErrorCode = NoAuth for /path. The server does not log the denial by default; if audit logging is enabled (3.6+), the audit log records the failed operation with the user, path, and session. If the client shows ConnectionLoss or SessionExpired, this is a different failure.

  2. Read the ACL on the znode. Run getACL /path from a session that has READ or ADMIN on the node. If no live client has access, use the super user (see Fixes) or read the ACL from a snapshot off-line. The ACL tells you which scheme and expression the client must match.

  3. Identify the client’s authenticated identity. For digest, the identity is user:<base64(SHA1(user:password))>. For sasl, it is the Kerberos principal. For x509, it is the client certificate subject. Compare this identity against the ACL expression exactly.

  4. Check whether addAuthInfo was called before the failing operation. In the Java client this is zoo.addAuthInfo("digest", "user:password".getBytes()) or the SASL equivalent at connection. A common bug is creating the ZooKeeper handle and immediately issuing operations before authentication completes.

  5. Check the server version. If you recently upgraded to 3.9.2, 3.8.4, or 3.7.3, exists() now checks READ. Clients that worked before by probing without READ will now fail. This is a behavior change, not a misconfiguration.

  6. Check for a Kerberos event correlation. With SASL/Kerberos, look for ticket renewal failures in client logs. ZOOKEEPER-4885 (open as of the 3.9.3 timeframe) describes a case where a non-SASL client is created after a Kerberos failure and never recovers, producing persistent NoAuth even after Kerberos is healthy again. The signal is NoAuth from a single client that began at the same timestamp as a renewal failure.

  7. Check the source distribution. If NoAuth comes from one host or one client identity repeatedly, it is almost certainly a misconfiguration or a probe. If it affects many clients at once, suspect a credential rotation or an ACL change.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
zk_auth_failed_count (mntr)Counts failed authentication handshakes. Distinct from ACL denials, but the two co-occur when the root cause is a wrong digest passwordNon-zero rate after a credential rotation, or sustained rate from a single source
Audit log result=failure entries (3.6+)Server-side record of failed operations when audit logging is enabledSustained or bursty rate from one source IP
cons four-letter outputPer-connection source IP and session ID. Maps a log entry back to a clientUnexpected source IPs appearing in the connection list
Server version (from srvr)Determines whether the ZOOKEEPER-2590 exists() ACL change appliesRecently upgraded to 3.9.2, 3.8.4, or 3.7.3, or later
zk_ensemble_auth_fail (mntr)Server-to-server authentication failures. A different problem (quorum auth), but easy to confuse with client NoAuthAny non-zero increment
zk_non_mtls_remote_conn_count (mntr)In x509-based deployments, counts non-mTLS remote connectionsNon-zero in an mTLS-required environment

Fixes

Client never called addAuthInfo

Call addAuthInfo on the ZooKeeper handle before issuing any operation against a protected znode:

ZooKeeper zk = new ZooKeeper(connectString, sessionTimeout, watcher);
zk.addAuthInfo("digest", "user:password".getBytes(StandardCharsets.UTF_8));
// now safe to operate on ACL-protected nodes

addAuthInfo is asynchronous in the Java client. The credential is sent to the server and the server responds. Operations issued before the response may still fail with NoAuth. If you need a synchronous guarantee, issue a no-op such as exists on a world-readable node and wait for it to succeed before issuing protected operations.

Digest credential mismatch

If the ACL was set with digest:user:<sha1> and the client presents user:wrongpassword, correct the password on the client. The hash is base64(SHA1(user:password)). Compute it locally to verify:

# Compute the digest auth hash for verification
echo -n "user:password" | openssl dgst -sha1 -binary | base64

Compare against the ACL expression visible via getACL /path. If you cannot recover the original password, your options narrow:

  • Reset the ACL using the super user (see below). Requires zookeeper.DigestAuthenticationProvider.superDigest to have been configured on the server.
  • Reconstruct the node: delete and recreate with the correct ACL. This is destructive. Ephemeral children vanish, watches fire, and any data on the node is gone. Do this only outside peak traffic and only after checking what depends on the node.

SASL/Kerberos mismatch

The ACL expression for SASL is the Kerberos principal (for example, zkclient@EXAMPLE.COM). Ensure the client’s JAAS config presents the matching principal. Common pitfalls:

  • The JAAS config file is not loaded (java.security.auth.login.config system property missing or pointing at the wrong file).
  • The principal in the keytab does not match the principal in the ACL.
  • The ticket has expired and renewal failed. ZOOKEEPER-4477 (fixed in 3.8.1) addressed a case where a single renewal failure prevented all future renewals on Java 9 and later.
  • ZOOKEEPER-4885 (open as of 3.9.3): after a Kerberos failure, the client falls back to a non-SASL client and never recovers, producing persistent NoAuth even after Kerberos recovers. The workaround is to recreate the ZooKeeper handle on the client side.

Rebuilding the ZooKeeper handle drops the session, which clears ephemeral nodes and watches. Do this only after confirming the client is stuck in the ZOOKEEPER-4885 pattern.

ACL set with auth before authentication

The auth scheme expands to “all identities this session has authenticated as.” If the node was created with an auth-scheme ACL before the creating session called addAuthInfo, the server rejects it with InvalidACL. If an empty or unmatchable ACL bypassed validation, the fix is to recreate the node with the correct ACL, or to use the super user to run setACL.

Version-driven exists() change

This is not a misconfiguration. If you upgraded to 3.9.2, 3.8.4, or 3.7.3 and clients now fail exists() where they previously succeeded, you have two options:

  • Grant READ on the probed nodes to the probing identity.
  • Change the client to not depend on exists() returning ok without READ.

Do not work around this with zookeeper.skipACL=yes. That disables all ACL checking on the server and removes a security control. The exists() change closes a probe that should never have been permitted.

Unauthorized access attempt

If the source IP is not in your client inventory, the NoAuth is doing its job. The fix is to prevent the source from reaching the ZooKeeper port: network policy, firewall rules, or removal of the pod or job. In parallel, audit your ACL posture. Confirm that sensitive subtrees (Kafka assignments, HBase region state, distributed lock nodes) carry restrictive ACLs, because the default world:anyone makes them readable and writable by any client that can reach the port.

Super user recovery

For any cause that requires changing an ACL you cannot reach, configure the super user escape hatch on the server:

-Dzookeeper.DigestAuthenticationProvider.superDigest=super:base64(SHA1(super:password))

Then connect as addauth digest super:password and bypass all ACL checks. This is the only recovery path for a node whose ACL has locked out every legitimate client. Configure it before you need it, and store the password in your secrets manager.

Prevention

  • Wire addAuthInfo into the connection path. Every client should call it as part of establishing the handle, before any application-level operation.
  • Set restrictive ACLs at creation time. Create znodes with explicit (scheme:expression, permissions) rather than relying on defaults.
  • Remember ACLs are not recursive. Children do not inherit the parent’s ACL. Protect each child explicitly, or set the creating client’s default ACL.
  • Configure superDigest before you need it. Document the password in your secrets store. You will need it during a lockout.
  • Track server version in monitoring. The exists() change in 3.9.2/3.8.4/3.7.3 is breaking. Know when you cross that boundary.
  • Keep audit logging on for sensitive subtrees. audit.enable=true in zoo.cfg (3.6+) logs mutations with the session and identity. NoAuth denials themselves are already in the server log.
  • Do not use skipACL in production. zookeeper.skipACL=yes disables all ACL checking. It is a debug escape hatch, not a configuration.

How Netdata helps

  • Correlate zk_auth_failed_count with server log spikes. Authentication handshake failures often precede or accompany NoAuth storms. Seeing both move together in one view narrows the cause from “client never authenticated” to “client authenticated as the wrong identity.”
  • Map connection source IPs to NoAuth bursts. Per-second collection on zk_num_alive_connections and the connection rejected counter lets you spot a single source IP producing a burst, the signature of a misconfigured rollout or a probe.
  • Track server version across the ensemble. Knowing which nodes are on 3.9.2 or later tells you immediately whether the exists() ACL change is in play, without grepping release notes mid-incident.
  • Surface ensemble-wide auth failures. zk_ensemble_auth_fail increments on server-to-server authentication failures, a different problem (quorum auth) but easy to confuse with client NoAuth. Having both on one dashboard prevents misdiagnosis.
  • Catch the Kerberos recovery trap. Sustained NoAuth from a single client that began at the same timestamp as a Kerberos ticket renewal failure matches the ZOOKEEPER-4885 pattern. High-resolution collection makes that timestamp correlation visible.