The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / zookeeper / zookeeper-no-authentication-open-access

Operations Guides

ZooKeeper with no authentication: the open-by-default coordination store

ZooKeeper ships with no authentication by default. Any TCP client that can reach port 2181 can open a session, read any znode, and write any znode whose ACL has not been explicitly restricted. The default ACL on znodes created without an explicit ACL is OPEN_ACL_UNSAFE: world:anyone with full cdrwa (create, read, write, delete, admin) permissions.

This matters because ZooKeeper is the coordination store for systems that treat its contents as authoritative: Kafka broker registrations and controller elections (pre-KRaft), HBase region assignment and master election, HDFS NameNode HA fencing state. An unauthenticated writer in those subtrees can silently corrupt cluster state, force leader changes, or trigger cascading failovers, and the writes succeed because the ACL permits them. There is no second layer that catches them.

The lockdown order matters: inventory first, then audit logging, then SASL authentication, then per-znode ACLs, then network segmentation. Skipping ahead to “turn on SASL” without an inventory is how teams discover, mid-rollout, that a critical client they forgot about can no longer connect.

What this enables

  • A documented inventory of every IP currently connected to 2181, so you can spot unexpected clients before you flip on auth and break them.
  • An audit trail of every create, setData, delete, setACL, and reconfig operation against critical znodes, with the client session and identity, for post-incident forensics.
  • A path from OPEN_ACL_UNSAFE to authenticated, per-znode ACL-controlled access, where credentials you actually trust are required to mutate coordination data.
  • Defense in depth: if one control fails (a leaked digest, a misconfigured security group, an upstream CVE), the others still constrain the blast radius.

Each layer fails independently and logs differently:

flowchart TD
  A[Client reaches 2181] --> B{Network ACL
allows source IP?} B -- no --> X[Blocked at firewall] B -- yes --> C{SASL auth
valid credentials?} C -- no --> Y[Session rejected] C -- yes --> D{Per-znode ACL
permits operation?} D -- no --> Z[NoAuth logged] D -- yes --> E[Mutation applied] E --> F[Audit log records
user, IP, path]

Prerequisites

  • ZooKeeper 3.6.x or newer. Audit logging and sessionRequireClientSASLAuth both require 3.6+. Current stable releases are 3.8.6 and 3.9.5.
  • Shell access to each ensemble node, with permission to edit zoo.cfg and rolling-restart the QuorumPeerMain process.
  • A known-client inventory: the list of IPs and CIDRs that should be connecting. Without it, the connection log is noise.
  • A rolling-restart maintenance window once authentication settings change. Existing clients using the open posture will be rejected when you enforce SASL.
  • Familiarity with the difference between client port 2181, AdminServer port 8080 (3.5+, on by default), and the inter-ensemble quorum and election ports.

Procedure

1. Confirm the current exposure

Read-only, safe to run on production.

# Check zoo.cfg path on your distro; common locations:
# /etc/zookeeper/conf/zoo.cfg  (Debian/Ubuntu)
# /etc/zookeeper/zoo.cfg       (RHEL/CentOS)
# /opt/zookeeper/conf/zoo.cfg  (tarball installs)

# Confirm no auth is configured on the client port
grep -E 'authProvider|sessionRequireClientSASLAuth|kerberos|sasl' /etc/zookeeper/conf/zoo.cfg
# Default install: returns nothing

# Confirm audit is disabled
grep -E 'audit.enable|audit.log.dir' /etc/zookeeper/conf/zoo.cfg
# Default install: returns nothing

# Confirm the four-letter-word surface (3.5.3+)
grep '4lw.commands.whitelist' /etc/zookeeper/conf/zoo.cfg
# Default: only 'srvr' is whitelisted

If the SASL/authProvider grep returns nothing and 2181 is reachable from outside the deployment network, you are running open by default.

2. Inventory who is connecting today

Read-only, but the Accepted socket connection log (logged at DEBUG level by NIOServerCnxnFactory) is verbose. It only appears in the server log if DEBUG logging is enabled; sample it, do not tail indefinitely on a busy node.

# Distinct source IPs recently connected, from the ZK server log.
# The line ends with "/<ip>:<port>" - strip the slash and port.
grep 'Accepted socket connection' /var/log/zookeeper/zookeeper.log \
  | sed -E 's#.*from /?([^: ]+):[0-9]+.*#\1#' | sort -u

# Live connection snapshot via four-letter-word 'cons'
# (must be in 4lw.commands.whitelist; expensive on busy nodes)
echo cons | nc localhost 2181 | head

Diff the source IPs against your known-client inventory. Anything outside the inventory is the starting point for investigation. Run cons once, off-peak; it is O(n) in connection count and will perturb latency on a loaded node.

3. Turn on audit logging (3.6+)

Edit zoo.cfg on each node, then rolling-restart:

audit.enable=true
audit.log.dir=/var/log/zookeeper/audit

After restart, every mutation (create, setData, delete, setACL, multi, reconfig) is written to zookeeper_audit.log with session, user (once SASL is on), and znode path. Reads are not audited by default.

Verify:

# WARNING: this writes to production ZK. Use a throwaway test path and clean up.
zkCli.sh -server localhost:2181 create /audit-test "hello"
tail -n 5 /var/log/zookeeper/audit/zookeeper_audit.log

# Clean up
zkCli.sh -server localhost:2181 delete /audit-test

Leave audit logging on from this point forward. It is the only forensic record you will have once you start locking things down, and the only way to answer “who mutated /kafka/controller?” after the fact.

4. Stand up SASL authentication

This is the disruptive step. Stage it as two flips, not one.

The built-in digest scheme (username/password, SHA-1 hashed) is the lowest-friction option. For each client identity, generate user:password, hash per the digest scheme, and configure ZooKeeper to load a JAAS file. Kerberos, SASL/PLAIN, and mTLS are alternatives; pick based on your existing identity infrastructure.

In zoo.cfg:

authProvider.1=org.apache.zookeeper.server.auth.SASLAuthenticationProvider
sessionRequireClientSASLAuth=true

authProvider.1 enables SASL handling on the server. sessionRequireClientSASLAuth=true (3.6+) forces all client sessions to authenticate; any client that does not present valid SASL credentials is rejected. Roll out authProvider first, confirm no production client is being rejected via zk_auth_failed_count, then flip sessionRequireClientSASLAuth.

Caveats from the upstream security guidance:

  • Digest transmits the password in the clear and stores an unsalted SHA-1 hash. Use SASL/Kerberos or mTLS where credential confidentiality matters.
  • Quorum peer auth (quorum.auth.enableSasl=true) is a separate control and is not enabled by default. CVE-2023-44981 (authorization bypass in SASL quorum peer authentication) affected this path through 3.9.0/3.8.2/3.7.1; if you enable it, run 3.9.1+, 3.8.3+, or 3.7.2.
  • The ip ACL scheme is spoofable. It trusts source IP. Use it only as a secondary check.

5. Apply per-znode ACLs

ACLs in ZooKeeper are not recursive. Each znode must have its ACL set independently. The default OPEN_ACL_UNSAFE (world:anyone:cdrwa) is permissive; lockdown means explicitly setting auth: or digest:user: ACLs on each critical subtree.

For the major ecosystems:

  • Kafka (ZooKeeper mode): set zookeeper.set.acl=true on brokers so the broker applies ACLs to its znodes at creation. Without this, broker-registered znodes remain world-writable even after SASL is on. Kafka is migrating to KRaft; for ZooKeeper-mode clusters this is still the documented lever (Kafka’s ZooKeeper Authentication docs).
  • HBase and YARN: these do not set ACLs for their znodes by default. Locking them down requires a tool like zkpolicy or an explicit per-znode ACL walk after SASL is enabled.
# WARNING: setAcl on production coordination paths is disruptive.
# Test on a non-production node first. Locking /kafka before brokers
# have working SASL credentials will break the cluster.
zkCli.sh -server localhost:2181
# Inside the CLI:
setAcl /kafka auth:cdrwa
setAcl /hbase auth:cdrwa
getAcl /kafka

Children must be walked and set individually. A parent ACL does not flow down, which is the single most common reason a “locked-down” ensemble is still writable by an anonymous client.

6. Tighten the surface around the client and admin ports

  • Four-letter-word whitelist (3.5.3+): explicitly list only what monitoring needs (mntr, ruok, isro, srvr). Do not whitelist cons, wchc, wchp, dump, or envi in production. They expose connection and watch detail and are O(n) expensive.
  • AdminServer (3.5+, port 8080, on by default): bind to localhost or an internal interface, and put it behind network policy. On 3.9.0+, consider admin.snapshot.enabled=false and admin.restore.enabled=false (both default to true) unless you actively use those endpoints; the snapshot/restore commands were the subject of CVE-2025-58457 (fixed in 3.9.4), and AdminServer IP-based auth was bypassable via spoofed X-Forwarded-For in CVE-2024-51504 (fixed in 3.9.3).
  • maxClientCnxns (default 60 per source IP) is a DoS guard, not an auth control. Do not raise it to work around an auth change.

7. Network segmentation

Port 2181 should be reachable only from known client CIDRs and the load balancer you control. Port 8080 (AdminServer) and the quorum and election ports should never be reachable outside the ensemble subnet. Apply this at the security-group or firewall layer, not only at ZooKeeper. Adjacent endpoints have had bypasses (CVE-2024-23944 persistent watchers, fixed in 3.9.2/3.8.4; CVE-2024-51504 AdminServer, fixed in 3.9.3; CVE-2026-24281 reverse-DNS hostname verification, fixed in 3.8.6/3.9.5), so do not treat any single control as sufficient.

Verifying it works

After the rolling restart and ACL walk:

# 1. Anonymous client should be rejected at session establishment.
# NOTE: four-letter-word commands like 'ruok' do NOT establish a session
# and are not subject to sessionRequireClientSASLAuth. They will still
# respond. The correct test for anonymous rejection is zkCli without JAAS:
#   zkCli.sh -server localhost:2181 ls /
# and confirm it fails with an auth error (verified: 4lw commands are
# processed before session establishment, so they bypass SASL enforcement).

# 2. Authenticated client should succeed
zkCli.sh -server localhost:2181  # with JAAS configured
ls /kafka

# 3. ACL on a protected path should be visible and restrictive
getAcl /kafka

# 4. Audit log should now record the authenticated user, not anonymous
tail -n 20 /var/log/zookeeper/audit/zookeeper_audit.log

Run a controlled failover of one dependent service (one Kafka broker, one HBase RegionServer) and confirm two things: it can re-register under the new ACL, and the audit log captures the re-registration with the expected identity. If re-registration fails, you have an ACL gap; do not continue the rollout.

Common pitfalls

  • ACLs are not recursive. setAcl /hbase auth:cdrwa does not protect /hbase/rs or any child. You must walk and set each znode, or use a tool that does.
  • The ip ACL scheme is spoofable. It matches source IP and trusts network-level controls. Never use it as the only control on sensitive paths.
  • Digest auth transmits passwords in the clear. Acceptable inside a segmented network, unacceptable across an untrusted one. Use Kerberos or mTLS there.
  • HBase and YARN do not set ACLs by default. Turning SASL on without a parallel ACL walk leaves their znodes world-writable. SASL gates session creation, not per-node writes against OPEN_ACL_UNSAFE.
  • AdminServer is on by default on 8080. If your firewall assumed ZK listens only on 2181, you have a second exposure point. Confirm admin.enableServer and the bind address.
  • Four-letter-word whitelist defaults to srvr only. Monitoring that depends on mntr silently breaks on upgrade to 3.5.3+. Explicitly whitelist what you need.
  • Forcing SASL kicks out existing clients. Apply authProvider first; only flip sessionRequireClientSASLAuth=true once all known clients have working credentials.

Signals to monitor

All metrics below are exposed via mntr: zk_auth_failed_count, zk_ensemble_auth_fail, zk_non_mtls_remote_conn_count, and zk_tls_handshake_exceeded since 3.6; zk_insecure_admin_count and zk_unsuccessful_handshake since 3.7.

SignalWhy it mattersWarning sign
zk_auth_failed_countCounts SASL/Digest auth failures. Should be zero in steady state.Sustained non-zero rate after rollout: a client did not get new credentials, or a brute-force attempt.
zk_ensemble_auth_failServer-to-server auth failures between quorum members.Any non-zero increment threatens quorum.
zk_insecure_admin_countAdministrative operations performed without authentication.Non-zero in a hardened deployment means a control is bypassed.
zk_non_mtls_remote_conn_countRemote connections not using mutual TLS.Non-zero in an mTLS-required environment.
zk_unsuccessful_handshake / zk_tls_handshake_exceededTLS handshake failures and timeouts.Spikes after cert rotation or mTLS rollout.
Accepted socket connection log entriesSource IPs reaching 2181.IPs outside the known-client inventory.
NoAuth / Permission denied log entriesACL denials.Repeated denials from a single source warrant investigation.
audit.enable=true + zookeeper_audit.logForensic record of mutations on critical paths.Mutation on a coordination subtree by an unexpected identity.

How Netdata helps

  • Per-second collection of zk_auth_failed_count and connection-state metrics from mntr, so a burst of auth failures (a misconfigured rollout or a probing attempt) surfaces in the same window as the underlying connection changes, not on a 5-minute scrape cadence.
  • The connection and log-derived signals (Accepted socket connection, NoAuth) can be correlated against per-second auth counters in a single dashboard, which is the diagnostic gap that turns a bad SASL rollout into a multi-hour incident.
  • ML anomaly detection on zk_num_alive_connections and zk_packets_received catches the indirect signal of a lockdown change: the client fleet disconnecting and reconnecting as the new auth posture takes effect.
  • Per-node dashboards make a rolling restart visible. You can watch each node cycle and confirm zk_server_state returns to leader or follower without an election storm, which is when a SASL rollout usually goes sideways.