The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kafka / kafka-broker-may-not-be-available

Operations Guides

Kafka Broker May Not Be Available: How To Fix It

Client logs show a warning like this:

WARN [Producer clientId=...] Connection to node -1 could not be established. Broker may not be available.

Python clients may raise kafka.errors.NoBrokersAvailable, while librdkafka-based clients report Connection refused against the broker address. The bootstrap server is often reachable; ping and telnet succeed, yet the client still fails. This happens because Kafka clients use the bootstrap connection only to fetch metadata. After that, they disconnect and try to open fresh TCP connections to the host:port pairs advertised in the metadata response. If those endpoints are unreachable, misconfigured, or secured differently than the client expects, the connection fails even though the bootstrap succeeded.

What this means

This is a client-side symptom of a failed second-hop connection. The bootstrap succeeded and the client received a metadata response containing broker endpoints. The direct connection to one of those brokers then failed. The failure can happen at multiple layers: TCP routing, DNS resolution, TLS handshake, SASL authentication, or because the target broker is genuinely down. In containerized environments, the most common root cause is an advertised.listeners configuration that returns an internal hostname or IP to external clients. The broker is healthy, but it is telling clients to connect to an address they cannot reach.

flowchart TD
    A[Client sees Broker may not be available] --> B{Bootstrap host:port reachable?}
    B -->|No| C[Fix client network or DNS]
    B -->|Yes| D{kafka-broker-api-versions.sh works from client?}
    D -->|No| E[Check bootstrap listener and firewall]
    D -->|Yes| F{Advertised endpoint reachable from client?}
    F -->|No| G[Fix advertised.listeners or NAT rules]
    F -->|Yes| H{Broker logs show TLS or SASL errors?}
    H -->|Yes| I[Fix handshake or auth config]
    H -->|No| J[Check broker process health and load]

Common causes

CauseWhat it looks likeFirst thing to check
advertised.listeners misconfiguration (Docker, K8s, NAT)Client receives metadata but cannot resolve or reach the advertised host:port. Connections work inside the container network but fail from outside.kafka-broker-api-versions.sh from the client’s network. Compare the returned broker list against what the client can reach.
TLS or SASL handshake failureTCP connection establishes but is immediately reset or hangs. Broker logs contain SSLHandshakeException or AuthenticationException.Broker logs for handshake or authentication errors. Verify client and broker protocol versions and SASL mechanisms match.
Firewall or security group dropTCP SYN to the advertised port times out. The bootstrap port is open, but the data port is not.nc -zv or Bash /dev/tcp to the advertised host:port from the client host.
Broker process down or overloadedConnection refused to the advertised port, or the broker appears in metadata but does not respond to API requests.Broker process liveness, port binding with ss -tlnp, and NetworkProcessorAvgIdlePercent via JMX.

Quick checks

# TCP reachability to bootstrap
nc -zv <bootstrap-host> <bootstrap-port>

# Bash built-in alternative
timeout 5 bash -c 'cat < /dev/null > /dev/tcp/<bootstrap-host>/<bootstrap-port>'

# Kafka protocol reachability from client
kafka-broker-api-versions.sh --bootstrap-server <bootstrap-host>:<bootstrap-port>

# Broker logs for listener and handshake errors
grep -E "advertised\.listeners|listeners=|ERROR.*SSLHandshakeException|ERROR.*AuthenticationException" /var/log/kafka/server.log

# Broker listening interfaces
ss -tlnp | grep java

# KRaft quorum health
kafka-metadata-quorum.sh --bootstrap-server localhost:9092 describe --status

# File descriptor limits
cat /proc/$(pgrep -f kafka.Kafka)/limits | grep "Max open files"

How to diagnose it

  1. Test basic TCP connectivity to the bootstrap server using nc or Bash /dev/tcp. If this fails, investigate routing, DNS, or firewalls on both sides before changing broker configuration.
  2. Run kafka-broker-api-versions.sh --bootstrap-server <host>:<port> from the client machine. This performs a full Kafka protocol handshake. If it fails, the client cannot reach the Kafka protocol layer. This often indicates a firewall blocking the port or the broker listening on an interface the client cannot reach.
  3. If the API versions check succeeds, the cluster is reachable at the bootstrap address. The issue is likely the second hop. Inspect the broker’s advertised.listeners in server.properties or broker logs. Ensure the advertised host:port is resolvable and routable from the client’s network. In Docker or Kubernetes, the default advertised address is often the pod IP or container hostname, which external clients cannot resolve.
  4. Test TCP connectivity from the client to the specific advertised host:port that the client is complaining about. Use nc -zv <advertised-host> <advertised-port>. If this fails while the bootstrap works, you have a NAT, routing, or advertised.listeners mismatch.
  5. Check broker logs at /var/log/kafka/server.log for SSLHandshakeException, AuthenticationException, or SASL mechanism mismatch errors. If TCP connects but the connection drops during negotiation, the issue is at the security protocol layer, not the network layer.
  6. On the broker, verify the process is running and bound to the correct network interface. A common misconfiguration is listeners=PLAINTEXT://localhost:9092, which prevents external connections entirely. Use ss -tlnp to confirm the listening address.
  7. Check whether the broker is saturated. Even a running broker can refuse connections if network threads are exhausted. Query the JMX MBean kafka.network:type=SocketServer,name=NetworkProcessorAvgIdlePercent. If the value is below 0.1 (10% idle), the broker is overloaded and may drop connections.
  8. In KRaft mode, verify quorum health with kafka-metadata-quorum.sh --bootstrap-server <host>:<port> describe --status. If there is no quorum leader, metadata may be stale and brokers may advertise incorrect endpoints.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Authentication failuresFailed TLS or SASL handshakes often present as broker unavailability to clients.Sustained AuthenticationException in broker logs.
Network Processor Average Idle PercentSaturated network threads cannot accept new connections or complete handshakes.Sustained value below 0.3 (30% idle).
Offline Partitions CountConfirms whether a broker is genuinely down and its partitions are leaderless.Nonzero value reported by the active controller.
Active Controller CountStale or missing metadata means clients may receive bad broker lists.Cluster-wide sum is not exactly 1.
Connection CountApproaching the file descriptor or thread limit causes silent connection refusal.Count exceeds 2x baseline or nears the OS FD limit.
Failed produce requestsDistinguishes a pure connectivity issue from a protocol or authorization failure.Sustained nonzero rate outside of rolling restarts.

Fixes

Fix advertised.listeners mismatches

Edit server.properties on each broker and set advertised.listeners to an address that clients in their respective networks can resolve and reach. In containerized environments, this usually means advertising the node IP, load balancer, or ingress hostname instead of the pod IP.

Use separate listeners for inter-broker and client traffic:

listeners=INTERNAL://0.0.0.0:9092,EXTERNAL://0.0.0.0:9093
advertised.listeners=INTERNAL://broker-1.internal:9092,EXTERNAL://kafka.example.com:9093
listener.security.protocol.map=INTERNAL:PLAINTEXT,EXTERNAL:PLAINTEXT

Map each listener name to the correct security protocol. After changing advertised.listeners, restart the broker. Clients must then use the external bootstrap server to ensure they receive the external advertised endpoints.

Tradeoff: Multiple listeners add configuration complexity and require careful firewall rules, but they eliminate the single biggest source of “broker may not be available” errors in NATed environments.

Do not rely on /etc/hosts workarounds on client machines. They are fragile and break when broker IPs change.

Fix TLS or SASL handshake failures

Ensure the client and broker support a common TLS version. Many modern Kafka deployments on recent JVMs negotiate TLS 1.2 and higher. If the client is restricted to TLS 1.3 and the broker only offers 1.2, the handshake will fail.

For SASL, verify the mechanism matches on both sides. For example, AWS MSK with IAM authentication requires the client to use the IAM mechanism, not SCRAM-SHA-256. Check broker logs for the exact AuthenticationException message.

Inspect certificate validity, trust stores, and whether the broker presents the full certificate chain. A missing intermediate certificate causes client-side trust failures that look like network timeouts.

Tradeoff: Stricter TLS cipher suites and certificate validation improve security but can break legacy clients.

Fix firewall or security group blocks

Open the advertised broker port from the client subnet to the broker host. Remember that the client connects to the advertised endpoint, not just the bootstrap endpoint. If you use a two-listener setup, ensure both the bootstrap listener port and the advertised data listener port are accessible.

Verify that return traffic is allowed. Stateful firewalls usually handle this, but asymmetric routing or strict ACLs can block response packets.

Recover a genuinely down or overloaded broker

If the broker process is not running, investigate why before restarting. Check related guides for disk exhaustion, OOM kills, or controller queue backups. A restart is not free: the broker loses its page cache, must re-fetch replicas to join the ISR, and triggers client reconnections across the cluster.

If the process is running but unresponsive, check for GC pauses via JMX heap metrics, disk I/O latency via iostat, or request queue saturation. If NetworkProcessorAvgIdlePercent is near zero, the broker is overloaded. Temporarily throttle producers with quotas or migrate partitions to reduce load.

Tradeoff: A rolling restart of a sick broker can restore service, but expect elevated UnderReplicatedPartitions and latency for several minutes to hours depending on partition size.

Prevention

  • Validate advertised.listeners from a client-equivalent network location after every broker deployment, container reschedule, or infrastructure change.
  • Monitor broker authentication failure rates via JMX or logs to catch TLS and SASL drift before clients fail.
  • Manage firewall rules and security groups in infrastructure-as-code so advertised ports remain open by default.
  • Avoid making a single broker the only bootstrap host for critical clients. If that broker is down, clients cannot even fetch metadata.

How Netdata helps

  • Correlate broker authentication failure rates with client connection errors to distinguish network issues from security misconfigurations.
  • Alert on low Network Processor Average Idle Percent to catch network thread saturation before clients time out.
  • Monitor Offline Partitions Count and Active Controller Count to confirm whether a broker is genuinely down or the cluster metadata plane is degraded.
  • Track broker process uptime and OS-level TCP metrics to spot file descriptor exhaustion or listen-queue overflows that silently reject connections.
  • Cross-reference client-visible errors with broker request queue size and request handler idle percentage to separate overload from connectivity failures.