The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kafka / kafka-offline-partitions-count

Operations Guides

Kafka OfflinePartitionsCount > 0: partitions with no leader and how to recover

When kafka.controller:type=KafkaController,name=OfflinePartitionsCount is nonzero, at least one partition has no active leader. Those partitions are completely unavailable: producers receive errors, consumers stall, and no data is written or read until a leader is elected. This is a data-plane outage.

This metric is only meaningful on the active controller. Non-controller brokers always report zero. If the controller itself is down, the metric may be stale or unreachable at the exact moment you need it. Brief spikes can occur during controller re-election or ungraceful broker shutdown, but any sustained nonzero value past 60 seconds is an active incident that requires immediate intervention.

Partitions go offline when the controller cannot elect a leader from the current In-Sync Replica set (ISR). The most common trigger is a broker failure that takes all replicas for a partition down, or an ISR that already shrank to a single leader which then fails. The path from healthy replication to offline partitions usually runs through UnderReplicatedPartitions, making under-replication the leading indicator and offline partitions the impact.

flowchart TD
    A[Broker failure or degradation] --> B[Follower falls behind]
    B --> C[ISR shrinks]
    C --> D[UnderReplicatedPartitions rises]
    D --> E[No ISR member available]
    E --> F[Controller cannot elect leader]
    F --> G[OfflinePartitionsCount > 0]

What this means

OfflinePartitionsCount measures partitions that currently have no leader. Without a leader, no broker is authorized to accept writes or serve reads for that partition. The controller elects a new leader from the ISR. If the ISR is empty or the controller cannot process the election, the partition stays offline.

The severity is PAGE if the metric is nonzero for more than 60 seconds and no broker or controller has an uptime below 600 seconds. That uptime gate filters out noise from cold starts and rolling restarts. During steady state, this metric must be zero.

Topics with replication.factor=1 are especially vulnerable. If the single broker hosting that partition goes down, there are no other replicas to promote. The partition remains offline until that specific broker returns.

Common causes

CauseWhat it looks likeFirst thing to check
All replicas for a partition are on failed brokersMultiple brokers unreachable; UnderReplicatedPartitions spiked just before the outageBroker process liveness and network reachability on the replica nodes
No ISR member available with unclean elections disabledunclean.leader.election.enable=false (default) and the only remaining broker in the ISR is downkafka-topics.sh --describe --unavailable-partitions to see the ISR state
Controller unable to elect a leaderEventQueueSize (ControllerEventManager) is growing; elections are delayed or stalledActive controller health and metadata store latency
Broker decommissioned without partition reassignment, or RF=1 broker downPartitions still assigned to a broker ID that no longer exists, or a single-replica partition on a failed nodeTopic replica assignment list against current live broker IDs

Quick checks

Run these safe, read-only commands to orient yourself.

# List unavailable partitions and their ISR state
kafka-topics.sh --bootstrap-server localhost:9092 --describe --unavailable-partitions
# Confirm which broker is the active controller
echo "get -b kafka.controller:type=KafkaController,name=ActiveControllerCount Value" | java -jar jmxterm.jar -l localhost:9999
# Read OfflinePartitionsCount on the controller broker
echo "get -b kafka.controller:type=KafkaController,name=OfflinePartitionsCount Value" | java -jar jmxterm.jar -l localhost:9999
# Check under-replicated partitions cluster-wide
kafka-topics.sh --bootstrap-server localhost:9092 --describe --under-replicated-partitions
# Inspect controller event queue depth
echo "get -b kafka.controller:type=ControllerEventManager,name=EventQueueSize Value" | java -jar jmxterm.jar -l localhost:9999
# Check for log directories taken offline due to disk errors
echo "get -b kafka.log:type=LogManager,name=OfflineLogDirectoryCount Value" | java -jar jmxterm.jar -l localhost:9999
# Verify broker processes are listening on the expected port
ss -tnp | grep :9092

How to diagnose it

  1. Confirm controller authority. Verify that the broker where you are reading OfflinePartitionsCount reports ActiveControllerCount = 1. If the controller is down, the metric may be stale or unreachable, and the cluster cannot self-heal.
  2. Identify the victims. Run kafka-topics.sh --describe --unavailable-partitions. Record the topic, partition, replication factor, and the current ISR for each offline partition.
  3. Map replicas to brokers. Cross-reference the offline partition replicas with broker liveness. If all replica brokers are down, the cause is broker failure. If some replicas are up but not in the ISR, the ISR already shrank before the leader was lost.
  4. Check the controller pipeline. Query EventQueueSize on kafka.controller:type=ControllerEventManager. If it is consistently above 100 or growing, the controller is backlogged and cannot process leader elections fast enough. Check LeaderElectionRateAndTimeMs for slowing election velocity.
  5. Correlate with under-replication. A cluster-wide rise in UnderReplicatedPartitions that began before the offline event indicates a cascading replication failure. Identify the common follower that dropped out first; its disk or network metrics usually reveal the root cause.
  6. Inspect broker log directories. An OfflineLogDirectoryCount greater than zero means a broker shut down a disk due to I/O errors. Partitions on that path are unavailable even if the broker process is running.
  7. Determine if unclean election is enabled. If unclean.leader.election.enable=false (the default since Kafka 0.11.0.0) and no ISR member is alive, the partition will stay offline until an ISR replica recovers. If UncleanLeaderElectionsPerSec is nonzero, unclean elections are occurring somewhere in the cluster; the newly elected leader may have truncated data.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
OfflinePartitionsCountDirect count of partitions with no leaderNonzero for more than 60 seconds outside of restarts
UnderReplicatedPartitionsLeading indicator that replication is degrading before leaders are lostNonzero and growing across the cluster
ActiveControllerCountWithout a controller, no leader elections can occurCluster-wide sum not equal to 1
EventQueueSize (ControllerEventManager)Backed-up queue delays elections and ISR updatesConsistently above 100 or growing without draining
UncleanLeaderElectionsPerSecConfirms silent data loss when enabled; inverse of offline when disabledAny nonzero delta
LogManager OfflineLogDirectoryCountDisk or filesystem failure removes partitions from serviceAny nonzero value
IsrShrinksPerSecVelocity of replicas leaving the ISRSustained nonzero outside maintenance windows

Fixes

Restore an ISR replica

If unclean.leader.election.enable=false (the safe default), the controller will only elect a leader from the ISR. The only way to recover without data loss is to bring a missing ISR member back online. Restart the failed broker and monitor IsrExpandsPerSec and UnderReplicatedPartitions as it catches up.

For partitions with replication.factor=1, the single broker is the only possible leader. Bring that broker back. If its storage is permanently lost, the partition data is unrecoverable. You must delete and recreate the topic or accept permanent data loss for that partition.

Resolve controller metadata backlog

If the controller event queue (EventQueueSize) is growing and offline partitions are piling up, do not restart additional brokers. Each restart generates more controller events and worsens the backlog. Check ZooKeeper request latency (in ZooKeeper mode) or KRaft quorum health (in KRaft mode). If the metadata store is degraded, the controller is bottlenecked by external latency. Once the metadata store recovers, monitor the queue drain rate. If the queue drains, allow the controller to work through the backlog before taking further action.

Reassign partitions from a decommissioned broker

If a broker was removed without reassigning its partitions, those partitions may have no live replicas. Generate a partition reassignment JSON and execute it with kafka-reassign-partitions.sh. Verify progress with the --verify flag. Do not decommission brokers without first moving their partitions to remaining cluster members.

Force an unclean leader election (last resort)

If no ISR member can be restored and availability is more important than data consistency, enable unclean.leader.election.enable=true. You can apply this at the topic level to avoid a broker restart, but a cluster-wide default change requires a rolling restart. A non-ISR replica will be elected leader and truncate its log to its own offset. Any data acknowledged by the previous leader but not replicated to the new leader is silently lost. Revert the setting to false immediately after recovery to prevent future data loss.

Prevention

  • Avoid replication.factor=1 for critical topics. A single broker failure guarantees an outage.
  • Keep unclean.leader.election.enable=false unless your application explicitly tolerates data loss.
  • Decommission brokers only after reassigning partitions. Verify with kafka-reassign-partitions.sh --verify before removing the node.
  • Monitor UnderReplicatedPartitions and IsrShrinksPerSec. These provide early warning of the cascade that ends in offline partitions.
  • Watch controller queue depth and metadata store latency. A healthy controller should process events near zero queue depth in steady state.
  • Set min.insync.replicas=2 when using acks=all and replication.factor=3. This prevents the ISR from shrinking to a single replica, reducing the chance of total leader loss.

How Netdata helps

  • Correlates OfflinePartitionsCount with UnderReplicatedPartitions, ActiveControllerCount, and broker process liveness on one timeline to distinguish controller issues from broker failures.
  • Surfaces controller-only metrics automatically from the active controller node; non-controller nodes are omitted so you do not read stale zeros.
  • Alerts on ControllerEventQueueSize and JVM GC pauses to catch controller bottlenecks before they delay leader elections.
  • Tracks per-broker disk I/O latency and OfflineLogDirectoryCount to pinpoint hardware degradation that removes partitions from service.
  • Visualizes IsrShrinksPerSec and IsrExpandsPerSec to flag the replication collapse that precedes offline partitions.