The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kafka / kafka-too-many-partitions-per-broker

Operations Guides

Kafka too many partitions per broker: controller load, recovery time, and the 4000 guideline

Every replica on a Kafka broker carries fixed overhead: file descriptors, memory-mapped index files, replication fetcher load, and controller metadata. The guideline of roughly 4000 partitions per broker is not a hard architectural limit. It is an operational warning based on a serial bottleneck: the controller processes partition state changes one at a time. When a broker with 5000 partitions fails, the active controller queues thousands of events. Until the queue drains, leader elections stall, metadata propagation slows, and clients see NOT_LEADER_FOR_PARTITION. Restarting the broker forces every replica to recover its logs and catch up from leaders before rejoining the ISR. Recovery time is a step function; a routine restart can become a multi-hour incident. Partition count is easy to ignore in steady state and catastrophic to discover during an outage.

What it is and why it matters

PartitionCount per broker is the total number of leader and follower replicas assigned to that broker. Read it from the JMX MBean kafka.server:type=ReplicaManager,name=PartitionCount. It changes only when topics are created, deleted, or reassigned.

Each partition adds overhead in four dimensions:

  • File descriptors: every log segment requires a .log file and at least one index file (.index, .timeindex). Hundreds of partitions with rolling segments exhaust the default ulimit -n of 1024 quickly. Production deployments should raise this to 100,000 or higher.
  • Memory: brokers memory-map index files and maintain metadata structures per partition. This overhead competes with the OS page cache even when message volume is low.
  • Replication fetcher load: followers catch up from leaders via replica fetcher threads. More follower partitions increase disk I/O, network utilization, and CPU consumption across the fetcher thread pool.
  • Controller event queue: every ISR change, leader election, and reassignment becomes an event that the active controller processes serially.

The 4000 partition guideline is a conservative threshold. The actual ceiling depends on hardware, disk latency, network bandwidth, and whether the cluster runs ZooKeeper or KRaft mode. Clusters on slower spinning disks or with constrained network throughput will hit controller and recovery limits well before 4000. KRaft removes ZooKeeper session bottlenecks and improves failover latency, but the active controller still processes partition events sequentially. A broker with 5000 partitions still generates roughly 5000 events on failure, and a restarted broker still must catch up all 5000 replicas regardless of the metadata quorum implementation.

How it works

The critical path is the controller event queue. Kafka maintains exactly one active controller. It holds a queue of pending events: ISR expansions and shrinks, leader elections, topic changes, and broker lifecycle updates. It processes them sequentially. There is no horizontal scaling for this queue.

When a broker fails:

  • For partitions where the dead broker was the leader, the controller elects a new leader from the remaining ISR.
  • For partitions where the dead broker was a follower, surviving leaders eventually shrink the ISR, which the controller also processes.

In total, a broker failure generates on the order of one controller event per partition it hosted. A broker with 5000 partitions therefore injects roughly 5000 events into a single-threaded queue. While those events drain, new failures, rolling restarts, or reassignment operations add more work. Queue depth grows, and LeaderElectionRateAndTimeMs spikes. Partitions that need new leaders remain offline until the controller reaches their events. Clients see NOT_LEADER_FOR_PARTITION and retry, amplifying load on the remaining brokers.

Broker restart compounds the problem. After restart, the broker loads every partition directory, replays unflushed messages from the recovery point, and rebuilds in-memory indexes. Only then does it open replica fetcher connections to leaders. Until the follower catches up to the high-water mark, the leader does not add it back to the ISR. With thousands of partitions, log recovery alone can take tens of minutes. The catch-up phase then keeps UnderReplicatedPartitions elevated for minutes or hours, depending on data volume, network bandwidth, and disk speed. During this window the cluster runs at reduced durability, and any additional failure risks unavailability.

flowchart TD
    A[Broker hosts 5000 partitions] --> B[Steady-state FD, memory, and fetcher overhead]
    A --> C[Broker fails]
    C --> D[Controller queues ~5000 events]
    D --> E[Sequential drain]
    E --> F[Leader elections delayed]
    F --> G[UnderReplicatedPartitions rises]
    F --> H[Clients see NOT_LEADER_FOR_PARTITION]
    C --> I[Surviving brokers absorb leadership traffic]
    A --> J[Broker restarts]
    J --> K[Catch-up from leaders takes minutes to hours]

Where it shows up in production

Rolling restarts leave the cluster under-replicated for hours. If one broker takes thirty minutes to catch up after restart, a ten-broker rolling restart creates a five-hour window of continuous under-replication. If the cluster uses a replication factor of two, that window is also a period of reduced redundancy: a single additional disk failure on the broker that hosts the other replica risks unavailability. A second full broker failure during that window doubles the controller queue depth and can leave some partitions without a viable leader.

Single broker death becomes a metadata stall. Operators watching only OfflinePartitionsCount may notice a lag between the broker failure and the count rising. That lag is the controller queue backing up. ControllerEventQueueSize reveals the real problem: thousands of pending events.

Untested recovery times. Teams that do not run failure tests often discover during incidents that a broker restart takes ninety minutes. They know their broker count and replication factor, but have never measured recovery duration. Without a baseline, on-call engineers cannot tell whether a restart is progressing normally or has stalled on a slow disk or network partition.

Leadership skew creates hot spots. Even balanced PartitionCount can hide leadership imbalance. A broker leading 80% of its assigned partitions handles far more produce and fetch traffic than a peer leading 20%. RequestHandlerAvgIdlePercent drops on the hot broker. If it fails, the traffic shift plus the controller event wave creates a compound outage. The remaining brokers must absorb the redirected produce traffic while also handling the controller event wave. If any survivor was already near its limit, it too can become a bottleneck.

Tradeoffs and common misuses

Partition count is often treated as a throughput knob. More partitions allow higher consumer parallelism and smaller per-partition batch latencies. Consumer groups scale parallelism to the partition count of the subscribed topics. Adding partitions to increase throughput is valid, but the operational tax is permanent and paid on every broker that hosts a replica.

  • Over-provisioning partitions “for future scale” creates permanent overhead. A topic with 64 partitions and replication factor three adds 192 replicas cluster-wide. If those replicas land on brokers that already hold 3500 partitions, the broker crosses the 4000 threshold for a topic that may not yet have any traffic.
  • Ignoring follower load is common because PartitionCount includes both leaders and followers. A broker with few leaders but many followers does not serve heavy client traffic, yet it still consumes disk I/O, network, and controller attention during failures. During a rolling restart, this broker still has to catch up thousands of follower partitions, extending the under-replicated window even though it was never a hot spot under normal load.
  • Looking only at cluster averages hides imbalance. A fifty-broker cluster with 200,000 total partitions averages 4000 per broker. If two brokers hold 8000 partitions due to a decommissioning or reassignment error, those two brokers dominate recovery time and controller load.
  • Assuming KRaft removes all limits. KRaft improves controller failover and raises total cluster partition limits, but the active controller still processes partition state changes serially. High partition churn from frequent broker failures or rapid reassignments can still back up the event queue.

Signals to watch in production

SignalWhy it mattersWarning sign
kafka.server:type=ReplicaManager,name=PartitionCountTotal partition overhead per broker>4000 partitions, or >20% deviation from cluster mean
kafka.server:type=ReplicaManager,name=LeaderCountConcentration of request-processing loadAny broker >30% above the cluster mean
kafka.controller:type=ControllerEventManager,name=EventQueueSizePending metadata operations on the active controller>100 sustained, or growing continuously above 1000
kafka.controller:type=ControllerStats,name=LeaderElectionRateAndTimeMsSpeed of failure recoveryElection time consistently >1s, or burst outside maintenance
kafka.server:type=KafkaRequestHandlerPool,name=RequestHandlerAvgIdlePercentBroker processing headroom<0.3 sustained
kafka.server:type=ReplicaManager,name=UnderReplicatedPartitionsDurability degradation during recoveryStays flat or grows after a broker restart, rather than trending down
Recovery duration after controlled shutdownReal-world operational readinessMeasured during game day; should recover within minutes, not hours

How Netdata helps

Netdata collects per-broker JMX metrics directly. Use it to spot partition-related imbalances without maintaining a separate metrics store.

  • Per-broker PartitionCount and LeaderCount charts show skew across all brokers on one dashboard.
  • Controller event queue depth on the active controller correlates with broker restarts and failures.
  • Correlate RequestHandlerAvgIdlePercent with LeaderCount to distinguish saturation caused by leadership density from raw throughput pressure.
  • Anomaly detection on UnderReplicatedPartitions and IsrShrinksPerSec flags recoveries that exceed historical baselines.
  • JMX ingestion without an external store means you can monitor a new broker as soon as its JMX endpoint is accessible and immediately see partition load and recovery behavior.

When a rolling restart begins, watch ControllerEventQueueSize on the active controller. If it climbs above 100 while the first broker is restarting, pause the rollout until the queue drains. Use per-broker PartitionCount to verify that the restarted broker rejoined with the expected replica set. If UnderReplicatedPartitions does not trend down within your baseline recovery window, investigate disk I/O or network throttling before proceeding to the next broker.