The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-range-count-too-high

Operations Guides

CockroachDB range count per node too high: Raft ticking as a scaling dimension

CPU utilization is climbing across one or more CockroachDB nodes, and the usual explanations do not fit. Query throughput is flat. Disk I/O looks healthy. Admission control is not queuing. L0 sublevels are in single digits. CPU keeps trending upward, and the only change is that the database holds more data.

The likely cause is the scaling dimension most teams never instrument: range count per node. Every range in CockroachDB is an independent Raft consensus group. Even idle, each range’s Raft state machine ticks at a fixed interval, consuming CPU separate from query execution, compaction, or any workload-driven process. This baseline overhead grows linearly with range count per node.

At 5,000 ranges per node, this overhead is rounding error. At 50,000, it is measurable. At 200,000, it is a different operating regime where Raft ticking alone consumes multiple CPU cores.

What this means

The entire CockroachDB keyspace is divided into ranges, each approximately 512 MiB by default (increased from 64 MiB in v20.1). Each range is replicated (typically 3x) and every replica participates in a Raft consensus group. One replica is the Raft leader, one is the leaseholder, and in the common case they are co-located on the same node.

Raft is not idle when there are no writes. Every range’s Raft state machine ticks at a fixed interval. The default is 500ms, controlled by the COCKROACH_RAFT_TICK_INTERVAL environment variable. Each tick involves processing heartbeat messages, checking election timers, and maintaining state. This work happens for every range on the node, every tick interval, regardless of traffic.

Rough guideline: approximately 1% of a CPU core per 500 to 1,000 ranges, just for Raft ticking. A node holding 50,000 ranges consumes roughly half a core to a full core before any query executes. At 100,000 or more, Raft ticking alone consumes multiple cores.

Two distinctions matter for diagnosis:

  • Replica count vs. lease count: The ranges metric reports total replicas per store. The replicas_leaseholders metric (Prometheus form of replicas.leaseholders) reports leases held. A node may hold 10,000 replicas but only 3,000 leases. Replica count drives Raft processing overhead. Lease count drives query serving load. Both consume CPU, but through different mechanisms.
  • Quiesced ranges still cost CPU: CockroachDB quiesces ranges that have had no proposals for several tick intervals, reducing heartbeat network traffic. But quiesced ranges still require processing when unpacked. Quiescence reduces but does not eliminate the marginal cost of maintaining inactive ranges.

When range count consumes enough baseline CPU to shrink headroom, the node becomes brittle. Burst traffic, compaction spikes, or GC pauses that previously fit within idle capacity now collide with the Raft ticking floor. If CPU saturation delays Raft heartbeats, the node loses leases or liveness. The cluster redistributes ranges to surviving nodes, increasing their range count and baseline CPU, accelerating the cycle.

flowchart TD
    A[Data grows, no nodes added] --> B[Range count per node rises]
    B --> C[Raft ticking CPU grows]
    C --> D[Less headroom for queries and compaction]
    D --> E[CPU saturates during bursts]
    E --> F[Raft heartbeats delayed]
    F --> G[Node loses leases or liveness]
    G --> H[Ranges redistributed to survivors]
    H --> B

Common causes

CauseWhat it looks likeFirst thing to check
Data growth without node additionsRange count rising proportionally across all nodes; CPU baseline climbing with no query throughput changeCompare current range count to historical trend; project when you cross 50,000 per node
Range imbalance (>20% deviation)One or two nodes with significantly more ranges than the mean; CPU asymmetric across nodesPer-store ranges metric; check for zone constraint conflicts or heterogeneous node sizes
New node not absorbing its shareAfter adding a node, range count on existing nodes stays flat or keeps risingCheck ranges_underreplicated and rebalance queue activity; verify the new node has capacity
Excessive splits from sequential keysRange count higher than expected for the data volume; many ranges far below 512 MiBCompare total ranges to expected count (total data / 512 MiB); check for sequential primary keys

Quick checks

# Check range and lease count per store
curl -s http://localhost:8080/_status/vars | grep -E '^ranges\b|replicas_leaseholders'

# Check CPU utilization
curl -s http://localhost:8080/_status/vars | grep sys_cpu

# Check for under-replication (rebalancing may be stuck)
curl -s http://localhost:8080/_status/vars | grep ranges_underreplicated

# Check store capacity to estimate expected range count
curl -s http://localhost:8080/_status/vars | grep -E 'capacity'

# Check lease transfer rate (elevated rate may indicate instability)
curl -s http://localhost:8080/_status/vars | grep leases_transfers

# Per-leaseholder range distribution (admin-only, expensive, diagnosis only)
# cockroach sql -e "SELECT lease_holder, count(*) FROM crdb_internal.ranges GROUP BY lease_holder ORDER BY count DESC;"

How to diagnose it

  1. Get per-store range counts. Scrape ranges per store. With replication factor 3, each unique range produces 3 replicas. The sum of ranges across all stores divided by the replication factor gives the cluster’s unique range count. Expected total unique ranges is approximately total data divided by 512 MiB.

  2. Check for imbalance. Calculate the mean range count per node. If any node deviates more than 20% from the mean, the allocator is not distributing evenly. Common causes: zone constraint conflicts, heterogeneous node sizes, or a node with insufficient disk space to accept more replicas.

  3. Compare range count to data volume. If your cluster holds 10 TB of data with RF=3, you have approximately 20,000 unique ranges. Across 5 nodes, each holds roughly 12,000 replicas. If actual range count is significantly higher, you have many small ranges from excessive splitting on sequential key patterns.

  4. Correlate CPU with range count. If CPU is elevated but query throughput, disk I/O, and compaction are normal, Raft ticking is likely a significant consumer. Profile with Go’s pprof endpoint (/debug/pprof/profile) to confirm. Look for Raft-related functions consuming a disproportionate share of CPU time.

  5. Check startup time. Higher range count means longer startup. A node that restarted in 30 seconds six months ago and now takes several minutes is carrying significantly more ranges. Track sys_uptime after restarts to observe the trend.

  6. Verify lease distribution. Compare ranges (replicas) to replicas_leaseholders per store. If one node holds disproportionately many leases, it bears both Raft overhead and query serving load, compounding CPU pressure. Note that crdb_internal.ranges queries are expensive and perform cluster-wide RPC fan-out. Use them for manual diagnosis only, never for continuous monitoring.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ranges (per store)Directly measures the Raft ticking CPU multiplierAny node approaching 50,000; trajectory heading there within months
replicas_leaseholders (per store)Determines query serving load; combined with high range count, compounds CPU pressureImbalance exceeding 20% deviation from mean
sys_cpu_user_ns, sys_cpu_sys_nsTotal CPU budget; baseline consumption from Raft reduces headroomSustained above 70-80% with no corresponding query throughput increase
ranges_underreplicatedIndicates rebalancing is stuck or blockedNonzero without corresponding operational event
leases_transfers_successHigh rate suggests instability from range-count-driven CPU pressureMore than 10x baseline without operational cause
sys_uptime (post-restart)Proxy for range count growth through startup timeStartup time increasing over months

Fixes

Add nodes

The most direct fix. Adding nodes distributes ranges across more machines, reducing per-node Raft overhead. The allocator automatically rebalances replicas and leases. On v23.1+, the load-based rebalancing objective (kv.allocator.load_based_rebalancing.objective) defaults to cpu.

Monitor the rebalancing process. Check that ranges_underreplicated converges to zero and that the new node’s range count trends toward the cluster mean. If the new node is not absorbing ranges, check for zone constraint conflicts or insufficient disk space.

Address range imbalance

If one or two nodes hold significantly more ranges than others (>20% deviation), investigate three common causes:

  • Zone constraint conflicts: Zone configs may restrict where replicas can be placed. Tight constraints prevent even distribution. Check with SHOW ZONE CONFIGURATIONS;.
  • Heterogeneous node sizes: Smaller nodes may be full, preventing the allocator from placing replicas even when aggregate free space exists.
  • Disk space: A node near its capacity limit cannot accept new replicas.

Reduce range count via merges

If your cluster has many small ranges from deleted data or sparse key regions, range merges can consolidate them. CockroachDB performs merges automatically for sufficiently empty adjacent ranges. If your data has large sparse regions (TTL-expired data, dropped tables), ensure the merge queue is active and not blocked.

Merges are only effective for sparse or deleted regions. Active data at the default 512 MiB range size cannot be merged further.

Review range size configuration

If you are on a version prior to v20.1 or have manually overridden range size, you may have smaller ranges than necessary. The default is approximately 512 MiB. Larger ranges mean fewer ranges for the same data volume, directly reducing Raft ticking overhead. Very large ranges increase the cost of Raft snapshots (up to the range size per snapshot) and can slow recovery.

Short-term mitigation: reduce background work

If you cannot immediately add nodes, reduce non-essential background work to free CPU for Raft processing. Pause bulk imports, defer schema changes, and ensure admission control is active. This buys time but does not address the underlying range count growth.

Prevention

  • Track range count per node as a capacity metric. It belongs alongside CPU, memory, and disk in capacity planning. The per-node target is below 50,000 for typical workloads.
  • Project range count growth from data growth. Total unique ranges is approximately total data divided by 512 MiB. If data grows 10% per year and node count is fixed, range count per node grows at the same rate. Plot the trajectory and plan node additions before crossing the threshold.
  • Alert on range imbalance exceeding 20% deviation. Persistent imbalance indicates allocator constraints that will compound as range count grows.
  • Watch startup time as a leading indicator. Increasing startup time signals growing range count before CPU metrics reflect the problem.
  • Account for replication factor in capacity planning. Increasing RF from 3 to 5 increases total replica count by 67% without adding data, directly increasing per-node range count.

How Netdata helps

  • Per-second range and lease metrics. Netdata scrapes ranges and leases_count per store at high frequency, giving immediate visibility when range count trends upward or rebalancing creates imbalance.
  • CPU correlation. Overlay range count with sys_cpu_user_ns and sys_cpu_sys_ns to confirm whether baseline CPU is growing in lockstep with range count. When Raft ticking is the driver, the correlation is tight and query throughput stays flat.
  • Startup time tracking. sys_uptime after restart events reveals whether startup time is creeping upward as range count grows.
  • Lease transfer and liveness correlation. High range count can destabilize a node during CPU bursts. Correlating leases_transfers_success and node liveness status with CPU saturation distinguishes range-count-driven instability from other causes.
  • Growth trajectory. Long-term retention of per-store range count lets you project when you will cross the 50,000 per node threshold, giving lead time to add capacity.

See CockroachDB monitoring with Netdata for setup details.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.