The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-hot-range-bottleneck

Operations Guides

CockroachDB hot range bottleneck: one leaseholder saturated while the cluster idles

One node runs hot while the rest of the cluster idles. CPU on that node is 80%+, but others hover around 20-30%. SQL latency is elevated, but only for specific tables. Transaction retry rates climb with writetooold as the dominant restart cause. No range is unavailable, no node has lost liveness, and disk I/O is within normal bounds. The cluster has aggregate capacity, but a single range is funneling all its traffic through one leaseholder.

This is a hot range bottleneck. The root cause is almost always a monotonic key pattern in your primary key or index that concentrates writes and reads onto a narrow slice of the keyspace. CockroachDB’s range-level architecture means all requests for a given range pass through its leaseholder. When one range receives disproportionate traffic, that leaseholder becomes a serial bottleneck for the entire workload, regardless of how much capacity the rest of the cluster has.

The asymmetry is the diagnostic fingerprint. If all nodes are equally loaded, the problem is elsewhere.

What this means

CockroachDB divides its keyspace into ranges of approximately 512 MiB each. Every range has multiple replicas (default 3) spread across nodes. One replica is the leaseholder, which serves all reads for that range and coordinates all writes through Raft consensus. This provides strong consistency, but means traffic for any single range is serialized through one node.

When a range receives disproportionate traffic from sequential key patterns, the leaseholder node saturates. The cluster cannot redistribute this load because traffic is pinned to one range and therefore one leaseholder. Adding nodes does not help. Scaling vertically on the saturated node helps only marginally. The fix requires changing the data access pattern so traffic spreads across multiple ranges.

Load-based splitting (kv.range_split.by_load.enabled, alias kv.range_split.by_load_enabled, on by default) can help by automatically splitting hot ranges when they exceed a QPS threshold (kv.range_split.load_qps_threshold, default 2500). But load-based splitting has structural limits: it cannot split a single-row hotspot, and it cannot keep up with moving hotspots like a write cursor that continuously advances through an index.

flowchart TD
    APP[Application traffic] -->|"Sequential keys concentrate on one range"| LH
    subgraph Node1 ["Node 1: Saturated"]
        LH[Hot range leaseholder]
        LH -->|"CPU 80%+, elevated latency"| SAT[Retry storms, P99 spikes]
    end
    subgraph Node2 ["Node 2: Idle"]
        R2[Cool range replicas]
        R2 -->|"CPU 20%"| IDLE2[Unused capacity]
    end
    subgraph Node3 ["Node 3: Idle"]
        R3[Cool range replicas]
        R3 -->|"CPU 20%"| IDLE3[Unused capacity]
    end
    LH -->|"Raft replication"| R2
    LH -->|"Raft replication"| R3

Common causes

CauseWhat it looks likeFirst thing to check
Timestamp-prefixed primary keysAll recent writes land in one range that advances slowly over timeSHOW CREATE TABLE on the affected table; check if PK starts with a timestamp column
Monotonically increasing externally generated IDsWrites concentrate at the tail of the index; one node spikes while others idleCheck whether the app generates ordered IDs (e.g., Snowflake-style without random bits)
Single-row counters or status fieldsOne row receives repeated UPDATE traffic; range is indivisibleLook for UPDATE ... SET counter = counter + 1 patterns
Sequence-based hotspotsnextval() creates a single-row bottleneck even with otherwise distributed keysCheck for SQL sequences used in INSERT paths; SERIAL defaults to unique_rowid() unless sql.defaults.serial_normalization is changed
Application querying “latest” recordsReads concentrate on a narrow range of recent dataExamine query patterns for ORDER BY ... DESC LIMIT N

Quick checks

Safe, read-only diagnostics. Run from any node with SQL or HTTP access.

# Check per-node CPU asymmetry via Prometheus metrics
curl -s http://localhost:8080/_status/vars | grep -E 'sys_cpu_(user|sys)_ns'

# Check leaseholder distribution across nodes (lease metrics: leases, leases_transfers_success)
curl -s http://localhost:8080/_status/vars | grep leases

# Check transaction restart causes (look for writetooold)
curl -s http://localhost:8080/_status/vars | grep txn_restarts
-- Identify large or heavily replicated ranges (admin-only, expensive cluster-wide RPC)
-- Per-range QPS is not exposed in crdb_internal.ranges; the DB Console Hot Ranges page ranks ranges by QPS.
SELECT range_id, start_pretty, end_pretty, lease_holder, range_size
FROM crdb_internal.ranges
ORDER BY range_size DESC
LIMIT 20;
-- Check primary key definition of the affected table
SHOW CREATE TABLE <your_table>;
-- List SQL sequences that may be causing single-row bottlenecks
SHOW SEQUENCES;

The crdb_internal.ranges query is expensive. It performs a cluster-wide RPC fan-out and should be used for diagnosis only, not continuous monitoring. Do not build automated alerting on crdb_internal tables; they are unsupported and may change between versions.

How to diagnose it

  1. Confirm the asymmetry. Compare per-node CPU utilization. One node above 2x the others, with similar range counts, strongly suggests a hot range. Check lease count per store to see if leaseholder distribution is skewed.

  2. Identify the hot range. Use the DB Console Hot Ranges page to rank ranges by QPS; crdb_internal.ranges does not expose per-range QPS. Any range showing more than 10x the average QPS is hot. Note the start_pretty/end_pretty keys to identify which key range is affected.

  3. Check the DB Console. The Hot Ranges page ranks ranges by QPS, CPU, reads, and writes. Compare the top range QPS against the load-based splitting threshold to confirm the range is a split candidate.

  4. Examine the primary key. Run SHOW CREATE TABLE on the affected table. Look for monotonic patterns: timestamp-prefixed keys, externally ordered IDs, or any column whose values always increase. SERIAL defaults to unique_rowid() (time-ordered) unless sql.defaults.serial_normalization is changed.

  5. Check the transaction restart breakdown. If writetooold is the dominant restart cause, it confirms write contention on hot keys. If readwithinuncertainty appears, that points to clock skew instead, which is a different problem. See CockroachDB clock offset high: catching clock drift before self-termination.

  6. Determine if load-based splitting can help. If the hot range covers multiple rows and the hotspot is not a single row, load-based splitting may eventually split it. If the hotspot is a single row or a continuously advancing cursor, splitting will not help. The application or schema must change.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Per-node CPU utilizationAsymmetry reveals a hot leaseholder before absolute thresholds fireOne node above 2x CPU of others with similar range counts
Per-node lease countShows whether leaseholder distribution is skewedSignificant deviation from even distribution across nodes
Transaction restart rate by causewritetooold confirms write contention on hot keyswritetooold restarts elevated above baseline
SQL service latency (P50 and P99)Hot ranges cause table-specific latencyP99 elevated for specific statement fingerprints while others remain normal
Per-range QPS (DB Console Hot Ranges page)Direct measurement of traffic concentrationAny range above 10x average QPS
Key Visualizer write distributionShows which key ranges receive the most writes over timeBright bands at sequential key boundaries

For continuous monitoring, use Prometheus metrics exposed at /_status/vars. Per-node CPU asymmetry and lease count distribution are safe for per-scrape collection. Per-range QPS requires crdb_internal.ranges, which is expensive and should be polled at low frequency (every 5-10 minutes at most).

Fixes

Monotonic primary keys: short-term split, long-term hash

Short-term: ALTER TABLE ... SPLIT AT forces a manual range split at a specified value, distributing the range’s traffic across multiple leaseholders. This provides immediate relief.

-- Manual split at a specific value (short-term mitigation only)
-- This changes range distribution; expect brief latency during rebalancing.
ALTER TABLE <table> SPLIT AT VALUES (<value>);

Automatic range merging may undo the split if the resulting ranges become small and low-traffic. If merging is a problem, consider raising range_min_bytes in the zone configuration for the affected table. Manual splits are a tactical fix, not a permanent solution.

Long-term: hash-sharded indexes (GA since v21.2; v22.1+ uses virtual shard columns that skip backfill) are the recommended structural fix for monotonic primary keys. They distribute writes across multiple buckets by hashing the key.

-- Hash-sharded primary key (available since v21.2)
CREATE TABLE events (
    id BIGINT SERIAL,
    data BYTES,
    PRIMARY KEY (id ASC) USING HASH WITH (bucket_count = 8)
);

Alternatively, UUID-based primary keys distribute writes across the keyspace without sequential concentration. Note that unique_rowid() (the default for SERIAL) is time-ordered, so it can still create a moving hotspot at high write rates.

Single-row hotspots: change the application pattern

If a single row is the bottleneck (for example, a global counter or status field updated by all transactions), the range is indivisible. Neither SPLIT AT nor hash-sharded indexes can help. The application pattern must change:

  • Move counters to a side table with multiple rows (sharded counter pattern) and aggregate on read.
  • Use application-level caching for frequently updated values.
  • Batch updates to reduce per-transaction contention on the hot row.

Sequence-based hotspots: cache or replace

SQL sequences create a single-row bottleneck because nextval() increments one row through its leaseholder. Even with a well-distributed primary key, the sequence itself is a hotspot.

-- Increase sequence caching to reduce leaseholder round-trips
ALTER SEQUENCE <seq_name> CACHE 100000;

For high-throughput insert paths, consider replacing sequences with UUID-based keys that distribute across the keyspace without a centralized counter.

Lease placement tuning for multi-region

For multi-region deployments, zone configurations can influence where leaseholders land:

-- Hint where leaseholders should live (not a hard constraint)
ALTER TABLE <table> CONFIGURE ZONE USING
    lease_preferences = '[[+region=us-east-1]]';

lease_preferences are hints, not hard constraints. The allocator may not satisfy them if the preferred region lacks sufficient replicas or is behind quorum. Do not rely on lease preferences as the sole fix for hot ranges; they address placement, not traffic concentration.

Prevention

  • Audit primary key design during schema review. Any monotonically increasing key pattern is a future hot range. This includes timestamp-prefixed keys, externally ordered IDs, and sequence-backed columns; SERIAL defaults to unique_rowid(), which is itself time-ordered. Prefer hash-distributed or UUID-based keys for high-write tables.

  • Watch for moving hotspots. Queue-like workloads that write to the tail of an index create a continuously advancing hotspot. Load-based splitting cannot keep up with this pattern. Hash-sharded indexes are the only effective structural fix.

  • Monitor per-node CPU asymmetry continuously. The earliest warning of a hot range is one node running significantly hotter than others. Alert on sustained asymmetry (above 2x for more than 10 minutes) even if absolute CPU is not critical.

  • Review hot ranges periodically. A weekly look at the DB Console Hot Ranges page, plus a crdb_internal.ranges check ordered by range_size, can reveal emerging hotspots before they cause user-visible degradation.

  • Use the Key Visualizer in DB Console. It shows write distribution across the keyspace over time. Bright bands at sequential key boundaries indicate hotspot formation. Regular review catches patterns that point-in-time metrics miss.

How Netdata helps

  • Per-node CPU asymmetry detection. Netdata collects sys_cpu_user_ns and sys_cpu_sys_ns per node with per-second granularity. A hot range shows up immediately as one node diverging from the others.

  • Lease count distribution. By tracking lease count per store, Netdata makes leaseholder imbalance visible at a glance. A sudden shift in lease distribution often accompanies hot range formation or lease rebalancing churn.

  • Transaction restart correlation. Netdata surfaces txn_restarts broken down by cause. Correlating a writetooold spike with per-node CPU asymmetry confirms a hot range diagnosis without querying crdb_internal tables during an incident.

  • SQL latency segmentation. Per-node sql_service_latency histograms let you see which nodes are experiencing elevated P99. When one node’s P99 diverges from the cluster, the leaseholder on that node is likely serving a hot range.

  • Historical baselining. Netdata retains high-resolution historical data, making it possible to identify when asymmetry began and correlate it with deployment events, schema changes, or traffic pattern shifts.

Netdata’s database monitoring brings these signals together with per-second metrics.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.