The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-monitoring-checklist

Operations Guides

CockroachDB monitoring checklist: the signals every production cluster needs

CockroachDB layers SQL execution on a replicated KV store backed by Pebble LSM trees, Raft consensus, and MVCC concurrency control. Each subsystem has distinct failure modes, and interactions between them create cascades that single-signal monitoring cannot catch. A cluster can show healthy CPU, adequate disk space, and sub-millisecond SQL latency while L0 sublevels climb toward write stalls or clock offset drifts toward the self-termination threshold.

This checklist defines four monitoring maturity levels (survival, operational, mature, expert) and maps every production signal to its level. Use it to audit current coverage, identify gaps before they cause incidents, and plan what to add as your cluster grows.

The four monitoring levels

The levels are cumulative. Each builds on the one below it.

flowchart BT
    L1["Level 1: Survival
5 signals: is it broken?"] L2["Level 2: Operational
16 more: is it healthy?"] L3["Level 3: Mature
20 more: leading indicators"] L4["Level 4: Expert
18 more: deep internals"] L1 --> L2 --> L3 --> L4
  • Survival: is the cluster fundamentally broken? If you have nothing else, have these.
  • Operational: is the cluster healthy, not just alive? A team running CockroachDB in production should have all of these.
  • Mature: leading indicators and storage internals that give 10 to 30 minutes of warning before degradation becomes an outage.
  • Expert: deep internals and per-range analysis. These are the signals operators add after repeated incidents, when a failure mode was invisible until too late.

Level 1: survival

SignalMetric / sourceWhat it tells youAlert posture
Node liveness/_status/nodes, /health?ready=1Whether the cluster considers each node alive. Loss triggers lease redistribution.PAGE: unexpected not-live transition sustained over 5 min, node not draining
Range unavailabilityranges_unavailableRanges with no leaseholder or lost Raft quorum. Any nonzero value means some keyspace is unreadable or unwritable.PAGE: any nonzero value sustained over 5 min
Disk spacecapacity_availableFree space per store. CockroachDB recommends keeping below 60 percent utilization (at least 40 percent free). Running out prevents compaction and triggers a storage death spiral.PAGE: under 10 percent available and decreasing. TICKET: under 30 percent
SQL connectivitySynthetic probe: cockroach sql -e "SELECT 1"Whether clients can connect and execute queries through the full pgwire path (TCP, TLS, SQL execution).PAGE: probe failure sustained over 2 min
Certificate expirationsecurity.certificate.expiration.* metrics from /_status/vars, openssl x509 -enddateTime until TLS certificates expire. Expiry breaks inter-node communication or client access.TICKET: under 30 days. PAGE: expired or under 24 hours with no replacement staged

Use /health?ready=1 for load balancer health checks. It returns 503 when a node is draining or quorum is lost. Plain TCP checks will route traffic to functionally impaired nodes.

The liveness mechanism is evolving across versions, transitioning toward store-level liveness and lease-based failure detection. Check the release notes for your deployment version to understand the exact heartbeat and expiry behavior.

Level 2: operational

Everything in Level 1, plus:

SignalMetric / sourceWhat it tells you
SQL statement latency (P50, P99)sql_service_latencyClient-visible query performance. P99 is where SLO violations live.
Transaction throughputsql_select_count, sql_insert_count, sql_update_count, sql_delete_countWorkload volume and composition. Sudden drops need correlation with error rate to distinguish traffic decrease from system failure.
Transaction restart/retry ratetxn_restarts (by cause)Contention level. writetooold means schema contention, readwithinuncertainty means clock skew, txnpush means application conflicts. Each requires a different response.
SQL error ratesql_failure_countActive failures visible to applications. Watch for XX000 (internal errors) and 53200 (resource exhaustion).
CPU utilizationsys_cpu_user_ns, sys_cpu_sys_nsCapacity headroom. Per-range Raft ticking is a hidden CPU multiplier: 50,000 ranges costs significantly more CPU than 5,000 at identical query load.
Memory RSSsys_rssProximity to memory limits. CGo allocations (Pebble block cache) grow independently of the Go heap.
LSM L0 sublevel countstorage_l0_sublevelsThe single most predictive storage signal. 0 to 5 is healthy. 10 to 20 means active degradation. Over 20 means write stalls are imminent or active.
Active client connectionssql_connsConnection pressure. Each connection consumes a goroutine and memory. Over 1,000 to 2,000 per node suggests client pooling is misconfigured.
Under-replicated range countranges_underreplicatedReplication safety margin. Sustained under-replication means the cluster cannot heal fast enough.
Clock offsetclock_offset_meannanosTime sync health. Over 80 percent of --max-offset (default 500ms, so over 400ms) triggers self-termination. Even steady 200ms offset doubles the uncertainty interval and causes persistent read restarts.
Inter-node RPC latencyround_trip_latencyNetwork health as CockroachDB experiences it. Includes kernel scheduling delay, so CPU-saturated nodes show elevated RPC latency even with a healthy network.
Disk I/O utilization and latencyOS: iostat -xz 1, /proc/diskstatsWhether storage is the bottleneck. On cloud volumes, watch against provisioned IOPS caps, not just percentage utilization.
WAL fsync latencystorage_wal_fsync_latencyMost direct measure of write-path health. WAL fsync is on the critical path of every write. Over 50ms on SSDs signals I/O saturation.
Job statuscrdb_internal.jobs, per-type Prometheus metricsWhether backups, schema changes, and other long-running jobs are progressing or stuck. Track backup duration trend, not just success.
Admission control queue depthadmission.wait_durations.*, admission.io.overloadWhether internal flow control is throttling. The store-write queue begins shaping traffic at approximately 5 L0 sublevels.
Node uptimesys_uptimeCold-start gating signal. Suppress most performance alerts during the first 10 minutes after process start. Block cache warmup takes 10 to 30 minutes.

Level 3: mature

Everything in Level 2, plus leading indicators:

SignalMetric / sourceWhat it tells you
LSM read amplificationrocksdb_read_amplificationSSTable files consulted per read. Name retained from RocksDB era. 10 to 15 is normal. Over 25 indicates compaction debt.
Compaction throughput and backlogrocksdb_compactions, storage_marked_for_compaction_files, storage_l0_num_filesWhether background storage maintenance keeps pace with writes. Upward backlog trend over hours means approaching write capacity.
Write stall countstorage_write_stallsPebble refusing writes because LSM state is unsafe. Rate over 1 per second sustained for over 1 minute means foreground writes are materially impaired.
Pebble block cache hit raterocksdb_block_cache_hits, rocksdb_block_cache_missesCache efficiency. Drop below 90 percent impacts read-heavy workloads. Sudden drop after restart is expected (cold cache).
SQL memory budget utilizationsql_mem_root_currentMemory available for query execution. Over 90 percent risks error 53200. Per-node, not per-query: one large query can starve all others.
Go GC pause duration and frequencysys_gc_pause_ns, sys_gc_countPauses approaching the liveness heartbeat interval (default 3s) risk node death. GC CPU over 15 percent indicates memory pressure.
Goroutine countsys_goroutinesConcurrency level. Monotonic growth without workload increase means leak. Over 100,000 is critical.
Range count per noderanges, leases_countPer-node overhead from Raft ticking. Over 50,000 ranges per node creates significant CPU cost independent of query load.
Raft snapshot raterange_snapshots_generated, range_snapshots_applied_initialReplication health. High rate beyond rebalancing means nodes cannot keep up with log application. Each snapshot transfers up to 512 MiB.
Lease transfer rateleases_transfers_successSystem churn. Each transfer creates a brief unavailability window. Elevated rate without operational cause indicates instability.
Intent count and bytesintentcount, intentbytesUnresolved write intents from uncommitted transactions. Growing count means abandoned transactions or intent resolution falling behind.
MVCC garbage bytesMVCC metrics via /_status/varsDead data awaiting GC. Grows silently and eventually consumes disk. Protected timestamps from CDC or backups can block GC entirely.
KV read/write latencyexec_latencyStorage layer latency isolated from SQL planning overhead. Rising without SQL changes means storage degradation.
Raft log commit latencyraft.process.logcommit.latencyWAL fsync path specifically. Over 50ms on SSDs is a strong I/O saturation signal.
Changefeed lag (if using CDC)changefeed_max_behind_nanosCDC consumer freshness. A stalled changefeed creates protected timestamps that silently prevent MVCC GC, causing unbounded disk growth.
Protected timestamp count and agespanconfig_kvsubscriber_protected_record_count, jobs_changefeed_protected_age_secRecords preventing MVCC GC. Age over 24 hours or count growing without active backup or CDC means a stalled operation is blocking cleanup.
File descriptor usagesys_fd_open, /proc/<pid>/fdFD exhaustion prevents new connections and SSTable opens. Production should have ulimit of at least 35,000 to 65,000.
Composite pattern alertingMulti-signal correlationsDetects cascades that individual thresholds miss. Example: L0 sublevels climbing plus write stalls plus lease transfers equals a compaction death spiral.
Per-table and per-index statisticscrdb_internal queries (diagnostic use only)Size, row count, and query patterns for capacity planning. Do not use crdb_internal tables for automated monitoring pipelines; they are unsupported and may change between versions.
Network throughput between nodesOS: /proc/net/devInter-node bandwidth saturation. CockroachDB multiplexes all traffic on a single port, so Raft heartbeats cannot be prioritized over data transfer.

Level 4: expert

Everything in Level 3, plus signals operators add after repeated incidents:

SignalWhat it tells you
Raft proposal drop rateDropped proposals mean silently lost writes (retried, but immediate latency impact). Rarely watched.
Per-range request rate distributionDetects hot ranges before they become obvious from cluster-level metrics. A range at 10x average QPS bottlenecks its leaseholder while other nodes idle.
Intent resolution throughputSeparate from intent count. Tells you whether cleanup is keeping pace during an intent accumulation cascade.
Lease preference violation countWhether the allocator can satisfy zone config lease preferences. Critical for multi-region latency SLOs.
Raft entry cache hit rateCache misses cause disk reads for log entries, a hidden latency source.
Compaction debt by LSM levelShows where in the LSM tree pressure is building, not just the aggregate.
Goroutine profiling (debug/pprof)Stack traces for all goroutines. Detects where goroutines are stuck (I/O, lock contention) before count becomes critical.
KV write batch size distributionChanges in write patterns that precede L0 pressure or compaction issues.
SQL plan cache hit rateCache misses mean re-planning, which is CPU-intensive.
Cross-range transaction percentageHigher percentage means more distributed coordination overhead and latency.
Time-series data retention pressureCockroachDB’s internal time-series data can itself become a storage burden on large clusters.
Node decommission progress rateEnsures decommissions complete before patience runs out. Stalled decommissions block maintenance.
Admission control token exhaustion rateHow frequently work is being delayed, broken down by queue type. Foreground queue exhaustion is more concerning than elastic queue.
Closed timestamp lagAffects follower read freshness in multi-region setups.
Queue processor error countsSplit, merge, replicate, and GC queue failures indicate internal scheduling problems.
Disk stall detection metricsstorage_disk_stalled (nonzero means the node may self-terminate) and storage_disk_slow (incrementing means slow disk operations).
Non-gateway SQL activityConnections arriving at nodes not designated as application gateways. May indicate unauthorized direct access or misconfigured routing.
Bulk data export activityUnexpected EXPORT or BACKUP operations. May indicate data exfiltration.

How to alert on these signals

Not every signal should trigger a page. The playbook distinguishes PAGE (wake someone up) from TICKET (investigate during business hours).

Gate on sustained duration. Most conditions require sustained presence (typically over 5 minutes) before alerting. Transient spikes from range splits, lease transfers, or brief I/O contention are normal operating noise.

Gate on node uptime. Suppress most performance alerts during the first 10 minutes after process start (sys_uptime). Cache warmup takes 10 to 30 minutes. Statistics collection after restart may produce suboptimal query plans temporarily.

Binary signals page immediately. ranges_unavailable is zero under normal conditions. Any nonzero value sustained over 5 minutes is live user impact. storage_disk_stalled is similarly binary: nonzero means the node may self-terminate.

Workload-shaped signals require context. CPU at 80 percent alone is not page-worthy. CPU at 80 percent plus liveness flapping plus elevated lease transfers is a cascade. Latency thresholds are workload-dependent: establish baselines per statement fingerprint and alert on deviation (over 2x rolling P99 is a reasonable starting point), not on absolute values.

Distinguish restart causes. Alarm on txn_restarts breakdown, not the aggregate. readwithinuncertainty appearing for the first time is a clock problem. writetooold climbing is a schema or contention problem. txnpush spiking is an application transaction design problem. Each routes to a different team.

Treat recovery activity with suspicion. Snapshot and rebalance storms during cluster healing compete with foreground traffic. Recovery I/O can degrade the cluster further. Monitor recovery rate and foreground latency together.

What most teams miss

These gaps appear most often in incident reviews:

  • L0 sublevel count uninstrumented until write stalls. The most common and damaging gap. L0 gives 10 to 30 minutes of warning before stalls, and teams consistently waste that window because the metric was never collected.
  • Clock offset not monitored proactively. NTP is treated as set-and-forget. The first sign of trouble is a node self-terminating at 3 a.m. Meanwhile, readwithinuncertainty restarts have been degrading tail latency for days.
  • Aggregate latency instead of per-statement-fingerprint latency. A new slow query among fast queries barely moves P99 but kills specific endpoints. Per-fingerprint tracking catches plan regressions and missing indexes.
  • Transaction restart causes not distinguished. Total retry rate is alarmed without breakdown, wasting diagnosis time because the three main causes need completely different responses.
  • MVCC garbage and protected timestamps unmonitored. Data gets deleted but nobody checks whether GC reclaims space. Protected timestamps from stalled CDC or hung backups silently prevent GC. Disk fills with dead data.
  • Range count not tracked as a scaling dimension. Operators monitor data volume and query throughput but miss that 200,000 ranges per node behaves fundamentally differently from 20,000 because of per-range Raft processing overhead.
  • TCP health checks instead of /health?ready=1. Load balancers route traffic to draining, write-stalled, or GC-thrashing nodes because the TCP check succeeds.
  • Backup duration trend not tracked. Backup succeeds so the check is green. But duration grew from 20 minutes to 3 hours. Eventually it exceeds the backup interval, creating overlapping backups or gaps.
  • Admission control ignored as a capacity signal. Regular AC queuing means zero burst headroom. The system is maintaining throughput by adding latency, not by having spare capacity.
  • Cluster averages trusted over per-node data. One hot leaseholder or overloaded region hides under healthy global metrics. Always monitor per-node.

How Netdata helps

  • Per-second collection from CockroachDB’s /_status/vars endpoint means brief write stalls and lease transfers are visible, not hidden by 15 to 30 second scrape gaps.
  • ML anomaly detection on storage_l0_sublevels, read amplification, and write stalls catches slow drift toward compaction death spirals before static thresholds fire.
  • Cross-signal correlation links clock offset spikes to readwithinuncertainty restart rates, or L0 growth to admission control queue depth, so you see the cascade instead of isolated symptoms.
  • Per-node dashboards prevent hot ranges and single-node degradation from hiding under cluster averages.
  • Cold-start awareness uses sys_uptime to suppress false positives during the warmup window after node restarts.

Netdata’s CockroachDB monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.