The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-disk-space-running-out

Operations Guides

CockroachDB disk space running out: capacity_available trends and the 20% rule

When a CockroachDB store runs out of disk space, the node cannot accept writes, cannot compact its LSM tree (the operation that would reclaim space), and enters a downward spiral that typically requires operator intervention. The capacity_available metric tracks remaining free space per store, but raw free space alone does not tell the full story. MVCC garbage, protected timestamps, compaction space amplification, and replica rebalancing all consume space faster than headline data growth suggests.

The 20% rule is the operational floor: keep at least 20% of each store’s capacity free at all times. CockroachDB documentation recommends this minimum, with 30% preferred to absorb compaction amplification and Raft snapshot staging. Below 15% free space, compaction may be unable to run, creating an unrecoverable loop where the system needs free space to free space.

This article covers how to read capacity_available trends, distinguish genuine data growth from MVCC garbage accumulation, estimate runway, and respond when a node approaches the threshold.

What this means

Disk space exhaustion in CockroachDB is an availability problem, not just a capacity problem. When a store fills past the point where compaction can function:

  • Writes stop. Raft log entries cannot be persisted. The node loses leadership of its ranges.
  • The cluster compensates by up-replicating. Surviving nodes receive new replicas, each requiring disk space. If those nodes are also low on space, one full node can cascade into multiple full nodes.
  • Compaction stalls. Pebble needs temporary free space to merge SSTables across LSM levels. A worst-case compaction of the largest level can require free space equal to that level’s size. Without headroom, the LSM tree cannot be optimized, read amplification climbs, and the node degrades further.
  • The node may not restart. If the WAL cannot be replayed due to insufficient space, the process cannot complete startup.
flowchart TD
    A[capacity_available declining] --> B{Free > 15%?}
    B -- No --> C[Compaction stalls]
    C --> D[L0 sublevels climb]
    D --> E[Write stalls]
    E --> F[Node loses Raft leases]
    F --> G[Cluster up-replicates ranges]
    G --> H[Remaining nodes lose free space]
    H --> I[Multi-node disk-full cascade]
    B -- Yes --> J[System stressed but stable]

The cascade risk is why disk space thresholds are set conservatively. A single node hitting 100% utilization is not an isolated incident. The cluster’s self-healing mechanism (up-replication) actively makes the problem worse by consuming space on other nodes to restore replica counts.

Common causes

CauseWhat it looks likeFirst thing to check
Genuine data growthcapacity_used rising proportionally to live_bytesCompare live_bytes trend to total bytes trend
MVCC garbage accumulationcapacity_used rising while live_bytes is flat or decliningCheck protected timestamp records and changefeed job status
Compaction falling behindstorage_l0_sublevels climbing alongside disk usage growthCheck disk I/O utilization and compaction throughput
Protected timestamp stallProtected timestamp age growing, no active backup or CDC job progressingjobs_changefeed_protected_age_sec and job status
Raft snapshot stagingDisk usage spike correlating with rebalancing or node recoveryrange_snapshots_generated rate
Schema change temp filesDisk growth during ALTER TABLE or CREATE INDEX backfillCheck crdb_internal.jobs for running schema changes

Quick checks

All of these are read-only and safe to run at any time.

# Check capacity metrics per store
curl -s http://localhost:8080/_status/vars | grep -E '^capacity'

# OS-level disk usage for the store directory
df -h /path/to/cockroach-data

# Compare live data to total bytes to estimate MVCC garbage ratio
curl -s http://localhost:8080/_status/vars | grep -E 'live_bytes|key_bytes|value_bytes'

# Check for protected timestamp records blocking GC
curl -s http://localhost:8080/_status/vars | grep -i protected

# Check compaction health
curl -s http://localhost:8080/_status/vars | grep storage_l0_sublevels

# Check for under-replicated ranges (indicates recovery in progress)
curl -s http://localhost:8080/_status/vars | grep ranges_underreplicated

How to diagnose it

  1. Identify which stores are approaching the threshold. Check capacity_available per store. The metric is per-store, not per-node. A node with multiple stores can have one store near capacity while another has ample headroom.

  2. Determine whether growth is data or garbage. Compare live_bytes to capacity_used. If live_bytes is stable but capacity_used is climbing, the growth is MVCC garbage: old versions and tombstones that have not yet been compacted away. This distinction determines your response. Garbage accumulation is reversible by fixing the GC blocker. Genuine data growth requires adding capacity.

  3. Check for protected timestamp records. Protected timestamps from CDC changefeeds, backups, and other features prevent MVCC garbage collection below their timestamp. A stalled changefeed or hung backup creates a protected timestamp that silently blocks GC. Disk usage grows unboundedly while live_bytes stays flat. This is one of the most insidious failure modes in CockroachDB because it is silent until the disk is full. Check spanconfig_kvsubscriber_protected_record_count and jobs_changefeed_protected_age_sec.

  4. Check compaction health. If storage_l0_sublevels is elevated, compaction is falling behind. Disk I/O is saturated or disk space is too low for compaction to proceed. L0 growth and disk space exhaustion compound each other: high L0 consumes more disk, and low disk prevents compaction from reducing L0.

  5. Estimate runway. Use linear extrapolation of capacity_available decline rate, but account for MVCC garbage separately. If GC is keeping up, the decline rate reflects true data growth. If GC is blocked (protected timestamps, GC TTL too long), the decline rate is inflated by garbage accumulation and will stabilize once the blocker is removed. Time to 80% utilization at the current rate gives your operational window. Anything under 30 days warrants a capacity plan.

  6. Check for cascade conditions. If ranges_underreplicated is nonzero and the cluster is actively up-replicating, surviving nodes are receiving replica snapshots that consume disk space. Verify that target nodes have adequate free space before the cluster fills them trying to heal.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
capacity_available per storeFree space remaining. The primary depletion metric.Below 20% with downward trend: ticket. Below 10% and declining: page.
live_bytes vs capacity_used ratioReveals MVCC garbage overhead. Diverging ratio means space is consumed by dead data.live_bytes flat while capacity_used climbs.
storage_l0_sublevels per storeCompaction health. Elevated L0 means compaction is falling behind, which both consumes extra disk and requires more free space to resolve.Above 10 sustained: compaction lagging. Above 20: write stalls imminent.
spanconfig_kvsubscriber_protected_record_countProtected timestamp records blocking GC. Each record prevents garbage collection below its timestamp.Records present without active backup or CDC operation.
jobs_changefeed_protected_age_secAge of protected timestamps from changefeeds. Growing age means GC is blocked.Above 24 hours, or growing while changefeed is stalled.
ranges_underreplicatedRecovery in progress. Up-replication consumes disk on target nodes.Nonzero during a disk space incident: cascade risk.
storage_write_stalls per storePebble refusing writes. May indicate compaction cannot proceed due to space constraints.Any nonzero in production during normal workload.

Fixes

Emergency: node has already filled up

If a node has shut down due to disk exhaustion, look for an emergency ballast file in the node’s storage directory. If one was pre-created (either manually via cockroach debug ballast or automatically during a prior startup), deleting it frees reserved space that may allow the node to start.

Warning: Only delete the ballast file on a stopped node, and only when the disk is genuinely full. Deleting it on a running node removes your emergency reserve.

Once the node is back, add permanent storage or reduce data immediately. The node will recreate the ballast file on the next successful startup, consuming the space you just freed.

MVCC garbage is the culprit

If live_bytes is flat but capacity_used is climbing, the growth is garbage. Steps:

  1. Check for stalled or failed changefeed jobs and backups. Cancel or resume them to release protected timestamp records that are blocking GC.
  2. Verify GC TTL (gc.ttlseconds) is not set excessively high. Longer TTL means more historical versions retained on disk.
  3. Wait for GC to run. It is asynchronous. Monitor live_bytes to capacity_used convergence over the following hours.
  4. After GC releases old versions, compaction must still run to reclaim physical space. Monitor storage_l0_sublevels to confirm compaction is progressing and not blocked by remaining space constraints.

Compaction blocked by low space

If the store is so full that compaction cannot proceed, the system is in a feedback loop. Options:

  1. Free space by any means available: delete the ballast file, clean up temporary files, rotate logs.
  2. Reduce write rate to slow L0 growth while compaction catches up. Pause bulk operations, tighten admission control.
  3. If the node is part of a healthy cluster, consider decommissioning it to move data to nodes with more capacity. Verify that target nodes have adequate free space first; the allocator avoids stores that are effectively full, but it does not reserve headroom for future growth.

Genuine capacity exhaustion

If live_bytes is growing proportionally to capacity_used, the data is genuinely outgrowing the provisioned storage:

  1. Expand the volume if the storage backend supports online expansion.
  2. Add nodes to the cluster. New nodes receive rebalanced replicas, spreading data across more storage.
  3. Review data retention. If old data can be archived or dropped, do so. Note: DROP TABLE and large DELETE operations do not immediately free space. MVCC tombstones must propagate through compaction before physical space is reclaimed. A large DELETE can temporarily increase disk usage before it decreases.

Prevention

  • Alert on capacity_available per store at two thresholds. Ticket at 20% free with a downward trend. Page at 10% free and declining. These thresholds give hours to days of runway before cascade risk becomes acute.
  • Track the live_bytes to capacity_used ratio over time. A diverging ratio means garbage is accumulating faster than compaction can reclaim it. Catch this before it fills the disk.
  • Monitor protected timestamp records. A stalled changefeed or hung backup can silently block GC indefinitely. Alert on protected timestamp age above 24 hours when no active job explains it.
  • Monitor L0 sublevels alongside disk usage. Elevated L0 means compaction is falling behind, which increases disk space pressure and requires more free space to resolve.
  • Plan capacity with linear extrapolation. Project capacity_available decline using the current growth rate. Account for MVCC garbage by tracking live_bytes growth as the true minimum floor. If the garbage-adjusted runway is under 30 days, plan storage expansion.
  • Verify free space before decommissioning nodes. The cluster moves replicas to surviving nodes during decommission. Ensure each surviving node has enough capacity to absorb the additional data plus compaction headroom.

How Netdata helps

  • Per-second capacity_available tracking per store lets you compute depletion rates with enough resolution to catch sudden growth spikes that 15-minute snapshots miss.
  • Correlate capacity with L0 sublevels, write stalls, and MVCC garbage metrics on the same timeline. When disk usage spikes, seeing whether L0 is also climbing tells you immediately whether compaction is the bottleneck or data growth is the driver.
  • Protected timestamp age and changefeed lag displayed alongside disk usage trends surface the silent GC-stall pattern before it fills the disk.
  • ML anomaly detection on capacity_used growth rate flags acceleration that static thresholds miss. A store growing 1 GB per day that suddenly starts growing 5 GB per day is an early warning, even if absolute free space is still above 20%.
  • Per-store granularity ensures you see the one store approaching capacity even when the cluster aggregate looks healthy.

Netdata’s CockroachDB monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.