The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-store-is-full

Operations Guides

CockroachDB store is full: emergency disk recovery and why deletes don't free space

A CockroachDB node’s store directory is full or nearly full, and the node may have already shut itself down to protect data integrity.

Deleting data will not help immediately. CockroachDB uses MVCC: a SQL DELETE writes tombstones at new timestamps rather than removing old data. Until those tombstones propagate through Pebble compaction, the old versions stay on disk. A large-scale delete can temporarily increase disk usage because the tombstones are written before old data is reclaimed.

The underlying problem is usually silent accumulation: MVCC garbage that GC never reclaimed, protected timestamps from stalled changefeeds or hung backups that blocked GC entirely, or capacity exhaustion that nobody was watching.

What this means

CockroachDB needs free disk space for three concurrent activities: Pebble compaction (which merges SSTables and reclaims obsolete data), Raft snapshot staging (up to 512 MiB per range during rebalancing), and WAL writes. When available space drops below approximately 15%, compaction loses the working room it needs to merge files. Below 10%, the system enters a death spiral it may not escape.

CockroachDB has built-in store fullness thresholds. At 92.5% utilization, the allocator stops rebalancing new replicas to that store. At 95%, it begins actively moving replicas away with high priority. When the storage engine can no longer write — a filesystem disk-full error, or available space below half the ballast size (512 MiB with the default 1 GiB ballast) — the node shuts down to protect data integrity.

Once a node shuts down due to disk pressure, you cannot simply restart it. WAL replay and compaction startup need free space that does not exist.

flowchart TD
    A[Store fills past 85%] --> B[Compaction space amplification spikes]
    B --> C[Compaction throughput drops]
    C --> D[L0 sublevels accumulate]
    D --> E[Reads slow, compaction slows further]
    E --> C
    D --> F[Free space drops below 15%]
    F --> G[Compaction cannot obtain working space]
    G --> H[Write stalls begin]
    H --> I[Node cannot commit Raft entries]
    I --> J[Liveness lost, leases transferred away]

The feedback loop is the core problem. Compaction reclaims space, but compaction itself needs free space to write merged output. When space is exhausted, the fix mechanism is also blocked.

Common causes

CauseWhat it looks likeFirst thing to check
Genuine data growthLive bytes and total bytes growing together; MVCC garbage ratio stablecapacity_used trend over weeks; is growth proportional to traffic?
MVCC garbage accumulationTotal bytes growing, live bytes stable or shrinkingkey_bytes + val_bytes - live_bytes per range; gc.ttlseconds setting
Protected timestamps blocking GCMVCC garbage growing, GC not progressing, protected timestamp records presentspanconfig_kvsubscriber_protected_record_count; changefeed job health
Raft snapshot stagingSudden spike during rebalancing or node decommissionrange_snapshots_generated rate; under-replicated range count
Schema change temp filesSpike correlates with ALTER TABLE or CREATE INDEX runningcrdb_internal.jobs for schema changes in running state
Time-series data growthSteady growth on a new or quiet clusterInternal metrics retention: 10s granularity for 10 days, 30m for 90 days

Quick checks

Run these read-only checks on a node that is still responding:

# Check capacity per store from the metrics endpoint
curl -s http://localhost:8080/_status/vars | grep -E 'capacity'

# OS-level disk space on the store directory
df -h /path/to/cockroach-data

# Check L0 sublevels (compaction health)
curl -s http://localhost:8080/_status/vars | grep storage_l0_sublevels

# Check write stalls (most severe storage signal)
curl -s http://localhost:8080/_status/vars | grep storage_write_stalls

# Check read amplification (compaction debt indicator)
curl -s http://localhost:8080/_status/vars | grep rocksdb_read_amplification

# Check protected timestamp records (GC blockers)
curl -s http://localhost:8080/_status/vars | grep -E 'protected_record_count|protected_age_sec'

# Check if the node is still running
curl -s http://localhost:8080/_status/vars | grep sys_uptime

If the node has already shut down, the metrics endpoint will not respond. Fall back to df -h on the store directory from the host OS.

How to diagnose it

  1. Determine which stores are critical. Check capacity_available per store. A node with multiple stores can have one at 98% and another at 60%. Focus recovery on the critical store.

  2. Separate data growth from garbage. If capacity_used is growing but live data is stable, the problem is MVCC garbage or protected timestamps, not data volume. Check the ratio of live bytes to total bytes per range to see how much on-disk data is garbage awaiting collection.

  3. Check for protected timestamps. This is the most insidious cause. A stalled changefeed or hung backup creates a protected timestamp record that prevents MVCC GC below that timestamp. Disk grows silently. Check spanconfig_kvsubscriber_protected_record_count and jobs_changefeed_protected_age_sec. If protected age exceeds 24 hours, investigate the owning job immediately.

  4. Assess compaction health. Check storage_l0_sublevels. Above 20 sustained means the store is already in or approaching a write stall. Check storage_write_stalls for active stall events. If L0 is high but decreasing, compaction is catching up. If it is high and rising, the death spiral is active and you need to reduce write pressure.

  5. Identify space consumers. If the node is still running, check which tables or ranges consume the most space. Large DELETE or UPDATE operations create tombstone pressure that temporarily inflates on-disk size before compaction reclaims it.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
capacity_availableCore metric: how much space remainsBelow 20% is yellow; below 10% with upward trend is critical
storage_l0_sublevelsCompaction health indicatorAbove 10 sustained means compaction lagging; above 20 means stalls imminent
storage_write_stallsActive write unavailabilityAny nonzero rate during normal workload
rocksdb_read_amplificationLSM tree organization healthAbove 25 sustained indicates compaction debt or tombstone accumulation
spanconfig_kvsubscriber_protected_record_countProtected timestamps blocking GCRecords persisting without active backup or CDC operation
jobs_changefeed_protected_age_secAge of CDC protected timestampsAbove 24 hours without an active changefeed is a silent disk growth trigger
MVCC garbage (total bytes minus live bytes)Dead data awaiting GCGrowing without bound over days
range_snapshots_generatedRebalancing and recovery I/OElevated without operational cause

Fixes

Delete the emergency ballast file

CockroachDB creates an emergency ballast file in each node’s store directory as a space reservation. The file exists specifically so that deleting it frees enough space to recover a node that shut down due to a full disk.

# Find the ballast file in the store directory
find /path/to/cockroach-data -name '*ballast*' -ls

# Delete it to free emergency space
# Non-destructive: this is a pre-allocated placeholder, not database content
rm /path/to/cockroach-data/<ballast_file_name>

This is non-destructive to data. The ballast file is a pre-allocated placeholder, not database content. After deleting it, the node should have enough free space to start and run compaction.

Once the node recovers, CockroachDB recreates the ballast automatically during periodic capacity calculations when there is sufficient free space (at least 4x the ballast size, or when more than 10 GiB would remain free). The node is not fully recovered until the ballast is reestablished.

Lower gc.ttlseconds and force garbage collection

If MVCC garbage is the problem, reducing the GC TTL accelerates reclamation. The gc.ttlseconds zone configuration parameter determines how long old MVCC versions are retained before becoming eligible for GC.

-- Lower GC TTL to accelerate garbage reclamation
-- WARNING: This reduces your ability to run AS OF SYSTEM TIME queries
ALTER RANGE default CONFIGURE ZONE USING gc.ttlseconds = 600;

After changing the setting, the MVCC GC queue processes eligible ranges automatically, but pace depends on GC queue throughput.

For emergency recovery after dropping a large table, the CockroachDB operational FAQ documents a manual procedure:

  1. Lower gc.ttlseconds to 600 on the relevant zone.
  2. Capture range IDs before dropping: SHOW RANGES FROM TABLE your_table;
  3. Drop the table.
  4. Manually enqueue each range in the MVCC GC queue via the DB Console Advanced Debug page with SkipShouldQueue checked.

This is aggressive and tedious for large tables. Use it only when standard GC cannot keep pace with the emergency.

Cancel stalled changefeeds and hung backups

If protected timestamps are blocking GC, address the owning job. A stalled changefeed keeps a protected timestamp that prevents all GC below its timestamp. Disk grows unboundedly while every other cluster metric looks healthy.

-- Check changefeed job health
SELECT job_id, status, running_status, created
FROM crdb_internal.jobs
WHERE job_type = 'CHANGEFEED'
  AND status NOT IN ('succeeded', 'canceled')
ORDER BY created DESC;

-- Cancel a stalled changefeed
-- WARNING: Disruptive to downstream consumers
CANCEL JOB <job_id>;

After canceling, the protected timestamp is released and MVCC GC resumes. Monitor spanconfig_kvsubscriber_protected_record_count to confirm the record is removed and watch disk usage decline as compaction reclaims the freed garbage.

Add storage capacity

The cleanest long-term fix. Add a new store to the node (if the deployment supports it) or add new nodes to the cluster. New nodes receive replicas via rebalancing, redistributing data and relieving pressure on the full store.

Adding capacity takes time. The cluster must generate and transfer Raft snapshots (up to 512 MiB per range) to the new storage. During this process, the recovering store still needs free space to operate. Start this process early, not when the store is already at 95%.

Run offline compaction

The cockroach debug compact command performs manual compaction on a stopped node. It can reclaim space from obsolete SSTables that online compaction has not yet processed.

# WARNING: Offline tool. The node must be stopped first.
# Stopping the node means its replicas lose a member and the cluster
# begins rebalancing. Use only when the node cannot start due to disk
# pressure and ballast file deletion was insufficient.
cockroach debug compact /path/to/cockroach-data

The tool has minimal logging output. Monitor the store directory size from the OS during the operation to confirm progress.

Prevention

  • Alert on capacity_available below 20% with an upward trend. Below 15%, compaction loses working space and the death spiral begins. Do not wait for the paging alert at 10%.

  • Monitor MVCC garbage as a ratio. Track total bytes minus live bytes relative to total bytes. If garbage exceeds 20-30% of total store size, GC is falling behind and accumulating a disk debt.

  • Monitor protected timestamp records. Check spanconfig_kvsubscriber_protected_record_count and jobs_changefeed_protected_age_sec. Protected age above 24 hours without an active operation is a silent disk growth emergency that will not trigger any other alarm.

  • Size for 30% free space minimum. CockroachDB documentation recommends at least 20% free at all times. 30% provides headroom for compaction space amplification and Raft snapshot staging during rebalancing.

  • Watch L0 sublevels proactively. Sustained values above 5-10 mean compaction is lagging. This is the earliest warning that the LSM tree is accumulating debt.

  • Audit changefeed health regularly. A stalled changefeed is the single most insidious cause of unbounded disk growth because it silently blocks GC while the cluster appears healthy from every other angle.

How Netdata helps

  • Per-second capacity tracking on capacity_available, capacity_used, and total capacity per store reveals the exact inflection point where disk growth accelerates, not a sampled snapshot every 15 or 30 seconds.

  • L0 sublevel correlation with disk space exposes the feedback loop between compaction debt and space exhaustion. If storage_l0_sublevels is climbing while capacity_available drops, you are watching the death spiral form in real time.

  • Protected timestamp and changefeed age metrics surface the silent GC-blocker pattern. Correlating spanconfig_kvsubscriber_protected_record_count with MVCC garbage growth and disk usage trends identifies the root cause before the disk is full.

  • Write stall detection at per-second resolution catches sub-minute stall events that standard scrape intervals miss.

  • ML anomaly detection on disk usage growth rates catches gradual capacity exhaustion days or weeks before the critical threshold fires.

Netdata’s CockroachDB monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.