The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-capacity-death-spiral

Operations Guides

Ceph capacity death spiral: an OSD fails and recovery has nowhere to go

An OSD fails on a cluster you have been running at 80% or higher. ceph -s shows recovery starting, then ceph health detail lists OSD_NEARFULL and BACKFILL_TOOFULL. The degraded PG count climbs, then flatlines. Recovery bytes per second sits near zero. Nothing is healing, and you are one more failure away from data loss.

This is the Ceph capacity death spiral. It is not a single fault: it is the intersection of tight capacity, CRUSH imbalance, and the hard thresholds Ceph uses to protect itself. Recovery needs spare space on the surviving OSDs. When those OSDs are already past the backfillfull ratio (default 0.90), Ceph refuses to push more data onto them, and backfill stalls for every PG that would target them.

The signature is BACKFILL_TOOFULL (or the health check PG_BACKFILL_FULL) alongside OSD_NEARFULL, a flat or rising degraded PG count, and recovery rate near zero. That combination means the cluster has lost its ability to self-heal until you add space, free space, or move the thresholds. This article walks through confirming the pattern and the tradeoffs of each response.

What this means

Ceph uses three capacity thresholds on each OSD, set as ratios of utilized space:

ThresholdDefault ratioWhat happens when crossed
nearfull0.85HEALTH_WARN. Backfill may be throttled.
backfillfull0.90New backfill operations to this OSD are refused. PGs targeting it enter backfill_toofull.
full0.95HEALTH_ERR. All writes to this OSD stop.

The expected ordering is nearfull < backfillfull < full. Violating that ordering raises OSD_OUT_OF_ORDER_FULL.

Recovery after an OSD failure is not free. The data from the failed OSD has to land on the surviving OSDs in the same pool and failure domain. If those OSDs are already at or above the backfillfull ratio, Ceph will not start new backfill to them. The PGs that need that target OSD sit in backfill_wait with the backfill_toofull flag set.

flowchart TD
    A[OSD fails] --> B[Recovery targets surviving OSDs]
    B --> C{Target OSDs above backfillfull?}
    C -- No --> D[Backfill proceeds, PGs heal]
    C -- Yes --> E[PGs enter backfill_toofull]
    E --> F[Degraded PG count flat or rising]
    F --> G[Recovery rate near zero]
    G --> H[Cluster stuck at reduced redundancy]
    H --> I[Second failure risks data loss]

The danger is the exposure window. With degraded PGs not healing, the cluster runs with fewer replicas than the pool size specifies. For a 3x replicated pool, some objects may be down to a single good copy. If the OSD holding that last copy fails before recovery unblocks, those objects become unfound, which is potential data loss.

A related trap: when OSDs are marked OUT, their capacity is subtracted from the cluster total. A cluster that looked comfortable at 78% can jump past nearfull once several OUT OSDs leave the denominator. Always check ceph osd df tree for the real picture, including which OSDs are IN versus OUT.

Common causes

CauseWhat it looks likeFirst thing to check
Running hot (>70% average)Cluster average near 80%, single OSDs already at 85%+ceph osd df tree: compare average to max OSD utilization
Uneven CRUSH distributionA few OSDs at 90%+ while average is 75%ceph osd df tree sorted by utilization variance
Sudden data ingestionPool usage spiked over hours or days before the OSD failureceph df detail and pool growth rate
OUT OSDs reduce total capacityNearfull appeared right after OSDs went OUTceph osd dump for in/out state, recompute real headroom
Forgotten recovery flagsnorecover or nobackfill in flags, recovery rate zeroceph osd dump | grep flags

Quick checks

Read-only and safe to run during an incident:

# Cluster status and recovery line
ceph -s

# Health detail including BACKFILL_TOOFULL and OSD_NEARFULL
ceph health detail

# Per-OSD utilization, variance, and which are past thresholds
ceph osd df tree

# Stuck PGs - look for backfill_toofull in the state column
ceph pg dump_stuck unclean

# Pool-level usage and max_avail
ceph df detail

# Check for recovery flags that would stall healing
ceph osd dump | grep flags

# Current PG state summary and recovery activity
ceph pg stat

If ceph health detail shows PG_BACKFILL_FULL or lists PGs in backfill_toofull, and ceph osd df tree shows surviving OSDs at or above 90% utilized, the pattern is confirmed.

How to diagnose it

  1. Confirm the thresholds are firing. Run ceph health detail and look for OSD_NEARFULL, OSD_BACKFILLFULL, and PG_BACKFILL_FULL. Note which OSDs are named in each check.

  2. Check per-OSD utilization, not just the average. Run ceph osd df tree. The cluster average can be 80% while three OSDs sit at 92%. Those three OSDs are the ones refusing backfill. Watch the gap between average and maximum utilization.

  3. Confirm recovery is stalled, not just slow. Sample ceph pg stat repeatedly and watch the recovering/backfilling PG count and recovery rate. Then run ceph pg dump_stuck unclean. If the degraded PG count is flat across samples and recovery bytes per second is near zero, recovery is stalled.

  4. Identify which PGs are blocked and why. Look for backfill_toofull or recovery_toofull in the PG states. For a specific stuck PG, run ceph pg <pgid> query to see which target OSD is too full to accept data.

  5. Check for administrative flags. Run ceph osd dump | grep flags. If norecover or nobackfill is set, recovery is intentionally stopped. If noout is set alongside down OSDs, they remain down+in and degraded PGs persist. Clear forgotten flags before assuming capacity is the only problem.

  6. Recompute real headroom. OSDs marked OUT have their capacity subtracted from the total. The cluster’s effective utilization may be higher than ceph df suggests if OSDs were marked OUT recently.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_pg_backfill_toofullPGs blocked because target OSD is past backfillfullAny nonzero value sustained, especially with degraded PGs
ceph_pg_recovery_toofullPGs blocked because target OSD is past the full ratioAny nonzero value means the cluster is at the hard stop
ceph_num_objects_degradedObjects with fewer replicas than pool sizeFlat or increasing trend instead of decreasing
ceph_pool_recovering_bytes_per_secWhether recovery is making progressNear zero while degraded count is high
ceph_health_detail{name="OSD_NEARFULL"}OSDs approaching capacity limitsActive, especially with individual OSDs above 85%
ceph_health_detail{name="OSD_BACKFILLFULL"}OSDs refusing new backfillActive means recovery will stall for PGs targeting these OSDs
Per-OSD utilizationDistribution imbalanceMax OSD utilization more than 10% above the average
ceph_osd_flag_norecover / ceph_osd_flag_nobackfillForgotten maintenance flagsSet while degraded PGs exist

Fixes

There is no fix that does not involve adding capacity, freeing capacity, or moving thresholds. Each has tradeoffs.

Reweight disproportionately full OSDs

If a few OSDs are at 92% while the rest sit at 80%, lowering the CRUSH weight of the fullest OSDs forces data to move elsewhere:

# Lower the reweight on a specific full OSD (0.0 to 1.0 scale)
ceph osd reweight <osd-id> 0.9

This only helps if the other OSDs have headroom. If the cluster average is already past backfillfull, reweighting just moves the problem. ceph osd reweight-by-utilization flattens distribution automatically, but watch the recovery impact on clients.

Free space

Delete non-critical data, old snapshots, or run RGW garbage collection. Snapshots consume space invisibly. Check ceph df detail for snapshot consumption by pool. For RGW deployments, forcing GC can reclaim space from deleted-but-not-collected objects.

This is the safest fix, but it depends on having data you can actually remove.

Add capacity

Add new OSDs. This is the only fix that increases the denominator, and it is the slowest if you do not have spare hardware. New OSDs receive backfill and reduce pressure on existing OSDs. Plan for the time it takes to provision, weight, and let the cluster rebalance onto them.

Temporarily raise the backfillfull ratio (dangerous, last resort)

If recovery is completely blocked and you cannot add or free space quickly, you can temporarily raise the backfillfull ratio to let Ceph push data onto OSDs it is currently refusing. This is risky. It lets OSDs fill closer to the full ratio, leaving less margin before writes stop entirely. Use this only as a bridge while you add capacity, and lower the ratio again as soon as recovery progresses.

# Temporarily raise the backfillfull ratio on a running cluster
ceph osd set-backfillfull-ratio 0.95

Treat this as buying time, not as fixing the problem. If an OSD hits the full ratio (0.95) while you are in this state, all writes to that OSD stop and the cluster enters HEALTH_ERR.

Prevention

The death spiral is preventable with capacity discipline and per-OSD monitoring.

  • Keep average utilization below 70%. Stay at least 20% below the backfillfull ratio (0.90) to leave room for recovery after a failure. That puts the operational ceiling around 70% average. Small clusters with few failure domains need even more headroom because losing one host removes a larger fraction of total capacity.
  • Monitor per-OSD utilization, not just cluster average. CRUSH does not guarantee even distribution. A cluster at 70% average with several OSDs at 85% is already at risk. Alert on maximum OSD utilization and on the variance between the fullest and average OSD.
  • Model failure scenarios explicitly. Ask whether the surviving OSDs can receive recovery data if the largest host fails. At 75% average with a host holding 10% of capacity, the remaining OSDs would need to absorb that data at roughly 83%, already past nearfull.
  • Do not dismiss nearfull warnings. Nearfull at 85% is the signal that a single failure will push surviving OSDs past backfillfull. Treat it as a capacity planning trigger, not background noise.
  • Watch recovery rate during any OSD failure. Degraded PGs with active recovery are expected. Degraded PGs with zero recovery rate mean the cluster cannot heal. That is the early signal of the death spiral, often before backfill_toofull even appears.
  • Track OSD IN/OUT state against total capacity. OSDs marked OUT reduce the cluster total. A cluster can tip into nearfull simply from OUT OSDs leaving the denominator, even without new writes.

How Netdata helps

Netdata’s Ceph collector surfaces the per-second signals that distinguish a normal recovery from a stalled one:

  • Per-OSD utilization shows CRUSH imbalance directly, not just the cluster average. The OSD at 92% while the average is 78% is the one that will block recovery.
  • ceph_pg_backfill_toofull and ceph_pg_recovery_toofull as gauges mark the exact moment recovery loses its target space. Correlate with degraded PG count to confirm the stall.
  • Degraded object count alongside recovery rate reveals whether healing is progressing. A flat degraded count with near-zero ceph_pool_recovering_bytes_per_sec is the death spiral signature.
  • ceph_health_detail labels (OSD_NEARFULL, OSD_BACKFILLFULL, PG_BACKFILL_FULL) map health checks to specific capacity conditions without parsing CLI output.
  • Recovery flag state (norecover, nobackfill, noout) on the same dashboard as PG states prevents a forgotten flag from consuming an hour of diagnosis.

Correlating per-OSD capacity with PG recovery state and recovery rate in one view shortens the path from “OSD failed” to knowing whether the cluster can still heal.