The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-norecover-nobackfill

Operations Guides

Ceph norecover and nobackfill: recovery intentionally, or accidentally, stopped

The symptom is familiar: PGs are stuck in degraded or undersized states, the cluster is HEALTH_WARN, but recovery throughput is zero. OSD load, network, and capacity all look normal. The cause is often two cluster-wide flags sitting in the OSD map: norecover and nobackfill. These flags are legitimate tools for protecting client I/O during recovery storms or maintenance, and a common source of “forgotten flag” incidents alongside noout.

The risk is operational, not technical. Nothing in ceph -s explicitly announces “recovery is paused by administrative action.” A degraded cluster with these flags set looks very similar to a degraded cluster where recovery is stalled for some other reason. The distinguishing signal is that the recovery rate is exactly zero while the flags are present.

What this means

norecover and nobackfill are cluster-wide OSD map flags that halt recovery machinery.

  • norecover blocks log-based recovery. When an OSD returns after a brief outage, Ceph normally uses the PG log to identify and replay only changed objects. With norecover set, this does not happen.
  • nobackfill blocks full PG content migration. Backfill copies entire PG contents rather than just changed objects, and is used when an OSD has been missing too long for the PG log to cover the gap, or when OSDs are added or removed.

Both flags are set and unset with ceph osd set and ceph osd unset. They take effect cluster-wide. In-progress recovery pauses when the flag is set and resumes from where it left off when the flag is cleared.

The companion flag norebalance is narrower. The official flag definition is “block osd backfill unless pg is degraded”: it prevents backfill of remapped-but-healthy PGs and does NOT block recovery of genuinely degraded PGs. So unsetting norecover+nobackfill alone is enough for degraded PGs to recover even with norebalance set. It is still worth checking all three flags together, because a PG can be both remapped and degraded, and the precise interaction depends on which replicas are missing.

These flags exist for good reasons. The two most common legitimate uses are recovery storm protection and maintenance windows. Multiple OSD failures can trigger massive data movement that competes with client I/O; setting norecover and nobackfill temporarily pauses that movement while you throttle or investigate. During planned OSD work, the flags prevent the cluster from overreacting to temporary state changes.

The health check OSDMAP_FLAGS fires when any of these flags are set. Treat that warning as a pointer, not noise.

Common causes

CauseWhat it looks likeFirst thing to check
Forgotten maintenance flagHEALTH_WARN, degraded PGs, zero recovery throughput, no recent OSD eventsceph osd dump | grep -E "norecover|nobackfill|norebalance"
Recovery storm throttle left onFlags set during a prior incident, never cleared; degraded count flat over hoursIncident records and current flag state
Script or automation set themFlags appear after a deploy or runbook step; no operator recall of setting themChange history and automation logs
Intentional pause still in effectMaintenance window still open; flagged in runbook with unset timeMaintenance ticket status

Quick checks

# Show current OSD map flags - authoritative check
ceph osd dump | grep -E "norecover|nobackfill|norebalance|noout"

# Cluster summary - look for degraded/recovering PG counts
ceph -s

# Health detail - OSDMAP_FLAGS fires when these flags are set
ceph health detail | grep -E "OSDMAP_FLAGS|norecover|nobackfill|norebalance"

# PG state summary - degraded count with zero recovery is the signature
ceph pg stat

# Stuck PGs by category - confirms degraded PGs are not progressing
ceph pg dump_stuck degraded
ceph pg dump_stuck undersized

# Per-pool recovery rate - confirms recovery is truly stalled
ceph osd pool stats

# OSD capacity state - rules out backfill_toofull as the actual cause
ceph osd df

How to diagnose it

  1. Confirm the flag is set. Parsing ceph osd dump for the flag names is more reliable than reading ceph -s. Any of norecover, nobackfill, or norebalance appearing here is a candidate cause.

  2. Verify recovery is actually stalled. Cross-check that degraded or undersized PGs exist (ceph pg stat) and that the recovery rate is zero or near-zero (ceph osd pool stats). If PGs are healthy, the flag is harmless. If PGs are degraded but recovery is progressing, the flag is not your bottleneck.

  3. Rule out other stall causes. Recovery can also be blocked by backfill_toofull (target OSDs too full), recovery_unfound or backfill_unfound (objects cannot be located), or downstream issues like network partitions. Check ceph health detail for these specific PG states.

flowchart TD
  A[Degraded PGs, recovery rate near zero] --> B{norecover or nobackfill set?}
  B -- Yes --> C{Maintenance window open?}
  C -- Yes --> D[Track ticket, unset after window]
  C -- No --> E[Throttle, then unset flags]
  B -- No --> F{backfill_toofull or unfound?}
  F -- Yes --> G[Capacity or unfound root cause]
  F -- No --> H[Check norebalance, network, OSD load]
  1. Check norebalance alongside the others. If norebalance is also set and you intend full recovery to resume, include it in the unset.

  2. Quantify exposure. Count degraded PGs and estimate the redundancy loss. While the flag is set, every degraded PG is one failure away from data loss. Flag-state plus degraded PGs is a ticket-worthy condition regardless of intent.

Metrics and signals to monitor

The metric names in this table are Netdata collector metric names (out of scope for this factual review); the underlying Ceph signals they reflect (OSD map flags, per-pool recovery byte rate, degraded and undersized PG counts, degraded object count) exist in the Ceph mgr/PG stats.

SignalWhy it mattersWarning sign
ceph_osd_flag_norecoverCluster-wide flag that halts log-based recoveryValue of 1 while ceph_pg_degraded > 0
ceph_osd_flag_nobackfillCluster-wide flag that halts full PG migrationValue of 1 while ceph_pg_degraded > 0
ceph_osd_flag_norebalanceHalts rebalance of remapped PGsSet alongside the other two, blocking some recovery
ceph_pool_recovering_bytes_per_secMeasures whether recovery is making progressZero or near-zero with degraded PGs present
ceph_pg_degradedFewer replicas than pool sizeNon-zero sustained for more than 300 seconds
ceph_pg_undersizedFewer copies than min_sizeNon-zero sustained for more than 300 seconds
ceph_health_detail{name="OSDMAP_FLAGS"}Surfaces any of the flags being setActive at all when flags are present
ceph_num_objects_degradedCluster-wide degraded object countNot decreasing while flag is set

Fixes

If the flag is accidentally set

Before unsetting, especially if the flag has been set for a long time, consider throttling recovery to avoid releasing a wave of accumulated backfill pressure. Operators have reported OSDs crashing immediately when nobackfill is unset after a long pause, because every pending backfill rushes to start at once.

# Throttle recovery before unblocking - safer for clusters with accumulated backlog
ceph config set osd osd_recovery_max_active 1
ceph config set osd osd_recovery_op_priority 1

# Unset the flags
ceph osd unset norecover
ceph osd unset nobackfill

# If norebalance was also set, unset it too
ceph osd unset norebalance

After unsetting, watch the recovery rate climb and the degraded PG count fall. Once recovery completes and the cluster returns to HEALTH_OK, return the throttle values to their previous settings.

If the flag is intentionally set

Document why. Record the expected unset time in your runbook or change management system. The alert for norecover or nobackfill set while degraded PGs exist should fire as a reminder, not be silenced permanently.

If the flag was set to weather a recovery storm and the storm has passed, unset it. The recovery storm pattern, where multiple OSD failures cascade into client-visible latency, is handled in more depth in Ceph capacity death spiral.

If OSDs crash on unset

This is a sign that accumulated backfill pressure is overwhelming target OSDs. Re-set the flag immediately, throttle more aggressively, and consider reweighting the fullest OSDs down to spread load across more targets.

# Re-set if OSDs crash on unset
ceph osd set nobackfill

# Throttle harder
ceph config set osd osd_recovery_max_active 1
ceph config set osd osd_recovery_op_priority 1

# Reweight the fullest OSDs down to spread backfill targets
# NOTE: osd id is the integer ID, not osd.<id>
ceph osd reweight <osdid> 0.9

There is no documented universal value for this response; reweighting the fullest OSDs down (0.9 here) plus a single recovery op at a time is field practice, and current Ceph versions also expose osd_max_backfills and per-pool target_max_objects as alternative backfill levers. Validate the exact numbers against your cluster’s capacity headroom.

Prevention

  • Treat OSD map flags as state, not commands. Every ceph osd set should have an accompanying ticket with an expected unset time. This is the same discipline that prevents the noout trap.
  • Alert on flag state correlated with PG state. The condition to alert on is not “flag is set” but “flag is set while degraded PGs exist.” That combination is always actionable.
  • Include flag checks in your runbook verification step. After any maintenance, the closing checklist should verify ceph osd dump | grep -E "norecover|nobackfill|norebalance|noout" returns only expected flags.
  • Monitor norebalance too. It is less obvious than norecover and nobackfill but can silently block recovery of remapped PGs.
  • Track flag dwell time. A flag set for more than 24 hours with no open maintenance ticket is almost certainly forgotten.

How Netdata helps

Netdata’s Ceph collector surfaces OSD map flags as ceph_osd_flag_{name} metrics, so you can alert on norecover and nobackfill state without parsing CLI output.

  • The correlation that matters most is ceph_osd_flag_norecover == 1 OR ceph_osd_flag_nobackfill == 1 combined with sum(ceph_pg_degraded) > 0. That composite is ticket-worthy regardless of whether the flag is intentional.
  • Pair flag state with ceph_pool_recovering_bytes_per_sec. If recovery rate is zero while the flag is set and degraded PGs exist, recovery is paused by administrative action rather than by a structural failure.
  • The ceph_health_detail{name="OSDMAP_FLAGS"} check surfaces as a labeled metric, giving you a single signal for “any interesting OSD map flag is currently set.”
  • ML anomaly detection on recovery rate helps distinguish a flag-paused flatline from a structurally stalled recovery. Both produce near-zero throughput, but the surrounding context differs.
  • Per-second collection means you see the exact moment a flag is set or unset, useful for correlating with change events and incident timelines.