The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-health-warn

Operations Guides

Ceph HEALTH_WARN: which warnings are noise and which are structural

HEALTH_WARN is an umbrella status, not a single condition. It fires during normal recovery after an OSD restart, and it also fires when the cluster is one step from data loss. Both look identical if you only watch the top-level status. The common mistake is blanket-silencing WARN because it fires too often during recovery, then missing the structural warnings that signal real danger.

The fix is to stop treating HEALTH_WARN as one thing. The ceph_health_detail metric exposes each individual health check with a name label that identifies the specific condition. That name is how you split transient recovery warnings from structural problems.

This reference covers which health check codes are expected during normal operations, which require immediate investigation, and how to mute individual codes without blinding yourself to the structural ones.

What HEALTH_WARN actually means

The monitors compute cluster health from a set of health checks. It surfaces as ceph_health_status: 0 for HEALTH_OK, 1 for HEALTH_WARN, 2 for HEALTH_ERR. WARN means the cluster is still serving data but a condition requires attention. ERR means I/O may be blocked or data is at risk.

The important detail is that HEALTH_WARN is an aggregation. Behind it is a list of specific health check codes, each with its own severity, message, and affected components. The Prometheus MGR module exposes this as ceph_health_detail with labels name and severity only (the message text appears only in ceph health detail JSON output). The value is 1 when the check is active and 0 when inactive.

The name label is the key. It identifies which specific health check is firing. OSD_NEARFULL, MON_CLOCK_SKEW, PG_DEGRADED, SLOW_OPS, and OSD_SCRUB_ERRORS are all distinct names with distinct meanings. Some are transient and expected. Some are structural and dangerous. The top-level WARN status does not tell you which is which.

The split: transient vs structural

The operational question is simple: will this WARN clear on its own, or does it require human action? The table below maps common health check names to their category.

Health check nameCategoryWhy
PG_DEGRADED, PG_RECOVERY, PG_BACKFILL, PG_PEERINGTransientExpected during and after OSD failures or CRUSH changes. Self-resolves as recovery completes.
OBJECT_MISPLACEDTransientObjects on non-optimal OSDs during rebalancing. Data is safe, just not ideally placed.
OSD_DOWN (with noout set, planned)TransientOSD is down for maintenance. Expected as long as noout is intentional and time-limited.
OSD_NEARFULLStructuralCapacity at or above 85%. Recovery requires spare space. Losing more OSDs pushes past backfillfull and blocks healing.
OSD_SCRUB_ERRORS or inconsistent PGsStructuralData corruption detected between replicas. If the good copy fails before repair, data is lost.
PG_FAILED_REPAIRStructuralAutomatic repair attempted and failed. Needs manual intervention to identify the authoritative copy.
MON_CLOCK_SKEW (persistent)StructuralClock drift beyond mon_clock_drift_allowed (default 0.05s). Can cause repeated Paxos elections and quorum instability.
SLOW_OPSStructuralOperations stuck beyond osd_op_complaint_time (default 30s). Something is blocking I/O.
OSD flapping pattern (no native health check)StructuralOSD cycling up and down. Each flap generates OSD map epochs and cascades peering load across the cluster.
noout flag set more than 24h (no native health check; monitor ceph_osd_flag_noout externally)StructuralThe noout trap. Forgotten flag prevents recovery from OSD failures. Degraded PGs accumulate silently.
LARGE_OMAP_OBJECTSStructuralRGW bucket index shards exceeded OMAP threshold. OSD hosting the index will degrade.
PG_NOT_DEEP_SCRUBBED (sustained)StructuralVerification debt. Corruption may be accumulating undetected.
POOL_NO_REDUNDANCY or TOO_FEW_OSDSStructuralCRUSH rules cannot place replicas across enough failure domains. Configuration error, not a runtime fault.

The dividing line is not the severity field. It is whether the condition self-resolves. Degraded PGs during active recovery are noise. Degraded PGs with zero recovery rate are structural. The same health check name can be either, depending on context.

flowchart TD
    A["HEALTH_WARN fires"] --> B["Read ceph_health_detail
by name label"] B --> C{"Check name and context"} C -->|"Degraded PGs +
recovery active"| D["Transient
Monitor recovery rate"] C -->|"Degraded PGs +
recovery stalled"| E["Structural
Recovery blocked"] C -->|"OSD_NEARFULL"| F["Structural
Capacity incident"] C -->|"OSD_SCRUB_ERRORS"| G["Structural
Data corruption"] C -->|"MON_CLOCK_SKEW
persistent"| H["Structural
Fix NTP or risk quorum"] C -->|"OSD_DOWN +
noout > 24h"| I["Structural
Noout trap"] C -->|"OSD flapping pattern"| J["Structural
Cascade risk"]

Transient warnings: monitor, do not mute

These conditions fire during normal cluster operations. They are expected after OSD restarts, during recovery, during rebalancing, and during planned maintenance. They should clear without intervention.

Degraded and recovering PGs during recovery. When an OSD fails or is added, PGs transition through degraded, recovering, and backfilling states. This is the cluster healing itself. The signal that matters is the recovery rate. If ceph_pool_recovering_bytes_per_sec is nonzero and the degraded count is decreasing, recovery is progressing normally.

Misplaced objects during rebalancing. ceph_num_objects_misplaced rises when CRUSH weights change or OSDs are added. Objects are on non-optimal OSDs but data is fully replicated. This clears as backfill completes.

Peering PGs after topology changes. PGs briefly enter peering after OSD state changes. This is normal. Peering that transitions to active+clean within minutes is expected.

OSD_DOWN with noout set for planned maintenance. If ceph_osd_flag_noout is 1 and the OSD down is intentional, the WARN is expected. The risk is forgetting to unset noout after maintenance. That moves it into the structural category.

Do not mute these codes globally. They are your signal that recovery is happening. The right approach is to alert on the combination of degraded PGs with stalled recovery, not on degraded PGs alone.

Structural warnings: investigate now

These conditions indicate genuine risk to data availability, integrity, or cluster stability. They do not self-resolve. Each one has a specific failure mode behind it.

OSD_NEARFULL. The cluster or individual OSDs have crossed the nearfull ratio (default 0.85). Ceph throttles backfill at this point. If another OSD fails, the cluster may not have enough spare capacity to recover, pushing past backfillfull (0.90) where recovery is blocked entirely. Check ceph osd df tree for per-OSD variance, not just the cluster average. The most-full OSD hits the wall first.

OSD_SCRUB_ERRORS and inconsistent PGs. Deep scrub found that replicas disagree on content, size, or checksums. This is silent data corruption from bit rot, firmware bugs, memory errors, or software bugs. The check ceph_health_detail{name="OSD_SCRUB_ERRORS"} is active. The per-PG inconsistency count is exposed per pool as ceph_pg_inconsistent and is visible via ceph pg dump | grep inconsistent. Do not blindly run ceph pg repair until you have identified which replica is corrupt. If the good replica fails before repair, the data is lost.

PG_FAILED_REPAIR. This is the escalation of inconsistent. Ceph tried to repair and could not. Manual intervention is required to identify the authoritative copy, force repair, or accept data loss. Work with support before forcing anything: ceph pg repair against the wrong replica propagates corruption.

MON_CLOCK_SKEW that persists. Clock drift between monitors beyond mon_clock_drift_allowed (default 0.05s) triggers this check. Transient skew after a MON reboot is normal and clears within minutes if NTP or chrony is healthy. Persistent skew is structural: it causes repeated Paxos elections, map distribution delays, and eventual quorum instability. If ceph_health_detail{name="MON_CLOCK_SKEW"} stays active for more than a few minutes, check NTP or chrony on the MON hosts.

SLOW_OPS. ceph_healthcheck_slow_ops counts operations stuck beyond osd_op_complaint_time (default 30s). These are not slow operations, they are stuck operations. Something is blocking I/O: disk failure, lock contention, peering delays, network partition, or resource exhaustion. Cluster-wide slow ops across many OSDs is far more concerning than isolated slow ops on one OSD.

OSD flapping (no native health check). An OSD cycling between up and down states. Each transition generates a new OSD map epoch that all OSDs must process, and triggers peering across every PG on that OSD. Flapping cascades: the peering overhead can slow healthy OSDs, causing them to miss heartbeats too. Set noout on the specific OSD to stop the peering storm, then investigate the root cause.

noout set for more than 24 hours. ceph_osd_flag_noout equal to 1 for an extended period is the single most common preventable Ceph outage. The operator set noout for maintenance, forgot to unset it, and the cluster stopped recovering from subsequent failures. Degraded PGs accumulate silently. The next failure may cause data loss.

LARGE_OMAP_OBJECTS. RGW bucket index shards have accumulated too many OMAP keys. This is the early stage of an OMAP storm. The OSD hosting the oversized index will degrade as RocksDB operations slow. Identify affected buckets and trigger resharding.

PG_NOT_DEEP_SCRUBBED sustained beyond the policy window. If deep scrubs are not completing within the configured interval, the cluster is accumulating verification debt. Corruption may be going undetected. Check whether noscrub or nodeep-scrub flags are set and forgotten.

POOL_NO_REDUNDANCY or TOO_FEW_OSDS. CRUSH rules cannot place replicas across enough failure domains. This is a configuration error, not a runtime fault, but it means data is not distributed as expected for failure tolerance.

How to read ceph_health_detail the right way

The name label on ceph_health_detail is how you split the two categories programmatically. Do not alert on the umbrella ceph_health_status alone. Alert on specific health check names, with context.

For inspecting active checks:

# Check which health checks are currently active
ceph health detail

# List active check names in machine-readable form
ceph health detail -f json-pretty | jq '.checks | keys'

The Prometheus metric ceph_health_detail{name="OSD_NEARFULL"} returns 1 when that check is active. Alert on specific names rather than on the aggregated status. For transient checks, add context. Alert on degraded PGs only when recovery rate is also near zero, which distinguishes stalled recovery from normal healing.

For muting, use per-code muting, never blanket suppression:

# Mute a specific check for a limited duration. Note: there is no
# `ceph health mute list` subcommand; muting always requires a code.
ceph health mute POOL_APP_NOT_ENABLED 1h

# Sticky mute persists for the full duration even if the condition clears
# (the --sticky flag is confirmed across Reef and Squid)
ceph health mute POOL_APP_NOT_ENABLED 24h --sticky

Muting is per-code. You cannot mute the entire HEALTH_WARN status, and you should not try to. The right pattern is to mute specific codes that are known-noise for your deployment, while leaving structural codes unmuted and alerting on them.

The checks that look like noise but are structural

Some health checks are easy to dismiss because they seem benign or because they fire during routine operations. These are the ones that cause outages when ignored.

noout flag. It looks like a maintenance artifact. It is the noout trap. A flag set for a one-hour maintenance window that was never cleared will silently prevent recovery for days or weeks. ceph_osd_flag_noout should alert if set for more than 24 hours without an accompanying maintenance ticket.

Clock skew. It looks like an NTP hiccup. Persistent MON_CLOCK_SKEW is a precursor to quorum loss. The default threshold is tight at 0.05s. VMs are particularly prone to drift. Do not disable the check. Fix NTP or chrony.

Scrub errors. They look like a one-time checksum mismatch. They are silent data corruption. If the consistent replica fails before repair, the corrupt copy becomes the only copy. Investigate promptly and identify which replica is wrong before running repair.

Nearfull. It looks like a capacity heads-up. It is the threshold past which recovery may be blocked. At 85% utilization, losing OSDs pushes the cluster toward backfillfull, where healing stops entirely. Treat nearfull as a capacity incident, not an advisory.

How Netdata helps

  • Per-check health detail metrics. Netdata collects ceph_health_detail with the name label for each health check code. This lets you alert on specific structural checks like OSD_NEARFULL, OSD_SCRUB_ERRORS, and MON_CLOCK_SKEW without firing on transient recovery checks.
  • Recovery rate correlation. Netdata surfaces recovery throughput alongside degraded PG counts on the same dashboards. Alerting on degraded PGs with near-zero recovery rate distinguishes stalled recovery from normal healing.
  • Flag state monitoring. Cluster-wide flags like noout and norecover are collected as metrics. This catches the noout trap before it causes an outage.
  • Per-OSD capacity and latency. Nearfull warnings on individual OSDs are visible alongside per-OSD commit and apply latency, so you can see whether a nearfull OSD is also a slow OSD.
  • Clock skew via health checks. ceph_health_detail{name="MON_CLOCK_SKEW"} is collected directly, so you do not need a separate NTP monitoring path to catch persistent skew.