The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-health-err

Operations Guides

Ceph HEALTH_ERR: reading the umbrella status and finding the real fault

ceph_health_status == 2 is the most severe top-level status Ceph reports, and one of the most over-paged signals in production. The reason is structural: HEALTH_ERR is an umbrella aggregation. It tells you that one or more child health checks crossed into ERR severity. It does not tell you which subsystem failed, whether I/O is actually blocked, or whether the condition will self-resolve when PGs finish peering.

Treating the umbrella as the page to action creates two failure modes. First, it duplicates the more specific child signals (OSD_FULL, PG down, MON quorum loss) that already have their own PAGE conditions with tighter gating, producing parallel escalations for the same incident. Second, it fires on transient cold-start states. A cluster-wide OSD restart will briefly push PGs through peering and incomplete, which trips HEALTH_ERR for tens of seconds before clearing. Operators who page on HEALTH_ERR without context spend their nights chasing ghosts.

The correct posture is to treat HEALTH_ERR as a TICKET that points you at ceph health detail. The child health checks you find there carry the tighter PAGE-worthy conditions. This article walks through how to read the umbrella, how to localise the actual fault via the name label, and how to filter cold-start false positives from genuine incidents.

What this means

HEALTH_ERR is computed by the MON cluster from a set of internal health checks. Each check has a severity (HEALTH_WARN or HEALTH_ERR) and a short code (the name label). When any active check has ERR severity, the umbrella reports HEALTH_ERR. When any check has WARN severity, the umbrella reports HEALTH_WARN.

The child checks that can push the umbrella to ERR are the ones that represent data risk or unavailability: capacity exhaustion (OSD_FULL, POOL_FULL), PG availability (PG_AVAILABILITY), and module failures (MGR_MODULE_ERROR). The full set of ERR-severity codes is version-dependent, so verify the list against the Ceph version in your cluster by checking ceph health detail or the official health-checks documentation for your release.

The critical operator insight is that HEALTH_ERR and its child checks fire simultaneously. When OSD_FULL goes active, the umbrella reports ERR and ceph_health_detail{name="OSD_FULL"} reports active at the same instant. They are not staged. The child check is the diagnosis; the umbrella is the announcement.

flowchart TD
  ERR["HEALTH_ERR
ceph_health_status = 2"] ERR --> Q["MON quorum lost"] ERR --> F["OSD_FULL / POOL_FULL"] ERR --> PG["PG_AVAILABILITY
(down / incomplete)"] ERR --> MDS["MDS_DAMAGED"] ERR --> MOD["MGR_MODULE_ERROR"] Q -.-> Q1["Child PAGE:
ceph_mon_quorum_status"] F -.-> F1["Child PAGE:
OSD_FULL detail"] PG -.-> PG1["Child PAGE:
ceph_pg_down /
ceph_pg_incomplete"] MDS -.-> MDS1["TICKET - CephFS only"] MOD -.-> MOD1["TICKET - check MGR"]

Because the child checks have their own per-signal PAGE conditions with longer sustain windows and proper cold-start gating, HEALTH_ERR itself is best treated as a TICKET (sustained > 120 seconds), not a PAGE. The umbrella is redundant with its children for the genuinely catastrophic cases and noisy for the transient ones.

Common causes

CauseWhat it looks likeFirst thing to check
Cluster fullOSD_FULL active; client writes return ENOSPC; reads still succeedceph osd df tree for the fullest OSDs
PG down or incompletePG_AVAILABILITY active; specific PGs unable to serve I/Oceph pg dump_stuck inactive 60 and ceph pg dump_stuck stale 60
MON quorum lossCLI commands slow or hang; new client connections failceph mon stat and ceph quorum_status
MDS damaged (CephFS)MDS_DAMAGED active; CephFS unresponsive or read-onlyceph fs status and the MDS daemon log
MGR module failureMGR_MODULE_ERROR active; dashboard, balancer, or metrics exporter degradedceph mgr module ls and the MGR log
Cold-start peeringERR appears within minutes of OSD/MON restart; PGs in peering or incompleteWait one to two minutes; re-check before escalating

Quick checks

These are read-only and safe to run on a production cluster.

# Inspect the active health checks - this is the primary diagnostic
ceph health detail

# Same data, machine-readable, for filtering by code or severity
ceph health detail --format json-pretty | jq '.checks'

# Top-level cluster status and quorum summary
ceph status

# Per-OSD capacity - look for the fullest OSD, not the average
ceph osd df tree

# Stuck PGs by category - pass a meaningful duration, not the default
ceph pg dump_stuck inactive 60
ceph pg dump_stuck stale 60
ceph pg dump_stuck undersized 300

# MON quorum state
ceph mon stat
ceph quorum_status

# CephFS state, if applicable
ceph fs status

# Confirm whether any checks have been muted
ceph health detail --format json-pretty | jq '.mutes'

The single most important command in this list is ceph health detail. The umbrella does not localise the fault. The detail output does.

How to diagnose it

  1. Read ceph health detail and extract every active check code. Each line names a subsystem. The name label is the localisation signal.
  2. Bucket each code by category: capacity (OSD_FULL, POOL_FULL), availability (PG_AVAILABILITY), control plane (MON_*), CephFS (MDS_*, FS_*), or module (MGR_MODULE_ERROR).
  3. Check cold-start context. If a cluster-wide OSD restart, MON restart, or host reboot happened in the last few minutes, transient peering and incomplete PGs may be driving the ERR. Typical peering completes within 60 to 120 seconds; large clusters may take longer. Wait for the child signal’s sustain window (300 seconds for ceph_pg_incomplete and ceph_pg_down) before escalating.
  4. For each remaining code, drill into the corresponding subsystem signal:
    • OSD_FULL: cross-check ceph osd df tree. A single full OSD blocks writes to every PG that maps to it.
    • PG_AVAILABILITY: identify the affected PGs with ceph pg dump_stuck, then ceph pg <pgid> query for the specific blocked PG.
    • MON-related codes: check ceph quorum_status and the per-MON ceph_mon_quorum_status metric.
    • MDS_DAMAGED or MDS_ALL_DOWN: check ceph fs status and the MDS daemon log.
    • MGR_MODULE_ERROR: identify the failing module with ceph mgr module ls and the MGR log.
  5. Correlate with the child signal’s own PAGE condition. The child signal is what tells you whether I/O is actually blocked, data is actually at risk, or whether the cluster is degraded but still functional.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_health_statusThe umbrella status; 2 is ERRTransition from 1 or 0 to 2 sustained > 120 s
ceph_health_detail{name="..."}Localises the active child checkAny check with severity="HEALTH_ERR" and value 1
ceph_mon_quorum_statusMON control-plane availabilitysum() < floor(count()/2) + 1 sustained > 300 s
ceph_pg_down, ceph_pg_incompletePG-level data unavailabilityAny value > 0 sustained > 300 s
ceph_osd_up, ceph_osd_inOSD daemon and placement stateMultiple OSDs down within the same failure domain
ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytesCapacity utilisationCrossing ceph_osd_nearfull_ratio (warn) or ceph_osd_full_ratio (ERR)
ceph_osd_flag_nooutForgotten maintenance flagSet for > 24 h while OSDs are down
ceph_osd_flag_norecover, ceph_osd_flag_nobackfillRecovery intentionally disabledSet while degraded PGs exist

The umbrella is the entry point. The child signals are the alerting surface.

Fixes

Cluster full (OSD_FULL, POOL_FULL)

When the cluster crosses the full ratio (default 0.95), Ceph refuses all writes cluster-wide. When an individual OSD reaches its full ratio, writes to every PG mapped to that OSD are blocked. Reads still succeed in either case. This is a hard stop.

  • Free space: delete snapshots, expired RGW multipart uploads, or non-critical data.
  • Add capacity: new OSDs, or reweight under-used OSDs higher.
  • Force RGW garbage collection: radosgw-admin gc process.
  • Last resort (risky): temporarily raise mon_osd_full_ratio to buy time while capacity is added. This is dangerous because it removes the safety margin during the window before capacity lands.

PG down or incomplete (PG_AVAILABILITY)

These PGs cannot serve I/O. The condition does not self-resolve without intervention.

  • Identify the failed OSDs in the acting set with ceph pg <pgid> query.
  • Check whether recovery is blocked by capacity (backfill_toofull), unfound objects (recovery_unfound), or operator flags (norecover, nobackfill).
  • If objects are genuinely unfound and the acting set is permanently gone, ceph pg <pgid> mark_unfound_lost revert|delete is the last-resort escape hatch. This command causes data loss. The action argument is required. Choose revert to pick a known-good copy or delete to accept data loss.
  • For incomplete PGs after a cold start, wait for peering before any destructive action.

MON quorum loss

Without a majority quorum, the cluster cannot accept map updates. Existing clients with cached maps continue briefly, then stall.

  • Check clock synchronisation on MON hosts: chronyc tracking or ntpq -p. Clock skew above mon_clock_drift_allowed (default 0.05 s) destabilises elections.
  • Check network connectivity between MON hosts.
  • If one MON is destabilising quorum, stopping its daemon temporarily can let the others form a stable majority. This is a control-plane action that further reduces quorum headroom until the remaining MONs reform; only do this when you are sure which MON is the outlier.
  • Inspect MON store size on each MON host; a bloated store slows elections and crash recovery.

MDS damaged (CephFS only)

MDS_DAMAGED indicates metadata corruption in the MDS journal or cache. Client I/O is not data-path for RADOS, but CephFS itself may be partially or fully unavailable.

  • Read the MDS daemon log for the damage reason before running anything.

  • Recovery typically involves cephfs-journal-tool. Run cephfs-journal-tool --help to see the subcommands supported by your Ceph version before running any repair.

  • Do not run destructive repair commands without first understanding which journal segment is damaged.

MGR module failure (MGR_MODULE_ERROR)

The MGR is still running, but a Python module has failed. Monitoring, dashboard, balancer, or PG autoscaler behaviour may be impacted, but client I/O is unaffected.

  • Identify the failing module from ceph health detail and ceph mgr module ls.
  • Try disabling and re-enabling the failing module.
  • Restart the MGR daemon as a last resort; failover to the standby should be quick.

Transient cold-start ERR

If ceph health detail shows PGs in peering or incomplete immediately after a coordinated restart, wait. Typical peering completes within 60 to 120 seconds; very large clusters may take longer. Escalate only if the condition persists past the child signal’s sustain window (300 seconds for ceph_pg_incomplete and ceph_pg_down).

Prevention

  • Alert on child signals, not the umbrella. ceph_health_status == 2 is a TICKET. The PAGE conditions belong to the specific subsystem signals (ceph_pg_down, ceph_pg_incomplete, MON quorum loss, OSD_FULL).

  • Gate HEALTH_ERR with a sustain window. 120 seconds filters most cold-start transients without delaying response to real incidents.

  • Maintain capacity headroom. Stay comfortably below the backfillfull ratio (default 0.90), not near it. A cluster that loses an OSD near the backfillfull threshold has no room to recover.

  • Watch the noout trap. ceph_osd_flag_noout set for > 24 hours while OSDs are down is a common preventable cause of cascading data-loss risk.

  • Keep MON clocks synchronised. Monitor NTP/chrony on MON hosts; do not disable the clock skew check as a workaround.

  • Run deep scrubs on schedule. noscrub and nodeep-scrub are maintenance flags, not a permanent configuration.

How Netdata helps

  • The Ceph collector exposes ceph_health_status per second, so the transition into HEALTH_ERR is visible immediately and the sustain duration is precise.
  • ceph_health_detail{name="..."} is collected as a labelled metric, so each child check is its own time series. The dashboard lets you correlate the umbrella going red with the specific code that drove it.
  • MON quorum (ceph_mon_quorum_status), PG states (ceph_pg_down, ceph_pg_incomplete), capacity (ceph_cluster_total_used_raw_bytes, ceph_osd_*_ratio), and OSD flags (ceph_osd_flag_noout) are in the same view, so the diagnostic flow from umbrella to subsystem is one click.
  • Anomaly detection on the umbrella and on the child signals catches the moment of transition without requiring static thresholds for every check code.
  • Cold-start transients show up as short-lived spikes in peering PG counts alongside the umbrella blip, which makes them easy to distinguish from sustained incidents.