The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-osd-nearfull

Operations Guides

Ceph nearfull: the 85% warning that decides whether the cluster can heal

OSD_NEARFULL or POOL_NEAR_FULL in ceph health detail means HEALTH_WARN. Clients are still reading and writing, and nothing looks broken yet. The nearfull ratio (default 0.85) is not a polite reminder to plan storage. It is the point where Ceph warns it is running out of the spare space it needs to heal itself.

One OSD failure at 85% can push surviving OSDs past backfillfull (0.90), at which point backfills refuse to start and recovery stalls. If another OSD fails in that window, you are left with degraded PGs that cannot be recovered, and the next stop is OSD_FULL at 95% where all client writes return ENOSPC.

What this means

Ceph has three capacity thresholds, all stored in the OSDMap after cluster creation:

ThresholdDefaultEffect
nearfull_ratio0.85HEALTH_WARN, backfill may be throttled
backfillfull_ratio0.90Target OSDs refuse backfill data, recovery blocks
full_ratio0.95All writes stop, clients get ENOSPC

Ceph enforces nearfull < backfillfull < full, and the OSD_OUT_OF_ORDER_FULL health check fires if you set them in the wrong order.

The non-obvious part: backfillfull does not produce a HEALTH_ERR. It produces backfill_toofull PG states and silently stalls recovery. A cluster sitting at 88% looks like it is functioning, but if an OSD fails, the PGs assigned to it cannot be redistributed because no surviving OSD will accept the backfill. The warning you ignored becomes the recovery that never starts.

flowchart TD
    A["Cluster at 85%+ (nearfull)"] --> B["OSD fails"]
    B --> C["PGs need to move to surviving OSDs"]
    C --> D{"Surviving OSDs under 90%?"}
    D -->|"Yes"| E["Backfill proceeds, cluster self-heals"]
    D -->|"No"| F["backfill_toofull blocks recovery"]
    F --> G["Degraded PGs accumulate"]
    G --> H["Second failure risks data loss"]

This cascade is why the headroom rule exists. To survive losing one failure domain (host, rack) worth of OSDs and still have recovery proceed, you need enough spare capacity that the redistributed data does not push surviving OSDs past backfillfull. The operator rule of thumb: stay at least 20% below backfillfull. With the default backfillfull at 0.90, that means keeping raw utilization below roughly 70-72%.

Common causes

CauseWhat it looks likeFirst thing to check
Organic data growthAll OSDs filling evenly, nearfull rising slowly over weeksceph df trend, daily growth rate
Unbalanced CRUSH distributionA few OSDs hit nearfull while cluster average is moderateceph osd df tree, max vs mean utilization
OSD loss reducing effective capacityNearfull appears after OSDs marked OUTceph osd tree, count of OUT OSDs
Snapshot accumulationPool grows but live data size does notceph df detail, snapshot counts per pool
RGW GC backlogRGW pools growing after bulk deletesradosgw-admin gc list --include-all
Reweight driftOne OSD oversized from past reweight imbalanceceph osd df, CRUSH weights

Quick checks

These are read-only and safe to run during production.

# Which OSDs and pools triggered nearfull
ceph health detail | grep -iE 'nearfull|near_full'

# Per-OSD utilization with failure domain hierarchy
ceph osd df tree

# Actual configured thresholds from the OSDMap (not ceph.conf)
ceph osd dump | grep -E 'full_ratio|nearfull_ratio|backfillfull_ratio'

# Pool-level usage and max_avail
ceph df detail

# PGs already blocked by capacity
ceph pg dump_stuck unclean | grep -iE 'toofull'

# Failure domains and any down/out OSDs reducing capacity
ceph osd tree

# Recovery flags that may be silently stopping healing
ceph osd dump | grep -E 'norecover|nobackfill|noout'

How to diagnose it

  1. Confirm the real thresholds. The defaults are 0.85 / 0.90 / 0.95, but your cluster may have been tuned. Run ceph osd dump | grep -E 'full_ratio|nearfull_ratio|backfillfull_ratio'. These values live in the OSDMap, not ceph.conf, so editing the config file has no effect on a running cluster. Use ceph osd set-nearfull-ratio, ceph osd set-backfillfull-ratio, and ceph osd set-full-ratio to change them.

  2. Check whether recovery is already blocked. Look for backfill_toofull or recovery_toofull PG states. If you see them alongside OSD_NEARFULL, the trap has already sprung: the cluster cannot heal because target OSDs are too full. This is more urgent than the nearfull warning itself.

  3. Find the fullest OSDs, not the cluster average. ceph osd df tree shows per-OSD utilization. The OSD that triggers nearfull is the most-full one, not the average. A cluster at 70% average with one OSD at 86% is already in nearfull. CRUSH does not guarantee even distribution, and variance of 10-20% between OSDs is common.

  4. Estimate runway. Track daily growth rate from ceph df samples over time: (total_bytes - used_raw_bytes) / daily_growth_rate. Use the worst-case growth, not the average. Account for replication: a 3x replicated pool uses 3x raw space per usable byte.

  5. Model losing your largest failure domain. If you lose the host with the most OSDs, can the remaining OSDs absorb that data without crossing backfillfull? This is the question the nearfull warning is actually asking. If the answer is no, you are already in the danger zone regardless of the current cluster-wide percentage.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_osd_nearfull_ratio / ceph_osd_full_ratioActual configured thresholds from OSDMapAlerts must use these, not hardcoded 0.85 / 0.95
ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytesRaw cluster utilizationTrending toward 0.70 with OSD failure risk
Per-OSD utilization (max, not mean)First OSD to cross threshold triggers the warningMax OSD above configured nearfull ratio
ceph_pg_backfill_toofull / ceph_pg_recovery_toofullRecovery already capacity-blockedAny nonzero count
ceph_pool_recovering_bytes_per_secIs recovery making progressZero with degraded PGs present
ceph_health_detail{name="OSD_NEARFULL"}The health check itselfActive
ceph_health_detail{name="POOL_NEAR_FULL"}Pool-level nearfull (quota or pool fullness)Active
ceph_osd_flag_norecover / ceph_osd_flag_nobackfillRecovery intentionally stoppedSet with degraded PGs

The most common monitoring mistake is alerting on a hardcoded 0.85. If an operator tuned mon_osd_nearfull_ratio to 0.80 for a small cluster, your 0.85 alert fires too late. Always alert against the actual configured ratio exposed as ceph_osd_nearfull_ratio.

Fixes

Reweight oversized OSDs (fast, temporary)

If a small number of OSDs are disproportionately full, reweight them down so CRUSH moves data elsewhere:

# Caution: triggers data movement. Do one OSD at a time and watch recovery impact.
ceph osd reweight <osd-id> 0.9

This only helps if other OSDs have room. If the whole cluster is tight, reweighting just moves the problem.

Add capacity (the real fix)

New OSDs are the only durable fix for genuine capacity exhaustion. Adding OSDs gives CRUSH more targets and lowers per-OSD utilization. Plan for the rebalance I/O impact during the addition, especially on clusters where client I/O and recovery share the same network or disks.

Delete data or snapshots

Check for snapshot accumulation, which consumes space invisibly:

ceph df detail

RBD and CephFS snapshots from weeks ago on heavily-written volumes can be enormous. Removing stale snapshots frees space, but the deletion itself generates I/O as Ceph flattens the deltas.

Force RGW garbage collection

If you run RGW and recently deleted large volumes of objects, the data may still be queued for GC rather than removed:

# Processes the GC queue once and exits
radosgw-admin gc process

Deleted objects and aborted multipart uploads are not removed immediately. A GC backlog can consume significant space that ceph df attributes to the pool without any obvious live-data cause.

Temporarily raise backfillfull_ratio (emergency only)

If recovery is already blocked by backfill_toofull and you need it to proceed while you add capacity, you can raise backfillfull_ratio slightly with ceph osd set-backfillfull-ratio. This is risky: it lets OSDs accept more data and pushes them closer to full_ratio, where writes stop entirely. Restore the default as soon as capacity is added. The same caution applies to raising full_ratio to unblock writes: it buys time but does not solve the underlying shortage.

Prevention

  • Alert against the configured ratio, not 0.85. Use ceph_osd_nearfull_ratio as the threshold in your alert, not a hardcoded constant.
  • Track per-OSD max utilization. The cluster average hides the OSD that crosses first. Alert when any OSD exceeds 80% of its configured nearfull ratio.
  • Keep raw utilization below roughly 70%. This is the 20% headroom rule below backfillfull. It is the difference between surviving an OSD failure and deadlocking recovery.
  • Model failure scenarios in capacity planning. Project what happens if you lose your largest failure domain. If the answer is crossing backfillfull, you need more capacity before the next failure, not after.
  • Watch for small clusters. The default thresholds assume enough OSDs that losing one is a small percentage of capacity. On a 3-5 node cluster, losing one node is a large fraction of total space, and the defaults are too aggressive. Consider lowering nearfull for small clusters so the warning fires while there is still time to act.
  • Track daily growth rate. Runway estimation only works if you know your growth rate. Sample ceph df regularly and trend it, then compute days-to-nearfull against the worst-case growth, not the average.

How Netdata helps

  • Per-second per-OSD utilization shows the OSD that crosses nearfull first, not just the cluster average that hides it.
  • Correlate nearfull with PG state. When OSD_NEARFULL fires, Netdata shows ceph_pg_backfill_toofull and ceph_pg_recovery_toofull on the same timeline, so you immediately know whether recovery is already blocked.
  • Recovery rate trending. ceph_pool_recovering_bytes_per_sec next to degraded PG counts tells you whether healing is progressing or stalled, which is the question that actually matters once capacity gets tight.
  • Configured-ratio awareness. Netdata surfaces ceph_osd_nearfull_ratio and ceph_osd_full_ratio as metrics, so alerts track the real threshold even if operators tune it away from the defaults.
  • Per-device-class breakdown. ceph_cluster_by_class_total_bytes separates HDD, SSD, and NVMe capacity, which matters when one device class fills faster than another.