The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-osd-fullness-imbalance

Operations Guides

Ceph OSD fullness imbalance: one OSD full while the cluster average looks fine

Cluster writes are blocked. ceph status shows HEALTH_ERR with OSD_FULL active. But ceph df reports the cluster at 65% utilized. Both are correct.

The Ceph full ratio (default 0.95) is enforced per-OSD, not cluster-wide. When any single OSD crosses that threshold, Ceph refuses writes for every PG that OSD serves. Because CRUSH distributes PGs across OSDs, one full OSD can block writes to a large fraction of PGs even when ninety-nine other OSDs have ample free space.

What this means

Three thresholds matter, all enforced per-OSD:

ThresholdDefault ratioWhat happens
nearfull0.85HEALTH_WARN, backfill may be throttled
backfillfull0.90OSD refuses to accept backfill data
full0.95ALL writes blocked cluster-wide

The cluster-wide utilization metric (ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytes) is an average. A single OSD at 96% in a 100-OSD cluster averaging 60% produces a cluster metric that looks healthy. The OSD_FULL health check fires on the individual OSD, and the resulting write block is cluster-wide.

CRUSH places PGs on OSDs proportional to OSD weights. It does not guarantee byte-level evenness. If weights do not match actual disk capacities, or if the balancer is off or not converging, some OSDs fill faster. The first OSD to hit the full ratio stops writes for every PG it serves as primary or replica.

The metric ceph_cluster_by_class_total_bytes (label device_class) shows capacity per class. If one device class is significantly fuller than another, the imbalance is structural, not random.

Common causes

CauseWhat it looks likeFirst thing to check
Mixed disk sizes, wrong CRUSH weightsLarge OSDs and small OSDs have the same CRUSH weight; small OSDs fill firstceph osd df tree, compare SIZE vs WEIGHT columns
Balancer off or not convergingceph balancer status shows a mode but no recent activity; STDDEV stays highceph balancer status
Override reweights left from old incidentREWEIGHT column shows values other than 1.0000 on some OSDsceph osd df tree, look at REWEIGHT column
Cluster was degraded when balancer should have runBalancer does not adjust while OSDs are down; after recovery, imbalance surfacesOSD status history, recovery timeline
Balancer optimizes by PG count, not bytesPG distribution looks even but byte utilization is skewedceph osd df tree, compare PGS vs %USE columns

The last row is the most subtle. The built-in balancer (in upmap mode on modern releases) optimizes by PG shard count, not by actual data bytes. An OSD with the “correct” number of PGs can still be fuller than peers if its PGs happen to contain more data. This is a documented limitation: the balancer cannot solve byte-level imbalance directly.

Quick checks

# Which OSDs are full, nearfull, or backfill_toofull
ceph health detail | grep -iE 'full|nearfull|toofull'

# Per-OSD utilization, hierarchical view
ceph osd df tree

# Balancer status
ceph balancer status

# Recovery flags that block rebalancing (nobackfill, norecover, norebalance)
ceph osd dump | grep flags

# Stuck PGs that may be blocked by full target OSDs
ceph pg dump_stuck unclean

The %USE column in ceph osd df tree shows per-OSD utilization. The VAR column shows each OSD’s deviation from the device-class average. The summary line reports STDDEV across the cluster.

The STDDEV threshold is vendor-documented Red Hat guidance (RHCS 6/7/8 Troubleshooting Guide): if ceph osd df ssd or ceph osd df nvme reports a standard deviation greater than 2.0 for a device class, the balancer may not be enabled or functioning correctly.

How to diagnose it

flowchart TD
    A[Writes blocked, OSD_FULL health check] --> B[Run ceph osd df tree]
    B --> C{One or few OSDs at 95%+?}
    C -->|No| D[Cluster-wide capacity issue]
    C -->|Yes| E[Fullness imbalance confirmed]
    E --> F{STDDEV or VAR high?}
    F -->|Yes| G[Check balancer status]
    F -->|No| H[Check CRUSH weights vs disk sizes]
    G --> I{Balancer on and mode upmap?}
    I -->|No| J[Enable balancer]
    I -->|Yes| K[Check for override reweights]
    K --> L{REWEIGHT not 1.0 on any OSD?}
    L -->|Yes| M[Reset override reweights]
    L -->|No| N[Consider PG count vs bytes limitation]
  1. Confirm the imbalance. Run ceph osd df tree. Identify the OSDs above 85% (nearfull), 90% (backfillfull), or 95% (full). A VAR above 1.5 means that OSD is roughly 50% fuller than the device-class average.

  2. Check the balancer. Run ceph balancer status. Confirm it is active and in upmap mode. If it is on but the cluster was recently degraded (OSDs down), note that the balancer does not run while the cluster is degraded. After recovery completes, it should begin correcting the imbalance.

  3. Check for override reweights. In ceph osd df tree, look at the REWEIGHT column. If any value is not 1.0000, someone set a manual override on that OSD, likely during a past incident. Override reweights conflict with the upmap balancer.

  4. Check CRUSH weights against disk sizes. Compare the SIZE column with the WEIGHT column. CRUSH weight should be proportional to disk size. If a 4TB OSD and an 8TB OSD both show weight 1.0, the 4TB OSD will fill roughly twice as fast.

  5. Check for backfill_toofull PGs. Run ceph pg dump_stuck unclean. PGs stuck in backfill_toofull mean recovery is blocked because the target OSD is above the backfillfull ratio (90%). Even after you free space on the full OSD, recovery may not proceed until target OSDs drop below 90%.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_health_detail{name="OSD_FULL"}Fires when any OSD crosses the full ratio; blocks all writesActive for any duration
ceph_health_detail{name="OSD_NEARFULL"}Fires at 85% per-OSD; early warning before the hard stopSustained for more than 5 minutes
ceph_pg_backfill_toofullPGs blocked from backfill because target OSD is too fullAny non-zero count
ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytesCluster-wide utilization averageAbove 0.80, approaching nearfull territory
ceph_cluster_by_class_total_bytes per device_classPer-device-class capacity; reveals structural imbalance between HDD and SSD tiersOne class significantly fuller than another
ceph_osd_weight per OSDCRUSH weight per OSD; should match disk capacityMismatch between weight and actual disk size
Per-OSD %USE from ceph osd df treeThe actual per-OSD utilization that drives the thresholdsAny OSD more than 10% above the device-class average

The cluster-wide utilization metric is necessary but not sufficient. If your monitoring pipeline only collects ceph_cluster_total_used_raw_bytes, you will not see this problem coming.

Fixes

Immediate: unblock writes

If writes are blocked right now, you need to get the full OSD below 95% quickly.

  • Reweight the full OSD down to reduce new CRUSH placements on it. This triggers recovery to move data off gradually:
# Temporary override reweight; range 0.0 to 1.0
# Triggers PG migration away from this OSD; expect recovery traffic
ceph osd reweight <osd-id> 0.80

Existing data does not move instantly, but the OSD stops accepting new placements and recovery gradually moves data off.

  • Delete data or snapshots. Run ceph df detail to find pools with high snapshot usage. Deleting old RBD or CephFS snapshots can free significant space immediately.

  • Force RGW garbage collection if you use RGW:

radosgw-admin gc process

Deleted RGW objects are queued for GC, not removed immediately. A backlogged GC queue consumes space invisibly.

  • As a last resort, temporarily raise the full ratio. This unblocks writes immediately but risks the OSD filling completely:
# WARNING: if the OSD reaches 100%, it will crash and may not recover cleanly
# Lower it back as soon as writes are unblocked
ceph osd set-full-ratio 0.97

Short-term: correct the imbalance

Once writes are unblocked, fix the structural imbalance.

  1. Reset any override reweights to 1.0:
ceph osd reweight <osd-id> 1.0

Override values fight the balancer.

  1. Ensure the balancer is on and in upmap mode:
ceph balancer on
ceph balancer mode upmap
  1. Wait. The upmap balancer moves PGs incrementally. It will not fix a large imbalance instantly. Monitor ceph osd df tree over hours. VAR values should converge toward 1.0.

  2. If the balancer is not converging and the cluster has mixed disk sizes, check whether upmap_max_deviation (default 5 PG shards, minimum 1) is too coarse for your environment. Lowering it below the default is not something the official documentation recommends; operator reports of excessive PG churn from low values are version-specific, so test any reduction on a non-production cluster first.

Structural: fix CRUSH weights

If the root cause is mismatched CRUSH weights (typically from mixed disk sizes), set CRUSH weights proportional to disk capacity.

CRUSH weights are in units of tebibytes. A 4TB disk should have a weight of approximately 4.0. An 8TB disk should have a weight of approximately 8.0. If all disks were assigned weight 1.0 regardless of size, small disks fill first.

# Check current CRUSH weights and hierarchy
ceph osd tree

# Set CRUSH weight for an OSD (weight is in TiB units)
ceph osd crush reweight osd.<id> <weight>

This changes the placement weight, which triggers PG migration. Do this gradually on production clusters. Changing CRUSH weights on many OSDs simultaneously causes a large recovery burst that competes with client I/O.

What not to do

Do not use ceph osd reweight-by-utilization as a substitute for the upmap balancer: it sets override reweights, and the Ceph docs state that any override reweight value conflicts with the balancer (with the balancer in use, all override reweights should be 1.0000). If you have been using it, stop, reset all override reweights to 1.0, and enable the upmap balancer.

Prevention

  • Monitor per-OSD utilization, not just cluster-wide averages. The ceph_cluster_total_used_raw_bytes metric will not warn you. The per-OSD view from ceph osd df tree is the only reliable leading signal.

  • Keep the balancer on in upmap mode at all times. The balancer does not run when the cluster is degraded. After any degradation event (OSD failures, maintenance), verify that the balancer resumes and the imbalance corrects.

  • Match CRUSH weights to disk capacity at deployment time. Mixing disk sizes without adjusting weights is the most common structural cause.

  • Clear override reweights after incidents. Any ceph osd reweight command sets a temporary override that persists until reset. If you reweight an OSD down during an incident and forget to reset it, the balancer cannot function correctly on that OSD.

  • Track the VAR column trend. A gradually increasing VAR on specific OSDs indicates drift the balancer is not correcting. Investigate before the OSD crosses nearfull.

How Netdata helps

  • Alerting on ceph_health_detail{name="OSD_NEARFULL"} catches the 85% per-OSD condition before the cluster-wide write block at 95%. OSD_FULL alerting is necessary but already too late to prevent the outage.

  • Correlating ceph_pg_backfill_toofull with OSD_NEARFULL or OSD_FULL health checks distinguishes a capacity cliff from a recovery stall. If backfill_toofull is non-zero while OSD_FULL is active, recovery is also blocked, not just writes.

  • Per-device-class metrics (ceph_cluster_by_class_total_bytes) surface structural imbalance between HDD and SSD tiers that the cluster-wide average hides.

  • The ceph_osd_weight metric exposes CRUSH weight per OSD. Correlating this with disk capacity reveals weight mismatches that cause uneven fill rates on mixed-size clusters.

  • Cluster-wide capacity metrics (ceph_cluster_total_used_raw_bytes, ceph_cluster_total_bytes) provide context. When per-OSD signals fire but the cluster-wide metric looks normal, the diagnosis is fullness imbalance.

Per-OSD utilization as a time-series metric depends on the Netdata Ceph collector revision (out of scope for this factual review); ceph osd df tree remains the authoritative per-OSD source of truth.