The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-rgw-gc-backlog

Operations Guides

Ceph RGW garbage-collection backlog: deleted data still consuming space

You deleted several terabytes of S3 objects, but ceph df shows raw usage barely moving. The cluster is approaching nearfull and the write freeze is coming. The deletes returned 204 to clients, so they succeeded from the S3 layer’s perspective, but the underlying RADOS objects are still on disk, queued behind the RADOS Gateway garbage collector.

This is one of the quietest contributors to “the cluster is full but we deleted everything”. RGW does not free object data inline on delete. It marks the head object as deleted, enqueues the data objects (tail segments, multipart parts) for asynchronous garbage collection, and relies on a background GC thread on each gateway to drain the queue. When that thread stops making progress, deleted data keeps consuming capacity indefinitely.

The signal that matters is ceph_rgw_gc_retire_object per RGW instance. When the cluster is nearfull and that counter stops moving, you are in this failure mode. This guide walks the diagnosis, the safe ways to force GC forward, and the dangerous shortcuts to avoid.

What this means

GC in RGW is a background worker that reads the shared RADOS GC queue (sharded across rgw_gc_max_objs OMAP shards in the gateway’s GC pool, default 32 shards), selects entries whose rgw_gc_obj_min_wait (default 7200 seconds) has elapsed, and issues RADOS deletes for the data objects. Each gateway processes shards it can lock, so the queue is cluster-wide rather than per instance. The counter ceph_rgw_gc_retire_object, exported per RGW instance, increments on each retire.

If rate(ceph_rgw_gc_retire_object[1h]) is approximately zero while the cluster is nearfull, deleted data is queuing without being reclaimed. Combine this with capacity signals (ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytes and the OSD_NEARFULL health check) and you have the diagnosis.

The operator playbook treats this as a TICKET when both conditions hold for more than an hour: cluster nearfull AND GC retire rate near zero. It is not a hard PAGE because the underlying data is still safe and reads succeed. The danger is the capacity trajectory toward a hard write stop.

flowchart TD
    A["S3 DELETE returns 204"] --> B["Tail and multipart objects\nenqueued for GC"]
    B --> C{"GC thread\ndraining?"}
    C -->|yes| D["ceph_rgw_gc_retire_object\nincrements"]
    D --> E["Capacity freed"]
    C -->|no| F["Queue grows"]
    F --> G["Raw usage flat or rising\ndespite deletes"]
    G --> H{"Cluster nearfull?"}
    H -->|yes| I["TICKET:\nstalled GC plus capacity"]
    H -->|no| J["Latent risk only"]

Common causes

CauseWhat it looks likeFirst thing to check
GC threads disabledceph_rgw_gc_retire_object flat across all RGW instances; queue grows unboundedceph config get client.rgw rgw_enable_gc_threads
GC lock contentionGC runs but slowly; RGW logs show failed to acquire lock on gc.Nradosgw-admin gc list --include-all size vs. retire rate
Min-wait delayRecently deleted objects only; rate recovers within roughly 2 hoursEntry timestamps vs. rgw_gc_obj_min_wait
I/O contentionGC rate drops during deep-scrub or recovery spikesCorrelate with ceph_healthcheck_slow_ops and recovery rate
Orphaned multipart partsGC queue draining, but ceph df still shows growthrgw-orphan-list scan for __shadow* and __multipart* objects
Full target OSDsGC issues RADOS deletes that fail or stallceph osd df for OSDs at backfillfull or full

Quick checks

# Check daemon responsiveness and capacity state
ceph health detail | grep -E 'NEARFULL|FULL|OSD_FULL|SLOW_OPS'

# Per-OSD utilization. The worst OSD matters more than the average.
ceph osd df tree

# Pool-level and raw bytes view
ceph df detail

# GC queue depth, including not-yet-eligible entries.
# Can be slow on a multi-million-entry queue; pipe to wc -l first if needed.
radosgw-admin gc list --include-all | head -50

# Per-instance GC retire counter. The daemon ID does not always match hostname;
# list daemons via `ceph orch ps` or systemctl to find the correct id.
ceph daemon rgw.$(hostname -s) perf dump | jq '.rgw | {gc_retire_object: .gc_retire_object}'

# GC-related tunables
ceph config get client.rgw rgw_enable_gc_threads
ceph config get client.rgw rgw_gc_obj_min_wait
ceph config get client.rgw rgw_gc_max_concurrent_io

All commands above are read-only.

How to diagnose it

  1. Confirm the cluster is actually nearfull. Look at ceph osd df tree, not just ceph df. A single OSD at backfillfull can block writes for every PG mapped to it, even when cluster average utilization is moderate.
  2. Confirm the GC counter is flat. rate(ceph_rgw_gc_retire_object[1h]) should be greater than zero on at least one RGW instance during any window where deletes are happening. If it is zero across every instance for an hour while deletes are happening, GC is stalled.
  3. Inspect the queue. radosgw-admin gc list --include-all shows entries the GC has not yet processed. Compare entry timestamps against rgw_gc_obj_min_wait. If every entry is younger than the min-wait, you do not have a stall, you have a recent delete burst. Wait it out.
  4. Check whether GC threads are running. At least one RGW in each zone must have rgw_enable_gc_threads = true. Read-only or edge gateways sometimes have this disabled, and configuration drift can leave the entire fleet without a GC worker.
  5. Look at RGW logs for lock contention. Repeated RGWGC::process() failed to acquire lock on gc.N lines indicate multiple gateways racing on the same GC shard lock. The work still gets done, just slowly and unevenly.
  6. Verify the OSD layer can actually delete. If the cluster is at OSD_FULL, RADOS delete operations cannot make progress. GC will appear to spin without retiring anything. Freeing capacity anywhere breaks the deadlock.
  7. Distinguish GC backlog from multipart orphans. If radosgw-admin gc list --include-all is small but raw usage keeps climbing, the growth is not the GC queue. It is orphaned __shadow* and __multipart* objects from incomplete uploads. Use the rgw-orphan-list tool to identify them. Do not delete shadow objects by hand.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_rgw_gc_retire_object (counter, per RGW instance)Direct measure of GC progressrate(...[1h]) approximately zero for more than 1 hour while deletes are happening
ceph_cluster_total_used_raw_bytes / ceph_cluster_total_bytesTrue cluster fullness vs. thresholdsTrending toward ceph_osd_nearfull_ratio (default 0.85)
ceph_health_detail{name="OSD_NEARFULL"}The health check that pairs with the TICKETActive alongside a flat GC rate
ceph_health_detail{name="OSD_FULL"}Writes are about to stop cluster-wideActive at all
ceph_healthcheck_slow_opsI/O contention that could starve GCSustained greater than zero
ceph_pool_recovering_bytes_per_secRecovery competing with GC for the same disksHigh while GC is flat
ceph_rgw_req / ceph_rgw_failed_reqSanity check on RGW frontend healthDrop to zero may indicate RGW is down, not just GC stalled

Fixes

Force GC forward manually

# Manually process the GC queue, including not-yet-expired entries.
# Safe to run from any RGW host.
radosgw-admin gc process --include-all

--include-all is what makes this an effective emergency lever. Without it, the command processes only entries whose rgw_gc_obj_min_wait has elapsed, which during a stall may be a no-op if the queue is stuck on something else. With it, GC processes every entry it can acquire a lock on.

Run it once and re-check radosgw-admin gc list --include-all and the retire counter. If the queue shrinks and the counter climbs, the issue was throughput. If it does not, the issue is lock contention, full OSDs, or orphans rather than a real queue backlog.

Raise GC concurrency

If GC is making progress but too slowly to outrun a capacity cliff, raise per-thread concurrency:

# Increase concurrent RADOS deletes per GC thread (default 10 in Quincy, Reef, and Squid).
ceph config set client.rgw rgw_gc_max_concurrent_io 20

This trades OSD I/O headroom for faster draining. Do not raise it during a deep-scrub window or while recovery is saturating disks. Watch ceph_osd_apply_latency_ms and ceph_healthcheck_slow_ops after the change.

Ensure at least one RGW has GC threads enabled

# Global default
ceph config get client.rgw rgw_enable_gc_threads

# Per-instance. Replace <id> with the daemon ID from `ceph orch ps` or systemctl.
ceph tell rgw.<id> config get rgw_enable_gc_threads

If every RGW in the zone has GC threads disabled (a common drift on read-only or edge gateways), set it explicitly on at least one and restart that daemon:

ceph config set client.rgw.rgw1 rgw_enable_gc_threads true
# Then restart rgw1 to pick up the change.

Break the full-cluster deadlock

If GC is stalled because OSDs are at OSD_FULL, RADOS deletes themselves cannot complete. You cannot GC your way out of this from RGW alone. Options, in order of safety:

  1. Add capacity (new OSDs). The cleanest fix.
  2. Delete data the cluster can actually reclaim: snapshots, non-RGW pools, anything where removal translates to immediate raw space.
  3. As a last resort, temporarily raise mon_osd_full_ratio. This is dangerous, buys hours not days, and must be reversed the moment capacity is freed. Treat it as a documented change.

Do not bypass GC casually

The radosgw-admin bucket rm --bypass-gc flag exists (documented on bucket rm; it skips GC bookkeeping for the deleted objects) and looks attractive in a capacity emergency. Do not reach for it without understanding the data-safety implications. Manual deletion of shadow and multipart objects is explicitly unsafe: you cannot reliably distinguish shadow objects of deleted parents from shadow objects of live ones. Prefer radosgw-admin gc process --include-all and let the GC machinery do the bookkeeping.

Prevention

  • Alert on the composite condition. Ticket when the cluster is nearfull AND rate(ceph_rgw_gc_retire_object[1h]) is approximately zero for more than one hour. Either signal alone is noise. Together they are the signature of this failure mode.
  • Verify GC threads on every RGW deployment. Make rgw_enable_gc_threads = true on at least one gateway per zone part of the deployment checklist. Drift here is the most common silent cause.
  • Do not change rgw_gc_max_objs after first deployment. Default is 32. Changing it reshards the GC queue and can strand existing entries.
  • Apply lifecycle policies for incomplete multipart uploads. Use AbortIncompleteMultipartUpload to prevent __multipart* and __shadow* objects from accumulating. GC does not reliably clean these up on its own.
  • Watch GC queue depth during bulk delete operations. Bucket lifecycle expirations, backup rotation, and test-data teardown can enqueue tens of millions of entries in hours. If retire rate cannot keep up, throttle the delete workload before capacity does it for you.
  • Track raw utilization on the worst OSD, not the cluster average. Capacity death spirals start with one backfillfull target, not with cluster-wide fullness.

How Netdata helps

  • Per-second ceph_rgw_gc_retire_object per RGW instance. A flat line on this counter is the leading indicator. Per-second resolution matters because GC retire events are bursty, and minute-level aggregation hides brief stalls.
  • Capacity signals beside the GC counter on the same timeline. ceph_cluster_total_used_raw_bytes, ceph_osd_nearfull_ratio, and the OSD_NEARFULL / OSD_FULL health checks should be read alongside GC retire rate. A flat GC line that is not moving capacity is benign. A flat GC line while raw bytes climb toward nearfull is the TICKET.
  • Pre-built health detail labels. The ceph_health_detail series with the name label surfaces OSD_NEARFULL, OSD_FULL, SLOW_OPS, and LARGE_OMAP_OBJECTS as individual booleans, so you can correlate GC stalls with capacity thresholds and OSD-layer contention without writing the queries yourself.
  • Anomaly detection on retire rate. Netdata flags a drop in GC retire rate before the absolute zero threshold is crossed, which gives lead time before the cluster reaches nearfull.
  • Per-pool and per-OSD capacity drill-down. When GC cannot make progress because target OSDs are full, the per-OSD utilization view identifies the blocker faster than ceph df alone.