The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-slow-requests

Operations Guides

Ceph slow requests (SLOW_OPS): operations blocked past osd_op_complaint_time

ceph_healthcheck_slow_ops greater than zero means operations on the cluster have crossed osd_op_complaint_time (default 30s) and are stuck, not merely slow. The SLOW_OPS health check surfaces them as a warning, and the metric itself is a live gauge pulled from ceph health detail.

Slow ops are a symptom of something downstream blocking I/O: a failing disk, a saturated BlueStore RocksDB DB, network timeouts between OSDs, or heavy deep-scrub on HDD during a maintenance window. The right first move is to read where in the pipeline each slow op is stuck before changing any config.

Cluster-wide slow ops across many OSDs are far more serious than isolated slow ops on a single OSD. A handful that clear within a minute or two while deep-scrub or recovery is running is often transient noise on HDD. Sustained slow ops across many OSDs for more than 120 seconds is a structural problem worth escalating.

What this means

Each OSD reports operations that have been in flight past osd_op_complaint_time. The monitor aggregates them and surfaces the count as SLOW_OPS in ceph health detail. The count is the current set of ops that have crossed the threshold. When they complete or are cancelled, the count drops. The metric ceph_healthcheck_slow_ops mirrors that count, so a value of zero means no operation is currently past 30s in flight.

The severity distinction matters. The playbook flags slow ops as TICKET (ceph_healthcheck_slow_ops > 0 sustained > 120s) rather than PAGE, because deep scrub and recovery on HDD clusters can produce transient slow ops that resolve on their own. PAGE-worthy escalation comes when slow ops cascade into PG down or incomplete states, which have their own dedicated alerts.

There is a related but distinct health check, BLUESTORE_SLOW_OP_ALERT, that tracks BlueStore-internal slow operations. It was introduced in Quincy v17.2.8, Reef v18.2.5, and Squid v19.2.1; clusters on older point releases will not see this health check. If you see it alongside SLOW_OPS, the cause is inside BlueStore rather than at the OSD op queue.

flowchart TD
  A[SLOW_OPS active] --> B{How many OSDs?}
  B -- One or two --> C[Check device:
iostat, SMART] B -- Many across failure domains --> D{Recovery or scrub running?} D -- Yes --> E[Throttle:
osd_max_backfills,
noscrub and nodeep-scrub] D -- No --> F{Slow op description?} F -- waiting for subops --> G[Network:
cluster NIC, TCP retransmits] F -- reached pg, stalled --> H[BlueStore:
bluefs slow_used_bytes,
RocksDB compaction] C --> I[Replace disk or reweight OSD] E --> J[Recovery resumes,
slow ops clear] G --> K[Fix link or split networks] H --> L[Migrate DB or compact RocksDB]

Common causes

CauseWhat it looks likeFirst thing to check
Disk I/O stall on one OSDSlow ops confined to one or two OSDs; ceph_osd_apply_latency_ms on that OSD is many times its peersiostat -xz on the backing device and SMART attributes
Op-queue saturation or recovery pressureSlow ops spread across many OSDs; high recovering and backfilling PG counts; recovery line in ceph -s is highceph pg stat and recovery rate; whether noscrub/nodeep-scrub would help
BlueStore kv_sync blocked or RocksDB compaction stallPeriodic latency spikes on specific OSDs; commit latency spikes between compaction eventsceph daemon osd.<id> perf dump bluefs and rocksdb sections
Inter-OSD network timeoutSlow ops reported “waiting for subops from ” across a failure domain; TCP retransmits risingCluster-network utilization via ip -s link, /proc/net/snmp retransmits
BlueStore DB spillover to slow deviceLatency cliff on specific OSDs with no disk errors; bluefs slow_used_bytes > 0ceph daemon osd.<id> bluefs stats
Deep scrub on HDDTransient slow ops that line up with the configured scrub window; PGs in active+clean+deep-scrubbing`ceph pg dump

Quick checks

# Triage: is SLOW_OPS the active check, and how many ops are stuck?
ceph health detail | grep -A 5 SLOW_OPS

# Live count from the monitor, machine-readable
ceph health detail -f json | jq '.checks.SLOW_OPS'

# Event timeline for every in-flight operation on a named OSD
ceph daemon osd.<id> dump_ops_in_flight

# Operations stuck past the blocked threshold
ceph daemon osd.<id> dump_blocked_ops

# Read the "where it's stuck" description in the OSD log
grep "slow request" /var/log/ceph/ceph-osd.<id>.log | tail -n 50

# Commit and apply latency per OSD, comparable across device classes
ceph osd perf

# Per-pool PG state, including deep-scrub and recovery activity
ceph pg stat

# Confirm whether recovery or scrub flags are silently set
ceph osd dump | grep flags

How to diagnose it

  1. Confirm scope. ceph health detail | grep -A 5 SLOW_OPS shows how many slow ops are in flight and which OSDs they belong to. If only one or two OSDs are listed, focus on those hosts. If the list spans many OSDs across multiple failure domains, treat it as a cluster-wide resource problem.

  2. Read where the ops are stuck. For each named OSD, ceph daemon osd.<id> dump_ops_in_flight shows the event timeline for every in-flight operation. The most recent event description tells you what each op is waiting on:

    • waiting for subops from <osd>: replication stall. The named replica is the bottleneck, not the primary.
    • waiting on pg: the PG is in a non-active state, usually peering or recovery.
    • waiting for rw locks: contention on a busy object, often many writers to the same PG.
    • reached pg followed by a long stall: the local device or BlueStore pipeline is the bottleneck.
  3. Correlate with per-OSD latency. ceph osd perf prints commit and apply latency per OSD. Compare each slow OSD to its peers on the same device class. An OSD whose latency is many times its peers’ is your failing component. Typical ranges under load: HDD OSDs show commit latency 10-50ms and apply latency 50-200ms; NVMe OSDs should show both under 5ms. These vary with DB device class and workload.

  4. Check what else the OSDs are doing. Slow ops that arrive during recovery or deep-scrub are usually resource contention, not failure. ceph pg stat shows how many PGs are recovering, backfilling, scrubbing, or deep-scrubbing. ceph -s shows the recovery I/O rate. If those line up with the slow ops, throttle or schedule.

  5. Look at BlueStore internals. If a single OSD has periodic latency spikes with no obvious device fault, inspect ceph daemon osd.<id> perf dump under the bluefs and bluestore sections, and ceph daemon osd.<id> bluefs stats. In the bluefs stats output, a nonzero slow_used_bytes means RocksDB has spilled from the fast DB partition to the slow data device. This is a cliff-edge failure with a 10-100x latency increase that is invisible in capacity metrics.

  6. Check the network. If slow op descriptions repeatedly say waiting for subops from <osd>, the named replica is the bottleneck. Use ip -s link show <cluster_iface> and cat /proc/net/snmp (the TCPRetransSegs counter) to see whether the cluster network is saturated or losing packets.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_healthcheck_slow_opsDirect count of ops past osd_op_complaint_time> 0 sustained for > 120s
ceph_osd_apply_latency_msTime to apply an op to the local storeSingle OSD > 5x the median of its device class
ceph_osd_commit_latency_msTime to durable commit, includes replicationSudden jump on dedicated WAL/DB indicates failing fast device
ceph_pg_recovering, ceph_pg_backfillingConfirms whether slow ops are recovery contentionHigh counts plus slow ops = expected, throttle if needed
ceph_pool_recovering_bytes_per_secWhether recovery is making progress or stalledZero with degraded PGs = stuck recovery, separate issue
PGs in scrub/deep-scrub state (check via ceph pg dump | grep scrub)Confirms scrub contention as cause of slow opsSlow ops coinciding with high scrub PG count = throttle or reschedule
TCP retransmits on cluster NICNetwork-driven replication stallsRising retransmits plus slow ops = link-layer problem

Fixes

Single failing disk

If slow ops are isolated to one OSD and its latency is 5x or more of its peers, the device is likely failing or saturated.

  • Confirm with iostat -xz /dev/<dev>: high %util, growing await, growing avgqu-sz.
  • Check SMART: smartctl -A /dev/<dev>. Nonzero Reallocated_Sector_Ct, Current_Pending_Sector, or Offline_Uncorrectable indicates active media errors.
  • If the disk is failing, drain and replace. Mark the OSD out and down cleanly, let recovery complete, then replace the media and redeploy.

Do not just restart the OSD. Restarting does not fix a failing disk and may disrupt clients further.

Recovery or scrub contention

If slow ops arrive while recovery or deep-scrub is running on HDD OSDs, throttle before considering anything else.

  • Reduce recovery pressure: ceph tell 'osd.*' injectargs '--osd_max_backfills 1 --osd_recovery_max_active 1'. This is a runtime change that takes effect immediately but does not persist across OSD restart. Use ceph config set osd <key> <value> for persistence.
  • Defer scrub during the incident: ceph osd set noscrub and ceph osd set nodeep-scrub. Unset them when the incident is over; leaving them set silently disables data integrity verification.
  • Confirm whether noout, norecover, or nobackfill are set and forgotten: ceph osd dump | grep flags.

RocksDB compaction stall or BlueStore kv_sync block

Periodic latency spikes on a specific OSD with no disk errors often point to RocksDB compaction. The bluefs and rocksdb sections of ceph daemon osd.<id> perf dump show compaction activity.

Manual compaction is a short-term lever:

# Compact the BlueStore RocksDB. The OSD must be stopped first;
# running this on a live OSD corrupts the DB.
systemctl stop ceph-osd@<id>
ceph-kvstore-tool bluestore-kv /var/lib/ceph/osd/ceph-<id> compact
systemctl start ceph-osd@<id>

This is a temporary mitigation, not a root-cause resolution. The underlying cause is often a mismatch between DB load and slow hardware, or DB spillover.

BlueStore DB spillover to slow device

If ceph daemon osd.<id> bluefs stats shows nonzero slow device usage, RocksDB has spilled from the fast DB partition to the slow data partition. Latency jumps 10-100x on every metadata operation.

  • Confirm with ceph daemon osd.<id> bluefs stats and check the slow device usage line.
  • Migrate the DB to a larger device using ceph-bluestore-tool. This requires OSD downtime and careful planning.
  • Short term, reweight the affected OSD down to reduce its object count: ceph osd reweight <id> 0.9.

Inter-OSD network timeout

If slow op descriptions consistently point to waiting for subops from <osd> across a failure domain:

  • Check whether the cluster network is shared with the public network. If so, recovery traffic is competing with client I/O.
  • Verify MTU consistency. Jumbo frames configured on some ports but not others cause silent fragmentation that looks like disk slowness.
  • Look for rising TCP retransmits in /proc/net/snmp (TCPRetransSegs).

mClock shard misconfiguration on HDD

In Squid 19.2.1 the mClock HDD shard defaults changed: osd_op_num_shards_hdd moved from 5 to 1 and osd_op_num_threads_per_shard_hdd moved from 1 to 5 (tracker.ceph.com/issues/66289). The same defaults were backported into Reef 18.2.5. Clusters on earlier point releases of Squid (below 19.2.1) or Reef (below 18.2.5) with the old defaults can exhibit slow requests on HDD OSDs under mClock scheduling; apply the new defaults via ceph config set after validating the change in a non-production tier.

Tuning osd_op_complaint_time

The default 30 seconds is appropriate for most deployments. Some operators lower it to 5-10 seconds for latency-sensitive workloads to get earlier visibility, but lowering the threshold increases noise from legitimate transient stalls on HDD clusters. Treat this as observability tuning, not a fix.

Prevention

  • Monitor per-OSD latency, not cluster averages. A single OSD with 5x the latency of its peers shows up only in per-OSD signals. Cluster-average latency looks fine while one OSD silently degrades.
  • Track bluefs slow_used_bytes. Spillover is a cliff-edge failure with no warning in capacity metrics.
  • Track RocksDB compaction stats. Compaction stalls are a common root cause of latency spikes but are almost never proactively monitored.
  • Separate public and cluster networks. Recovery traffic can saturate a shared NIC; the fix is physical separation, not throttling alone.
  • Schedule deep-scrub windows explicitly. osd_scrub_begin_hour and osd_scrub_end_hour exist for a reason. Document them and alarm on noscrub and nodeep-scrub set for more than 24 hours.
  • Validate DB partition sizing at deployment. Aim for 4-5% of the data partition for replicated pools, larger for EC or heavy omap usage. Revisit when workload changes.
  • Watch for forgotten noout. The “noout trap” is the single most common preventable Ceph outage.

How Netdata helps

  • ceph_healthcheck_slow_ops is collected per second, so you can pinpoint the exact second the count crosses zero and correlate it with what else changed.
  • Per-OSD ceph_osd_apply_latency_ms and ceph_osd_commit_latency_ms let you isolate the failing OSD against its device-class peers instead of staring at a cluster average.
  • PG state metrics (ceph_pg_recovering, ceph_pg_backfilling) and recovery rate metrics let you confirm or rule out recovery contention before changing config.
  • Co-located host-level disk metrics (%util, await, avgqu-sz) and network metrics (TCP retransmits, interface saturation) appear next to the Ceph metrics, so you can attribute slow ops to a device, a network, or a scheduler without pivoting tools.