The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-osd-latency-high

Operations Guides

Ceph OSD commit and apply latency: reading per-OSD latency outliers

Two metrics tell you more about per-OSD health than almost anything else in Ceph: ceph_osd_commit_latency_ms and ceph_osd_apply_latency_ms. They are per-OSD gauges exported by the manager’s Prometheus module, labeled by ceph_daemon, and they measure the OSD’s internal I/O time, not the latency your clients experience end-to-end. A cluster can show healthy aggregate latency while one OSD quietly runs 5x slower than its peers of the same device class. The clients whose objects land on that OSD see tail latency spikes; everyone else sees normal performance. Cluster-wide averages hide this. Per-OSD comparison surfaces it.

This article covers what the two metrics actually measure, why they are equal on BlueStore, what normal looks like by device class, and how to read a single-OSD outlier versus a cluster-wide rise.

What these metrics measure

Both metrics are reported per OSD, labeled ceph_daemon, in milliseconds. They come from OSD perf counters surfaced through the manager’s Prometheus module.

ceph_osd_commit_latency_ms measures the time to persist a write to the WAL/journal path. On a BlueStore OSD with a dedicated WAL or DB device, this reflects the performance of that fast device. On FileStore, when commit latency spikes on one OSD while apply latency stays flat, the WAL/DB path is the bottleneck.

ceph_osd_apply_latency_ms measures the time to apply the write to the main data device: HDD, SATA SSD, or NVMe depending on your tier. High apply latency on a single OSD points at the data device itself. Causes include device saturation, heavy recovery or scrub load, or failing media.

Neither metric includes client-to-OSD network round-trip time. They are internal to the OSD’s processing pipeline. Client-observed latency includes network RTT, replication overhead to peer OSDs, and queuing in the messenger layer. A cluster can look fast at the OSD level while clients experience high latency from network congestion or replica slowness. Correlate OSD latency with slow ops and client-reported latency before concluding the OSD is the root cause.

BlueStore collapses commit and apply

If you are running BlueStore (the default since Luminous), ceph osd perf will show commit and apply latency as identical values on every OSD. This is expected behavior, not a bug: BlueStore feeds both values from the same internal commit-latency counter (BSPerfTracker::update_from_perfcounters assigns the same counter to both).

FileStore is not supported in Reef (18.2.0) and later: the backend was removed from the source tree, and its release notes state “FileStore is not supported in Reef”. On Reef and Squid every OSD runs BlueStore.

BlueStore writes directly to raw block devices and manages its own write-ahead log through BlueFS. Unlike FileStore, where the journal device and the data filesystem were distinct I/O paths with distinct performance characteristics, BlueStore’s commit and apply operations go through the same underlying stack. The two perf counters remain in the interface for compatibility with FileStore-era tooling, but they report the same number.

The practical consequence: on a BlueStore-only cluster, do not try to diagnose a WAL-versus-data-device split from these two metrics alone. You cannot. If you need to isolate WAL/DB performance from data device performance on BlueStore, you need deeper signals: BlueFS slow device usage, RocksDB compaction stats, and OS-level iostat on the individual block devices backing the OSD.

For FileStore clusters still in the wild: commit latency reflects the journal device and apply latency reflects the XFS data device. A meaningful gap between them is diagnosable. High commit with normal apply means the journal device is saturated or failing.

Normal ranges by device class

These ranges are operational baselines under sustained client load, not hard guarantees:

Device classCommit latencyApply latency
HDD10-50ms50-200ms
NVMeless than 5msless than 5ms

SATA SSD falls between these tiers. Exact numbers depend on drive model, firmware, queue depth, and workload mix. Small random writes are the worst case for all device classes.

What matters more than absolute values is comparison within a device class. An HDD OSD at 80ms apply latency is unremarkable if its peers on the same host and same device class sit at 60-90ms. The same OSD is a problem if every other HDD OSD is at 20ms. The signal is the outlier relative to peers, not the absolute number against a universal threshold.

Reading outliers: single OSD vs cluster-wide

The interpretation fork is the most important part of this article. There are two patterns, and they point at completely different root causes.

A single OSD at 5x the same-device-class median, sustained for more than 300 seconds, points at a local problem. Local means: this OSD’s hardware, this OSD’s DB/WAL device, this OSD’s configuration. The rest of the cluster is fine. Common causes include a failing disk (check SMART attributes), a saturated or failing DB/WAL NVMe, RocksDB compaction stalls, BlueStore DB spillover (metadata leaking from the fast DB partition to the slow data partition), or a CRUSH hot spot concentrating a heavy-write workload on one OSD.

A cluster-wide rise (most or all OSDs of a device class climbing together) points at a systemic problem. Systemic means: the network, the recovery load, the scrub schedule, or the cluster’s overall capacity pressure. Common causes include cluster-network congestion saturating the inter-OSD link, an active recovery storm competing with client I/O, deep-scrub running across many OSDs simultaneously, or capacity pressure forcing write amplification.

flowchart TD
    A["OSD latency elevated"] --> B{"One OSD or
cluster-wide?"} B -->|"Single OSD
5x median, >300s"| C["Local problem"] B -->|"Cluster-wide rise"| D["Systemic problem"] C --> C1["Failing disk
check SMART"] C --> C2["DB/WAL NVMe
saturation or failure"] C --> C3["RocksDB compaction
or DB spillover"] C --> C4["CRUSH hot spot"] D --> D1["Cluster network
congestion"] D --> D2["Recovery storm
competing with clients"] D --> D3["Deep scrub
across many OSDs"] D --> D4["Capacity pressure
or write amplification"]

The diagnostic action for each fork differs. For a single-OSD outlier, isolate that OSD: check its device health, its BlueFS stats, its RocksDB state, and its CRUSH weight. For a cluster-wide rise, step back: check recovery rate, scrub schedule, network utilization on the cluster interface, and capacity. Debugging a cluster-wide latency rise by investigating individual OSDs wastes time. Every OSD looks slow because the problem is above them.

The point-in-time snapshot trap

ceph osd perf and the Prometheus metrics are point-in-time snapshots, not histograms or rolling averages. The value at any given sample is the instantaneous latency at the moment of sampling.

This has two consequences. First, a single high reading may be a transient spike: a compaction event, a momentary write burst, a scrub pass. One sample is not an outlier. You need sustained elevation to call it a problem. The 5x device-class median sustained for more than 300 seconds threshold exists to filter out these transient spikes.

Second, because these are snapshots, you cannot derive percentiles from them. You cannot tell whether an OSD has a high P99 with a low mean, or a uniformly elevated latency. If you need distribution data, use the OSD perf histogram via the admin socket (ceph daemon osd.<id> perf histogram dump, available since Luminous), which provides bucketed latency. The Prometheus metrics give you the scalar. The histogram gives you the shape.

When sampling manually with ceph osd perf, run it several times over 10 or more seconds rather than reading a single output. Sustained elevation across samples is the signal. A single jump is noise.

Getting the data

The fastest CLI path:

# Check per-OSD commit and apply latency
ceph osd perf

This prints two columns, commit_latency(ms) and apply_latency(ms), for every OSD. Scan for the OSD that stands out from its same-device-class peers.

For a specific OSD’s deeper perf-counter state:

# Inspect a specific OSD's operation latency perf counters
ceph daemon osd.5 perf dump | jq '.osd | {
  op_latency: .op_latency,
  op_r_latency: .op_r_latency,
  op_w_latency: .op_w_latency,
  op_rw_latency: .op_rw_latency
}'

These are the OSD perf counters that exist; there are no op_commit_latency or op_apply_latency counters in the OSD perf dump. Each entry carries avgcount, sum, and avgtime sub-fields, so average latency between two samples is delta sum / delta count (or read avgtime directly).

ceph_osd_commit_latency_ms and ceph_osd_apply_latency_ms are not derived from these counters: they are the OSD’s objectstore-level commit/apply latency, reported in osd_stats and printed by ceph osd perf as point-in-time values (already in ms).

To check for BlueStore DB spillover, a common cause of latency cliffs that does not show up in capacity metrics:

# Check if RocksDB has spilled to the slow data device
ceph daemon osd.5 perf dump | jq '.bluefs | {
  slow_used_bytes: .slow_used_bytes,
  slow_total_bytes: .slow_total_bytes
}'

Any nonzero slow_used_bytes means metadata has spilled from the fast DB partition to the slow data partition. This is a performance cliff. RocksDB operations that ran at SSD or NVMe speed now run at HDD speed. The symptom is sustained elevated commit and apply latency on that OSD only, while its peers remain normal.

Signals to watch in production

SignalWhy it mattersWarning sign
ceph_osd_commit_latency_ms per OSDReflects WAL/journal device pathSingle OSD sustained at 5x device-class median
ceph_osd_apply_latency_ms per OSDReflects data device pathSingle OSD sustained at 5x device-class median
Cluster-wide latency patternDistinguishes local from systemicAll OSDs of a device class climbing together
ceph_healthcheck_slow_opsOperations stuck beyond osd_op_complaint_time (default 30s)Nonzero count correlates with latency root cause
BlueFS slow_used_bytesDB spillover detectionAny nonzero value on an OSD with a latency outlier
Recovery rate (ceph_pool_recovering_bytes_per_sec)Recovery competes with client I/OHigh recovery rate with elevated latency
Scrub state (active+clean+scrubbing)Scrub saturates disk I/OLatency spikes correlated with scrub schedule

How Netdata helps

  • Netdata’s Ceph collector scrapes ceph_osd_commit_latency_ms and ceph_osd_apply_latency_ms per ceph_daemon at per-second resolution, so you see the outlier OSD the moment it diverges from its peers.
  • Per-second collection makes the “sustained for more than 300 seconds” threshold practical. You can watch the outlier form in real time and confirm it is sustained, not a transient spike from a compaction event or scrub pass.
  • Correlating per-OSD latency with ceph_healthcheck_slow_ops, recovery rate, and scrub activity in a single timeline view lets you distinguish a single-OSD hardware problem from a cluster-wide recovery or scrub impact without switching tools.
  • ML anomaly detection flags OSDs whose latency deviates from their own historical baseline, which catches slow degradation (a disk failing over weeks) that a static 5x-median threshold would miss until the OSD is badly degraded.
  • Device-class labeling lets you compare each OSD against peers of the same class automatically, which matters because HDD and NVMe have different normal ranges.