The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-osd-up-down-in-out

Operations Guides

Ceph OSD up/down vs in/out: the four states and mon_osd_down_out_interval

Most Ceph operations references describe an OSD as “down” as if it were a single state. It is not. An OSD carries two independent flags: up/down (is the daemon alive?) and in/out (does CRUSH place data on it?). Combined, that produces four states, and the most dangerous one, down+in, is invisible if you only look at one flag.

The two flags are linked by a clock. After an OSD goes down, the monitors wait mon_osd_down_out_interval (default 600 seconds) before automatically flipping it to out. That 10-minute window is the gap between “OSD stopped” and “recovery starts.” That timer, and the flags on either side of it, are the vocabulary every recovery, flapping, and noout runbook assumes.

This is a reference, not a troubleshooting flow. It maps the four states, explains the timer that transitions between them, and identifies the signals you need to distinguish a transient blip from a recovery in progress.

What it is and why it matters

The up/down and in/out flags live in the OSD map and are maintained by the monitors. They answer different questions.

  • up/down (ceph_osd_up): is the OSD daemon running and responding to heartbeats? This is daemon liveness. The monitors mark it down after peer OSDs report the daemon as unresponsive past osd_heartbeat_grace (default 20 seconds), with mon_osd_min_down_reporters (default 2) reporters from different subtrees at mon_osd_reporter_subtree_level (default host).
  • in/out (ceph_osd_in): is the OSD participating in CRUSH placement? This is data placement. in means new PGs can be assigned to it and existing PGs expect data to live there. out means CRUSH has excluded it.

The two are independent in principle but linked by mon_osd_down_out_interval. An OSD that goes down is not immediately out. The monitors wait the configured interval first, giving the operator (or systemd) a chance to bring the daemon back before the cluster starts the expensive work of moving its data elsewhere.

This distinction matters because every recovery runbook, every noout warning, and every flapping cascade is described in terms of these four states. If you read “OSD is down” without knowing whether it is also in, you do not yet know whether recovery has started.

How it works

The four states, with the operational meaning of each:

Stateup/downin/outMeaning
up + inupinNormal. Daemon running, CRUSH places data here.
down + indowninDaemon stopped, but CRUSH still expects data here. The mon_osd_down_out_interval clock is running. PGs are degraded but recovery has not started.
up + outupoutDaemon running but excluded from CRUSH. Intentional maintenance, or recently returned from outage and pending auto-mark-back-in.
down + outdownoutDaemon gone and CRUSH has moved on. Data is being or has been recovered to other OSDs.

The dangerous state is down+in. The data is “expected” on a daemon that is not answering. PGs served by that OSD are degraded, and the cluster has not yet committed to recovery. This is the window where a second failure in the same failure domain can cause data loss.

stateDiagram-v2
    [*] --> UP_IN
    UP_IN: up + in (normal)
    DOWN_IN: down + in (degraded, no recovery yet)
    UP_OUT: up + out (maintenance)
    DOWN_OUT: down + out (recovering)
    UP_IN --> DOWN_IN: daemon stops / heartbeat timeout
    DOWN_IN --> UP_IN: restart within mon_osd_down_out_interval
    DOWN_IN --> DOWN_OUT: mon_osd_down_out_interval expires (600s), unless noout set
    DOWN_OUT --> UP_OUT: daemon restarts after auto-out
    UP_OUT --> UP_IN: ceph osd in, or mon_osd_auto_mark_auto_out_in
    UP_IN --> UP_OUT: ceph osd out (manual)
    note right of DOWN_IN
        noout flag blocks
        transition to DOWN_OUT
    end note

The transitions, in order of how often you will see them:

  1. Daemon stops. OSD moves from up+in to down+in. Peers mark it down via heartbeat. The mon_osd_down_out_interval timer starts.
  2. Daemon restarts within the window. OSD moves back to up+in. PGs recover via PG log replay, which is relatively cheap. No backfill.
  3. Timer expires without restart. OSD moves to down+out. CRUSH recalculates placement; recovery and backfill begin to other OSDs.
  4. Daemon returns after auto-out. OSD moves to up+out. Because mon_osd_auto_mark_auto_out_in defaults to true, it will be marked back in automatically, triggering backfill to rebalance data onto it. This is the expensive path: a brief outage that exceeded 600 seconds produces a full backfill in both directions.
  5. Manual ceph osd out. Moves up+in to up+out without taking the daemon down. Used for planned evacuations.
  6. Manual ceph osd in. Reverses it.

Verified syntax: per-OSD noout is ceph osd add-noout osd.<id> (and ceph osd rm-noout osd.<id>); per-CRUSH-bucket noout is ceph osd set-group noout <bucket> (comma-separated flags from {noup,nodown,noin,noout}) with ceph osd unset-group to clear it. No preconfigured “group” object is needed; the target is a CRUSH bucket or device class.

The noout flag blocks transition 3. With noout set, a down+in OSD stays down+in indefinitely. PGs remain degraded and recovery never starts. This is intentional during brief maintenance, and the cause of the “noout trap” when it is left set after maintenance ends.

A related but separate concept is ceph_osd_weight. Setting an OSD’s reweight to 0 excludes it from placement regardless of its in/out state. This is a soft eviction used during investigation. It stops new placements without triggering the full OUT transition machinery, and it does not produce the same OSD map churn.

Where it shows up in production

Host reboot during maintenance. You reboot a host with 10 OSDs. All 10 go down+in. You have 600 seconds to bring them back before auto-out fires and the cluster starts moving terabytes of data. This is exactly why operators set noout (or per-OSD ceph osd add-noout osd.<id>) before planned reboots. Set it, do the work, unset it.

Brief network blip on the cluster network. OSDs lose heartbeats to peers, get marked down+in. Network recovers in 30 seconds. OSDs come back up+in. PG log replay handles it. No backfill. The 600-second window did its job.

OSD segfault. Daemon dies, systemd does not restart it (or restarts it slowly). OSD sits at down+in for 600 seconds, then transitions to down+out. Recovery starts. If this was a 3x-replicated pool and another OSD in the same PG’s acting set fails during the 600-second window, you have a data availability problem.

Flapping OSD. Daemon oscillates between up+in and down+in faster than the 600-second timer can expire. Each flap generates a new OSD map epoch, triggering peering across hundreds of PGs. The OSD never reaches down+out, so recovery never starts, but the peering storms damage cluster performance. The fix is ceph osd add-noout osd.<id> to stop the peering cascade, then investigate the underlying cause.

Forgotten noout. Maintenance finishes, noout is left set. Over the next week, three OSDs fail. All three sit at down+in. PGs accumulate in degraded state. The cluster appears to be limping but functional, until a fourth failure tips some PGs into incomplete. This is the most common preventable Ceph outage.

When this matters

The 600-second window is a tuning dial, not a constant. mon_osd_down_out_interval defaults to 600s and most clusters leave it alone. Shortening it speeds recovery start (less exposure) but increases the chance that a transient blip triggers an expensive full backfill. Lengthening it gives transient issues more time to self-resolve but extends the degraded window. Do not tune it without thinking about both sides.

mon_osd_adjust_down_out_interval (default true) scales the interval automatically. When enabled, the monitors increase mon_osd_down_out_interval when an OSD “appears to be laggy” (the option’s documented purpose). This can help clusters with a few slow OSDs but can mask deteriorating hardware. If you are trying to understand why a particular OSD took longer than 600 seconds to be marked out, this is why.

mon_osd_down_out_subtree_limit (default rack) blocks auto-out at scale. If the monitors detect that all OSDs within a rack are down, they will not auto-mark them out. This prevents a rack-wide outage from triggering a recovery storm that would saturate the surviving racks. If your CRUSH hierarchy is flat and lacks a rack level, this safeguard may not apply. Verify your topology if you run flat.

noout is a surgical instrument, not a default. Per-OSD noout (ceph osd add-noout osd.<id>) is safer than the cluster-wide flag because the cluster-wide version is easy to forget. Per-CRUSH-bucket noout is the right tool for maintenance on a whole host or rack.

Setting mon_osd_down_out_interval to 0 disables auto-out entirely. The OSD_NO_DOWN_OUT_INTERVAL health check warns about this. Some operators do it deliberately on clusters where they want manual control over every OUT transition. If you do this, your OSD-down alerting needs to be solid, because the safety net is gone.

Signals to watch in production

SignalWhy it mattersWarning sign
ceph_osd_up per OSDDaemon liveness. The raw up/down flag.Transition to 0. Aggregate count of ceph_osd_up == 0.
ceph_osd_in per OSDCRUSH participation. The raw in/out flag.Transition to 0 not explained by planned maintenance.
ceph_osd_flag_nooutWhether the auto-out transition is blocked.Set for more than 24 hours without an open maintenance ticket.
OSD dwell time in down+inHow long the cluster has been exposed with degraded PGs and no recovery.Any OSD with up == 0 AND in == 1 for more than a few minutes outside maintenance.
OSD map epoch rateRate of topology changes. Each up/down transition increments the epoch.Sustained rate of more than a few per minute indicates flapping.
ceph_health_detail{name="OSD_FLAPPING"}Pattern detection on top of up/down transitions.Active. Flapping does not always show as a clean down count.
ceph_osd_weight per OSDSoft eviction via reweight 0, independent of in/out.Unexpected reweight of 0 not tied to an investigation.

The single most useful correlation is ceph_osd_up == 0 AND ceph_osd_in == 1, sustained. That is the down+in window, and its duration is the cluster’s exposure to a second failure without recovery in progress.

How Netdata helps

  • The Ceph collector exposes ceph_osd_up and ceph_osd_in per OSD, so you can alert on the four states directly rather than relying on the cluster-wide down count.
  • The down+in window (ceph_osd_up == 0 AND ceph_osd_in == 1) is the most operationally meaningful state. Per-second granularity lets you see exactly when an OSD crossed into it and whether it recovered or transitioned to down+out.
  • Correlating ceph_osd_flag_noout with down OSDs identifies the noout trap before it causes a cascading failure.
  • OSD map epoch churn is visible as a rate on the OSD state metrics. Sustained churn indicates flapping before the OSD_FLAPPING health check fires.
  • ML-based anomaly detection on per-OSD state transitions surfaces flapping patterns that simple thresholds miss, useful for the marginal disks that fail slowly.
  • Per-pool PG state metrics sit alongside the OSD state metrics, so you can confirm whether a down+in OSD is actually producing degraded PGs and whether recovery has started.