The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-pg-inconsistent

Operations Guides

Ceph PG inconsistent (OSD_SCRUB_ERRORS): scrub found replica divergence

A scrub or deep-scrub finished comparing replicas for a placement group and found they disagree. The cluster surfaces this as OSD_SCRUB_ERRORS (often paired with PG_DAMAGED) and the affected PG sits in active+clean+inconsistent. Client reads still succeed because Ceph serves them from a consistent replica, but at least one copy in the acting set is corrupt.

The danger is not the symptom. Reads work, and Ceph did what it was designed to do: detect silent divergence. The danger is that the corruption was found, not fixed, and the window during which an uncorrupted replica survives is your margin of safety. If the OSD holding the good copy fails before you repair, the object becomes unreadable or unwritable.

The second danger is the repair command itself. ceph pg repair {pgid} resolves the inconsistency by copying data from the primary to the replicas. If the primary is the corrupt copy, repair propagates the corruption to the surviving replicas and destroys the good data. The Ceph documentation explicitly warns against running pg repair before you have identified which replica is authoritative.

The safe path: identify the PG, enumerate the inconsistent objects, determine which OSD holds the corrupt copy, then repair (or manually remove the bad object) once you know the source of truth.

What this means

The inconsistent PG state is set when a scrub finds that replicas disagree on object data, size, metadata, omap (extended attributes or key-value data), or snapset information. A light scrub compares metadata. A deep-scrub additionally checksums object data, which is what catches bit rot that left all metadata intact.

When Ceph detects an inconsistency, it does not silently fix it. The PG remains active+clean+inconsistent. Reads are served from the primary or another consistent replica. Writes continue, but the inconsistency record is preserved until a repair resolves it.

Health checks you will see:

Health checkMeaning
OSD_SCRUB_ERRORSActive when any scrub or deep-scrub has reported errors.
PG_DAMAGEDActive when PGs are in a damaged state, including inconsistent.

Relevant Prometheus metrics from the MGR module:

  • ceph_pg_inconsistent (per pool_id): count of PGs in the inconsistent state. Alert on this.
  • ceph_pg_failed_repair (per pool_id): count of PGs where automatic repair was attempted and failed. Escalation signal, not a normal state.
  • ceph_pool_objects_repaired (counter, per pool_id): cumulative count of objects repaired.

Severity is TICKET for sum(ceph_pg_inconsistent) > 0 of any duration. Reads still work, but this is a “before the good replica’s OSD dies” ticket, not a “next week” ticket. If automatic repair has been tried and failed (ceph_pg_failed_repair > 0), it becomes urgent manual intervention.

flowchart TD
    A[Scrub finds inconsistency] --> B[PG: active+clean+inconsistent]
    B --> C[Health: OSD_SCRUB_ERRORS / PG_DAMAGED]
    C --> D{Run ceph pg repair?}
    D -- BLIND --> E[Primary may be corrupt]
    E --> F[Corruption propagates to good replicas]
    D -- DIAGNOSE FIRST --> G[rados list-inconsistent-obj]
    G --> H[Identify bad replica via SMART + shard errors]
    H --> I[Safe repair or manual object removal]

Common causes

CauseWhat it looks likeFirst thing to check
Silent bit rot on a diskDeep-scrub reports data_digest_mismatch on one OSD’s copy; SMART shows reallocated or pending sectors.smartctl -A on the OSD device and ceph device get-health-metrics.
Disk firmware bugInconsistencies appear across multiple OSDs of the same model or firmware batch, often after a firmware update.Vendor advisories, firmware version, correlation across OSDs.
Non-ECC or failing RAM on an OSD hostdata_digest_mismatch or omap_digest_mismatch on recently written objects, no SMART errors.dmesg for EDAC / memory errors, host RAM type.
Torn write from power loss without barriersInconsistency appears after a host crash or power event, on the OSD that was writing.Recent power events, WAL integrity.
Ceph or BlueStore bugInconsistency appears without hardware cause, often clustered on a version or workload pattern.Ceph tracker, version, workload signature.
Metadata-level inconsistency (snapset, omap)rados list-inconsistent-obj returns no objects but PG remains inconsistent.rados list-inconsistent-snapset {pgid}.

Quick checks

These commands are read-only and safe to run.

# Confirm which PGs are inconsistent and from which health check
ceph health detail | grep -E 'OSD_SCRUB_ERRORS|PG_DAMAGED|inconsistent'

# List inconsistent PGs cluster-wide
ceph pg ls inconsistent

# Per-pool: list PGs flagged inconsistent
rados list-inconsistent-pg {pool}

# Acting set and primary for a specific PG
ceph pg {pgid} query | jq '.acting, .up, .info'

# Object-level inconsistencies inside a PG
rados list-inconsistent-obj {pgid} | jq .

# Snapset inconsistencies (when object list is empty)
rados list-inconsistent-snapset {pgid} | jq .

# Recent scrub results for the PG
ceph pg {pgid} query | jq '.info.stats.scrub_stats'

# SMART health on every OSD in the acting set
ceph device ls-by-daemon osd.{id}
ceph device get-health-metrics {devid}

If rados list-inconsistent-obj returns “Operation not permitted”, the client key lacks caps. It needs allow r on the pool. Re-run as client.admin or grant the missing cap.

If rados list-inconsistent-obj returns an empty object list for a PG the cluster still marks inconsistent, the inconsistency is at the PG metadata level. Check rados list-inconsistent-snapset {pgid} and the inconsistency payload in ceph pg {pgid} query.

How to diagnose it

The diagnosis has one goal: figure out which OSD holds the corrupt copy before issuing any write.

  1. Identify the inconsistent PGs and their acting sets.

    ceph pg ls inconsistent -f json | jq '.pg_stats[] | {pgid, state, acting}'
    
  2. For each inconsistent PG, list the inconsistent objects and read the error codes.

    rados list-inconsistent-obj {pgid} | jq '.inconsistencies[0]'
    

    The errors array on each entry tells you the failure mode:

    ErrorMeaning
    data_digest_mismatchObject content differs between replicas.
    size_mismatchObject sizes differ.
    omap_digest_mismatchOMAP (extended attributes / key-value) differs.
    read_errorOne replica returned a read error, often a bad sector.
  3. Cross-reference each shard in the shards array against the acting set from step 1. Each shard entry names the OSD (osd) and whether its copy was considered authoritative. The OSD that is not authoritative, or that reports read_error, is the suspect.

  4. Pull SMART for every OSD in the acting set, not just the suspect. The goal is independent evidence about which device is unhealthy.

    ceph device ls-by-daemon osd.{id}
    ceph device get-health-metrics {devid}
    smartctl -A /dev/{device}
    

    Rising Reallocated_Sector_Ct, Current_Pending_Sector, or Offline_Uncorrectable is strong evidence that this OSD’s copy is the corrupt one.

  5. Check host-level memory health if SMART is clean and the inconsistency pattern matches recent writes.

    dmesg -T | grep -iE 'edac|memory|hardware error'
    

    Non-ECC RAM on an OSD host, or a failing DIMM with EDAC reporting correctable errors that occasionally become uncorrectable, can produce data_digest_mismatch with no disk fault.

  6. Decide on the authoritative copy. If two replicas agree and one disagrees, the two agreeing copies are authoritative. If the primary is the outlier, blind ceph pg repair will propagate the corruption.

  7. Only after you know the source of truth, proceed to repair.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_pg_inconsistent (per pool)Direct count of PGs with scrub-found divergence.Any non-zero value, any duration.
ceph_pg_failed_repair (per pool)Auto-repair attempted and failed.Non-zero. Escalate to manual repair.
ceph_health_detail{name="OSD_SCRUB_ERRORS"}Health check firing on scrub errors.Active.
ceph_health_detail{name="PG_DAMAGED"}Umbrella check for damaged PGs.Active.
ceph_pool_objects_repaired (counter)Confirms repair is making progress after you trigger it.Flat when you expect it to increase.
ceph_health_detail{name="PG_NOT_DEEP_SCRUBBED"}Verification debt. Deep scrubs falling behind means corruption goes undetected longer.Active for more than 24 hours.
Per-OSD commit and apply latency on the acting setCorrelates a failing device with the bad replica.Latency outliers on the OSD holding the suspect copy.
SMART attributes on OSD devicesHardware-level precursor to bit rot.New reallocated or pending sectors.

Fixes

The right fix depends on which replica is corrupt. The single rule: never run ceph pg repair until you have evidence about which copy is authoritative.

Standard case: primary is the good copy, one replica is corrupt

This is the safe path the repair command was designed for.

# Re-verify the inconsistency
rados list-inconsistent-obj {pgid} | jq .

# Run repair
ceph pg repair {pgid}

# Watch progress: ceph_pool_objects_repaired should increase
# and the PG should return to active+clean
ceph pg {pgid} query | jq '.info.stats'

Repair triggers a recovery-style operation that overwrites the inconsistent copy with the primary’s data. On BlueStore, internal checksums make authoritative copy selection more reliable than on legacy FileStore, but the primary bias still exists when checksums are unavailable, which is why the diagnosis step is non-negotiable.

Dangerous case: primary is the corrupt copy

This is the scenario called out in the official docs and vendor knowledge bases. The primary OSD has the wrong digest and the two replicas agree with each other. Running ceph pg repair here overwrites the good replicas with corrupt data from the primary.

The fix is to remove the bad object from the primary OSD using ceph-objectstore-tool, then trigger repair so the primary pulls the object back from a good replica.

WARNING: this procedure is destructive and requires stopping the OSD. The exact ceph-objectstore-tool invocation is version-dependent. On Reef and later (FileStore removed, BlueStore only), refer to the current ceph-objectstore-tool documentation for the remove and list syntax. Test on a non-production OSD first.

High-level steps:

  1. Stop the primary OSD: systemctl stop ceph-osd@{id}.
  2. Use ceph-objectstore-tool with --op remove against the OSD’s data path, the PG ID, and the object ID to remove the bad object from the primary’s local store.
  3. Mark the object as missing on that OSD so the next read pulls it from a peer.
  4. Start the OSD: systemctl start ceph-osd@{id}.
  5. Trigger repair: ceph pg repair {pgid}.

Repair does not resolve: failed_repair

If ceph_pg_failed_repair is non-zero for the PG, automatic repair was attempted and failed. Common reasons:

  • The inconsistency is at the snapset or PG metadata level, not the object level, so the object-copy repair path does not apply.
  • The OSD that holds the authoritative copy is down or unfound.
  • An underlying device error prevents the read.

Run rados list-inconsistent-snapset {pgid} for metadata-level inconsistencies, and check whether the authoritative OSD is up and reachable. If the authoritative copy is on a down OSD, recovering that OSD (or declaring the object lost via ceph pg mark_unfound_lost revert|delete) may be required.

Erasure-coded pools

Repair semantics differ for EC pools. The rados list-inconsistent-obj output includes shard-level information, and repair uses the EC plugin’s reconstruction logic. The primary bias concern still applies: if the OSDs holding the corrupt shard are also the ones the primary would consult as authoritative, blind repair can propagate bad data. Walk the same diagnosis path: enumerate shards, find the disagreeing one, verify SMART, then repair.

Automatic repair

osd_scrub_auto_repair (default false) enables automatic repair when scrub finds errors, up to osd_scrub_auto_repair_num_errors (default 5). It applies to BlueStore and EC pools. If you have this enabled, your monitoring of ceph_pg_inconsistent becomes more important, not less: auto repair quietly fixes single-object divergences, which means the underlying cause (bad RAM, marginal disk, firmware bug) can keep producing new inconsistencies without you ever seeing a ticket. Track ceph_pool_objects_repaired as a rate. Any non-zero rate is a sign to investigate.

Prevention

  • Keep deep scrubs running. Suppressing noscrub and nodeep-scrub is the single most common way operators miss corruption until it spreads. If you must suppress for performance, set a duration limit and alert if the flags are set for more than 24 hours.
  • Watch scrub recency. ceph_health_detail{name="PG_NOT_DEEP_SCRUBBED"} active for more than 24 hours means verification debt is accumulating and corruption can grow undetected.
  • Use ECC RAM on OSD hosts. Non-ECC RAM produces data_digest_mismatch with no disk fault and is one of the hardest causes to diagnose.
  • Monitor SMART aggressively. Reallocated_Sector_Ct, Current_Pending_Sector, and Offline_Uncorrectable are leading indicators of the bit rot that scrub eventually catches.
  • Do not enable osd_scrub_auto_repair without also alerting on ceph_pool_objects_repaired rate. Silent auto repair masks the underlying hardware cause.
  • Make the operator rule explicit: never run ceph pg repair without first running rados list-inconsistent-obj. Bake it into your runbook so the 3 a.m. operator does not have to remember.
  • Track Squid-specific scrub scheduling issues. Squid 19.2.x introduced scrub scheduling changes that can leave deep scrubs taking far longer than expected, extending the window for undetected corruption. If scrubs stall after an upgrade, check whether osd_scrub_disable_reservation_queuing needs to be set to true.

How Netdata helps

  • The Ceph collector surfaces ceph_pg_inconsistent, ceph_pg_failed_repair, and ceph_pool_objects_repaired per pool, so divergence is visible the moment scrub reports it rather than the next time someone runs ceph health detail.
  • ceph_health_detail with labels for name (OSD_SCRUB_ERRORS, PG_DAMAGED, PG_NOT_DEEP_SCRUBBED) lets you alert on the specific integrity check, not just the umbrella HEALTH_WARN.
  • Per-OSD commit and apply latency correlate with the bad replica. When scrub flags a PG, latency outliers on the OSDs in its acting set point at the failing device.
  • Per-second collection shortens the window between “scrub found something” and “operator saw it”, which matters when the only good replica is one OSD failure away from disappearing.
  • Correlating scrub health with host-level SMART and memory signals on the same view turns “PG inconsistent” into “OSD 47 has new reallocated sectors and elevated commit latency” in one place, rather than across three dashboards.