The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-mds-damaged

Operations Guides

Ceph MDS_DAMAGED: metadata journal or cache corruption

MDS_DAMAGED means the CephFS Metadata Server has found damaged metadata, either in its journal or in on-disk structures read from the metadata pool, and has deliberately refused to continue serving the rank. CephFS may be partially or fully unavailable, and recovery is not the usual automatic failover path. A standby that takes over would replay the same suspect journal, so the cluster parks the rank until a human intervenes.

Treat MDS_DAMAGED as a careful, manual recovery situation, not a restart-the-daemon situation. The wrong first move can convert recoverable journal corruption into permanent metadata loss. The playbook alert condition is ceph_health_detail{name="MDS_DAMAGE"} active for more than 120 seconds (the health check code is MDS_DAMAGE; MDS_DAMAGED is not a Ceph health check name).

This article covers how to read the damage, decide between journal recovery, scrub repair, and backup restore, and avoid the two failure modes that turn a damaged rank into a damaged filesystem: blindly resetting the journal, and running ceph pg repair on the metadata pool without identifying the authoritative copy.

What this means

The MDS keeps the CephFS namespace in an in-memory cache and writes a metadata journal to a RADOS pool. Each rank has a single journal replayed on failover. When the MDS encounters corrupt or missing metadata while reading from the metadata pool, it reports the MDS_DAMAGE health check. The monitor-generated cluster-level health message is “N mds daemon(s) damaged” (e.g. “1 mds daemon damaged”), with one detail line per damaged rank, and the daemon-reported check surfaces per-rank damage entries.

There are three distinct shapes of damage, each needing a different response:

flowchart TD
    A[MDS_DAMAGE active] --> B{damage ls shows what?}
    B -->|backtrace entry| C[Backtrace damage]
    B -->|journal read failure| D[Journal corruption]
    B -->|table object read failure| E[Bootstrap race or OSD op timeout]
    C --> F[Often benign: scrub with repair]
    D --> G[cephfs-journal-tool: inspect, export, recover_dentries]
    E --> H[Fix underlying cause, then ceph mds repaired]
    F --> I[ceph mds repaired]
    G --> I
    H --> I
    I --> J[Online MDS scrub: recursive, repair, force]
    J --> K[Rank returns to active]

Backtrace damage is a reverse-link inconsistency between an inode’s data objects and its path. It is frequently benign and repaired by an online MDS scrub. Journal corruption is more serious: the journal events themselves cannot be replayed cleanly, and you may lose metadata mutations that were journaled but not yet flushed to the backing store. Bootstrap-race damage happens when the MDS cannot read its table objects during startup because PGs in the metadata pool are not yet active, or because rados_osd_op_timeout is set aggressively enough to fail those reads.

MDS_DAMAGED is a refusal state, not a crash. The MDS could keep serving, but it has chosen not to, because continuing would risk spreading corruption to clients. Your job is to identify which kind of damage it is, repair or discard the affected metadata, then tell the monitor the rank is safe to start again with ceph mds repaired.

Common causes

CauseWhat it looks likeFirst thing to check
Backtrace damage (often root inode)damage ls shows damage_type: "backtrace", frequently ino: 1; cluster otherwise stableceph tell mds.<fs>:<rank> damage ls
Journal corruption from interrupted commitMDS refuses to start rank; journal inspect reports truncation or bad magic; recent MDS crash or host power eventcephfs-journal-tool --rank=<fs>:<rank> journal inspect
Bootstrap race or rados_osd_op_timeoutMDS goes damaged immediately after start; metadata pool PGs not yet active+clean, or rados_osd_op_timeout setceph config get mds rados_osd_op_timeout; PG state in metadata pool
Snaptrim under heavy loadDamage appears after snapshot trimming activity; recent snaptrim events in OSD logs`ceph osd dump
Hardware fault on metadata pool OSDSMART errors or inconsistent PGs in the metadata pool; OSD latency outliersceph pg dump filtering metadata pool PGs; SMART on OSDs hosting metadata pool

Quick checks

Run these read-only commands before touching anything. They tell you which kind of damage you are dealing with and whether the underlying metadata pool is healthy.

# Confirm the health check and the rank
ceph health detail | grep -iE 'mds.*damage|MDS_DAMAGE'

# List active MDS daemons and rank state
ceph fs status
ceph mds stat

# List damage entries from a still-running MDS admin socket
ceph tell mds.<fs_name>:<rank> damage ls

# Check the metadata pool PG states for inconsistencies
# (replace <metadata_pool_id> with your pool's numeric ID)
ceph pg dump | awk '$1 ~ /^<metadata_pool_id>\./ && $NF !~ /active\+clean/'

# Look for OSDs hosting the metadata pool that are down or degraded
ceph osd tree | grep -E 'down|host'
ceph osd dump | grep -E 'flags|nearfull|full'

# Check for an aggressive rados_osd_op_timeout that can cause spurious damage
ceph config get mds rados_osd_op_timeout 2>/dev/null || echo "not set"

# Check for snaptrim activity that correlates with the damage timestamp
ceph osd dump | grep -E 'nosnaptrim|nodeep-scrub'

# Look for OOM kills or disk errors on the MDS host around the damage event
dmesg -T | grep -iE 'oom|ceph-mds|I/O error' | tail -n 50

Do not run ceph pg repair against the metadata pool yet. On a metadata pool, the primary copy is not necessarily authoritative, and pg repair can overwrite good metadata with bad. Identify the cause first.

How to diagnose it

  1. Confirm which rank is damaged and the damage type. Start with ceph health detail and ceph fs status. The cluster-level message names the rank. If any MDS for that filesystem is still responsive, run ceph tell mds.<fs>:<rank> damage ls to enumerate per-rank damage entries. Each entry has a damage_type (backtrace, frag, dentry, inode, etc.), an ino, and an id you will need later for damage rm.

  2. Distinguish backtrace damage from journal corruption. Backtrace entries with ino: 1 (the root inode) are a known pattern in older CephFS versions where the root backtrace was never written. Journal corruption shows up differently: the MDS log contains lines like failed to read journal or bad journal magic, and cephfs-journal-tool journal inspect reports the journal as unreadable or truncated.

  3. Check whether the damage is real or a bootstrap artifact. If the MDS went damaged immediately after starting, look at the metadata pool PG states. If PGs were still peering when the MDS tried to read its table objects, the MDS may have failed reads and marked itself damaged. This matches tracker #51866 (mds daemon damaged after outage): a globally set rados_osd_op_timeout caused the MDS to fail table reads while the metadata pool PGs were still peering, and the damage resolved with ceph mds repaired <rank> once the PGs reached active+clean. The read-timeout code path is unchanged in current releases.

  4. Export the journal before doing anything destructive. Before journal reset, event recover_dentries, or any damage rm, back up the journal: cephfs-journal-tool --rank=<fs>:<rank> journal export <backup.bin>. This is the insurance policy. If a later step makes things worse, you can re-import the original journal.

  5. Check the underlying OSD and PG health of the metadata pool. If PGs in the metadata pool are inconsistent or down, fix that first. MDS damage on top of unhealthy RADOS is a much harder recovery.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_health_detail{name="MDS_DAMAGE"}Headline signal. Sustained active state means the rank will not self-heal.Active for more than 120 seconds.
ceph_health_detail{name="MDS_ALL_DOWN"}If all MDS daemons are down on top of damage, CephFS is fully unavailable.Any non-zero value.
ceph_health_detail{name="FS_DEGRADED"}Indicates failover did not produce a healthy active set.Active alongside MDS_DAMAGED.
ceph_mds_mem_rssMemory pressure that can precede an OOM kill and a damaged restart.Sustained growth toward mds_cache_memory_limit.
ceph_mds_slow_reply (counter)Slow MDS replies under load can mask developing corruption, especially during subtree rebalancing.increase(ceph_mds_slow_reply[5m]) > 0.
ceph_pg_inconsistent on the metadata poolScrub-found inconsistencies in the metadata pool are a direct precursor to MDS damage.Any non-zero value on the CephFS metadata pool.
ceph_osd_flag_nosnaptrimShould be set temporarily during recovery; if set long-term, snapshot trimming debt accumulates.Set for more than 24 hours without a recovery ticket.

Fixes

Recovery paths group by damage type. The wrong path for the wrong type makes things worse.

Backtrace damage (often benign)

If damage ls reports damage_type: "backtrace" and the inode is the root (ino: 1) or another inode whose data objects still exist, run an online MDS scrub with repair and force. The standard pattern from the disaster-recovery documentation:

# Bring the rank up enough to scrub. Mark the rank repaired so the
# monitor lets the MDS start, then scrub the affected subtree.
ceph mds repaired <role>
ceph tell mds.<fs_name>:0 scrub start / recursive,repair,force

# Once the scrub completes and reports the backtrace repaired,
# remove the damage entry so the health check clears.
ceph tell mds.<fs_name>:0 damage rm <id>

<role> is the filesystem role (for example 0 or <fs_name>:0). <id> is the damage entry ID from damage ls. Backtrace-only damage with intact data objects is the most recoverable MDS_DAMAGED scenario.

Journal corruption

When the journal itself is unreadable, you need cephfs-journal-tool. The workflow is: inspect, export, recover what you can, then reset.

# Assess journal health.
cephfs-journal-tool --rank=<fs>:<rank> journal inspect

# Back up the journal before any modification.
cephfs-journal-tool --rank=<fs>:<rank> journal export journal-backup.bin

# Recover dentries from the journal into the metadata pool.
# 'summary' shows what would be recovered; use 'apply' to write.
cephfs-journal-tool --rank=<fs>:<rank> event recover_dentries summary
cephfs-journal-tool --rank=<fs>:<rank> event recover_dentries apply

# WARNING: Only after recover_dentries. Reset is destructive: any
# journal events not recovered above are lost.
cephfs-journal-tool --rank=<fs>:<rank> journal reset --force

After journal reset, mark the rank repaired and start it.

Tradeoff: recover_dentries writes inodes and dentries from the journal into the backing store only if they are higher-versioned than what is already there. You will lose metadata mutations that were never journaled, and you may see orphaned inodes that need a later cephfs-data-scan pass to clean up. Run an online scrub after the rank is active to verify integrity.

Bootstrap race or rados_osd_op_timeout induced damage

If the damage appeared at MDS startup, with metadata pool PGs still peering, the journal is probably intact. The fix is to remove the cause, then clear the damaged flag:

# Remove the aggressive timeout if set.
ceph config rm mds rados_osd_op_timeout

# Wait until the metadata pool is active+clean.
ceph pg dump_stuck unclean

# Tell the monitor the rank is safe to start.
ceph mds repaired <role>

The same pattern applies when an MDS co-located with OSDs starts before the metadata pool is ready. The durable fix is startup ordering: delay MDS startup until PGs are active. There is no official systemd unit override for this; teams typically add an ordering dependency (unit After=/Wants=) or a wrapper that waits for the metadata pool before starting the MDS.

When to restore from backup instead

Prefer journal recovery when:

  • The damage is backtrace-only.
  • journal inspect reports the journal is mostly readable and only a small tail is corrupt.
  • The filesystem has few in-flight mutations since the last scrub.

Prefer a backup restore of the metadata pool (or a full filesystem restore) when:

  • journal inspect reports the journal as unreadable end-to-end.
  • The metadata pool PGs are also inconsistent or down.
  • recover_dentries recovers an implausibly small fraction of the expected inodes.
  • You have a known-good metadata pool snapshot more recent than the corruption event.

The alternate-pool recovery procedure (build a fresh metadata pool and reconstruct metadata from the data pool using cephfs-data-scan) is documented but explicitly marked as not extensively tested in the Reef documentation. Treat it as a last resort behind journal recovery and backup restore.

Prevention

  • Do not set rados_osd_op_timeout on MDS nodes. It can cause the MDS to fail reading its table objects during bootstrap and mark itself damaged even when the journal is fine. The default (no timeout) is correct for MDS.
  • Run MDS scrubs regularly. Backtrace damage accumulates quietly and is trivially fixed by a scrub with repair. A weekly or monthly online scrub of the root subtree catches it early.
  • Keep deep scrubs running on the metadata pool. noscrub and nodeep-scrub left set on the metadata pool is a common precursor to MDS damage going undetected until the MDS reads the bad object.
  • Size the MDS cache to the working set. A chronically over-limit MDS spends time evicting under cap pressure, which slows journal flushes and widens the window where an interrupted commit can corrupt the journal.
  • Pause snaptrim during heavy MDS load. Operators have reported journal corruption following snaptrim under load; ceph osd set nosnaptrim during known-heavy CephFS windows removes that variable.
  • Disable MDS subtree rebalancing during incidents. Subtree export under load can cause slow requests that mask underlying journal corruption. Setting mds_bal_interval to 0 during recovery simplifies the picture.
  • Back up the journal before any destructive action. This is the single highest-leverage habit. A journal export takes seconds and has saved recoveries that would otherwise have become restores.

How Netdata helps

  • The Netdata Ceph collector surfaces ceph_health_detail per check, so MDS_DAMAGED appears as its own labeled time series alongside MDS_ALL_DOWN, FS_DEGRADED, and the umbrella ceph_health_status. You see the exact second the rank went damaged and which other checks fired at the same time.
  • Per-second collection lets you correlate the MDS_DAMAGED transition with preceding signals: a spike in ceph_mds_mem_rss (OOM-kill precursor), an inconsistent PG appearing in the metadata pool, or an OSD going down on the host that holds the journal pool.
  • ceph_mds_slow_reply and per-operation MDS latency histograms let you confirm whether the damage was preceded by cap pressure or subtree rebalancing slowdowns, which changes the recovery plan.
  • Anomaly detection on MDS memory, request latency, and OSD commit latency can flag slow drift that often precedes an MDS crash-induced journal corruption, hours before the health check fires.
  • The metadata pool PG state metrics are tracked per pool, so you can see whether inconsistent or down PGs preceded the MDS damage, which determines whether you are dealing with real metadata corruption or a bootstrap-race artifact.
  • Annotations on the chart timeline let you mark the moment you ran cephfs-journal-tool journal export or ceph mds repaired, so the recovery is auditable after the fact.