The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-mds-cache-oversized

Operations Guides

Ceph MDS_CACHE_OVERSIZED: cache above limit and the cap recall that follows

Health-check names verified: ceph health detail reports these as MDS_CACHE_OVERSIZED and MDS_CLIENT_RECALL (with MDS_CLIENT_RECALL_MANY when many clients are affected); the MDS_HEALTH_* forms are the internal enum names in the Ceph source (mds_metric_name), not the displayed codes.

MDS_CACHE_OVERSIZED fires when an active Metadata Server daemon’s in-memory inode and capability cache crosses mds_cache_memory_limit * mds_health_cache_threshold. It is almost always workload-driven: a client has touched enough files in a short enough window that the MDS is caching more metadata than its budget allows.

The warning itself is benign for a few seconds. What matters is the cascade. To shrink its cache, the MDS revokes client capabilities (caps), one batch per session per second, capped by mds_recall_max_caps (default 30,000). If clients keep acquiring caps on new inodes faster than the MDS can recall them, you see a second warning (MDS_CLIENT_RECALL), the recall throttle counters climb, ceph_mds_mem_rss keeps rising, and at the end of the path sits client eviction.

The terminal state is not the cache warning. It is one or more CephFS clients blacklisted because they failed to return enough caps within session_timeout. From the client side that looks like I/O suddenly failing with EBLACKLISTED. From the operator side it looks like a “slow client” eviction log entry, hours after the original MDS_CACHE_OVERSIZED was dismissed as noise.

What this means

The MDS keeps the hot portion of the CephFS namespace in RAM: inodes, dentries, directory fragments, open file metadata, and the per-client capabilities that allow clients to cache reads and writes safely. Every inode cached on the MDS side typically has matching state on one or more clients, gated by caps. The cache is not an MDS-local structure; it is a distributed cache coordinated through cap messages.

mds_cache_memory_limit (default 4 GiB as of Reef/Squid) is a soft budget. The MDS does not hard-stop at it; it starts trimming. The actual recall trigger is mds_cache_reservation (default 0.05, meaning recall begins at 95% of the limit). When the cache crosses 95%, the MDS starts asking clients to return caps so it can evict the corresponding inodes. mds_health_cache_threshold (default 1.5) is only the warning threshold: if the cache still reaches 150% of the limit, the MDS_CACHE_OVERSIZED health check fires.

That gap, between 95% and 150%, is where most incidents actually live. By the time the warning fires, the recall machinery has already been running flat-out and losing. Two more signals usually accompany it:

  • MDS_CLIENT_RECALL active, meaning a client is failing to release caps at the rate the MDS wants.
  • ceph_mds_server_session_recall_throttle and ceph_mds_server_global_recall_throttle counters rising, meaning the MDS is bumping into its own recall throttles (mds_recall_max_caps per session, decay-limited globally).

If the gap is not closed, ceph_mds_server_cap_revoke_eviction ticks upward as the MDS evicts clients that exceed mds_max_caps_per_client (default 1M caps) or miss session_timeout (default 60s).

flowchart TD
    A["Client scans tree
find, rsync, du"] --> B["Cache crosses
95% of limit"] B --> C["MDS recalls caps
mds_recall_max_caps per session"] C --> D{"Recall rate beats
acquisition rate?"} D -- yes --> E["Cache shrinks
warning clears"] D -- no --> F["MDS_CACHE_OVERSIZED fires
at 150% of limit"] F --> G["MDS_CLIENT_RECALL fires
recall_throttle counters rise"] G --> H["Client crosses
mds_max_caps_per_client or
misses session_timeout"] H --> I["Client evicted
EBLACKLISTED on client"]

The fork at “Recall rate beats acquisition rate?” is the decision point. The rest of the article is about determining which side you are on and what to do about it.

Common causes

CauseWhat it looks likeFirst thing to check
Recursive scan workloadMDS_CACHE_OVERSIZED correlates with a find, rsync, du -sh, or backup job. One client session dominates ceph_mds_caps.ceph tell mds.<id> session ls for the session with the largest num_caps.
Genuine working set larger than cacheSteady-state ceph_mds_mem_rss near the limit even with no scans. MDS_CLIENT_RECALL flickers but does not escalate. Slow but sustained recall throttle rate.Compare ceph_mds_inodes_with_caps against mds_cache_memory_limit.
Slow or stuck clientOne or two sessions have many recalled_caps and old recall timestamps. Cache size flat; the MDS is waiting on them.ceph tell mds.<id> session ls for stale recall timestamps; client-side logs for cap return stalls.
Standby-replay false positiveWarning on a standby-replay MDS with “0 inodes in use by clients, 0 stray files” in the health message. No active sessions on that daemon.ceph fs status and the rank state of the warning daemon.
mds_cache_memory_limit misconfigured under RookMemory request set in the CephFilesystem CRD but no limit; effective cache limit stays at the 4 GiB default.ceph config get mds mds_cache_memory_limit on the MDS host.

Quick checks

Run these read-only on the cluster and on the active MDS host. None of them modify state.

Both forms are valid on current releases: ceph tell mds.<id> ... is routed through the monitors and works from any host (the daemon must be up); ceph daemon mds.<id> ... connects to the admin socket directly and therefore only works on the MDS host while the daemon is running. The mix below uses whichever form suits the context.

# Top-level: is the warning still active, and which daemon is firing it?
ceph health detail | grep -E 'MDS_CACHE_OVERSIZED|MDS_CLIENT_RECALL'

# Cache size and inode pressure for the MDS
ceph tell mds.<id> cache status
ceph tell mds.<id> perf dump | jq '.mds_mem'

# Live sessions, sorted by caps held (look for the dominant client)
ceph tell mds.<id> session ls | jq 'sort_by(-.num_caps) | .[] | {id, num_caps, recalled_caps}'

# Effective config on this daemon (compare against what you think you set)
ceph config get mds mds_cache_memory_limit
ceph config get mds mds_cache_reservation
ceph config get mds mds_recall_max_caps
ceph config get mds mds_max_caps_per_client

# Confirm the daemon rank: is this the active MDS, or standby-replay?
ceph fs status

# Process memory for ground truth (rss in KB)
ps -eo pid,rss,args | grep ceph-mds

If the warning daemon is standby-replay and ceph health detail reports zero inodes in use, skip to the “Standby-replay false positive” subsection. This is a documented false positive on Quincy (17.2.x) and earlier, tracked in Red Hat Bugzilla #1951348 and #1944148 (both open and describe exactly this standby-replay oversized-cache warning).

How to diagnose it

  1. Confirm the cache is actually oversized on an active MDS. Cross-check ceph tell mds.<id> cache status against mds_cache_memory_limit. If the daemon is standby-replay, treat it as the false-positive case regardless of the cache size value.
  2. Find the dominant session. Sort session ls by num_caps. In the scan-driven case, one or two client addresses account for the bulk of the caps. Note those IPs.
  3. Check whether recall is making progress. Pull ceph_mds_caps, ceph_mds_inodes_with_caps, ceph_mds_server_session_recall_throttle, and ceph_mds_server_global_recall_throttle over a few minutes. If ceph_mds_caps is flat or rising while the throttle counters rise, recall is losing.
  4. Look for the eviction precursor. Check increase(ceph_mds_server_cap_revoke_eviction[5m]). Any non-zero value means the MDS has already evicted at least one client; expect EBLACKLISTED errors on the client side and entries in ceph osd blocklist ls.
  5. Identify the workload. The dominant client IP from step 2 should map to a host you can inspect. The usual suspects are find /, recursive rsync, backup jobs, indexing services, and CI artifact pipelines. A short perf dump baseline before and after the job starts shows the cap acquisition rate.
  6. Rule out the misconfiguration cases. Verify mds_cache_memory_limit is what you expect, especially under Rook, where a pod memory request without a limit does not propagate to the MDS. The behavior is tracked as rook/rook#8143 (“MDS configuration mds_cache_memory_limit is not set as expected”).

Metrics and signals to monitor

The signal names below are Netdata collector metric names (out of scope for this review); the underlying Ceph mds_server perf counters they map to were verified in the MDS source: cap_revoke_eviction is a named counter, and per-session and global recall throttles exist in Server::recall_client_state.

SignalWhy it mattersWarning sign
ceph_health_detail{name="MDS_HEALTH_CACHE_OVERSIZED"}Cache above 150% of the configured limit. MDS is already recalling and losing.Active for more than a few minutes.
ceph_health_detail{name="MDS_CLIENT_RECALL"}A client is failing to return caps fast enough.Active at all. Correlates with the next eviction.
ceph_mds_mem_rssGround-truth memory consumption of the MDS daemon.Trending toward host RAM limit; OOM kill is the cliff.
ceph_mds_capsTotal caps granted to all clients. Proxy for the distributed cache size.Rising while MDS_CLIENT_RECALL is active.
ceph_mds_inodes_with_capsInodes pinned in the MDS cache by outstanding caps.Tracking toward the configured limit.
ceph_mds_server_session_recall_throttle (counter)Per-session recall batches limited by mds_recall_max_caps.Monotonic increase during the incident.
ceph_mds_server_global_recall_throttle (counter)Global recall rate-limiter hits.Monotonic increase during the incident.
ceph_mds_server_cap_revoke_eviction (counter)Clients evicted for not returning caps.Any non-zero value. Use increase() over a short window.
ceph_mds_reply_latency_sum / _countEnd-to-end MDS request latency.Rising while cache pressure rises; client-visible stall.
ceph_mds_slow_reply (counter)MDS replies breaching the slow-reply threshold.Any non-zero value over a 5-minute window.

Fixes

The fix is one of: raise the limit because the working set genuinely needs it, change the workload so it does not pin the whole namespace, or address the specific edge case (slow client, standby-replay false positive, Rook misconfiguration). Restarting the MDS is almost never the right first move. Note that in Squid (19.2.0+) ceph mds fail and ceph fs fail require the --yes-i-really-mean-it confirmation flag when the MDS is active with a MDS_TRIM or MDS_CACHE_OVERSIZED health warning.

Raise mds_cache_memory_limit for a genuine working set

If ceph_mds_inodes_with_caps is consistently near the limit outside of any scan, and MDS_CLIENT_RECALL is steady background noise rather than an escalation, the working set is larger than the budget. Cache-related options are runtime-updatable; no MDS restart is required.

# Inspect current value
ceph config get mds mds_cache_memory_limit

# Raise it (example: 8 GiB). Apply live.
ceph config set mds mds_cache_memory_limit 8589934592

Two cautions. First, the MDS host must have the RAM. ceph_mds_mem_rss should track the new limit within a few minutes; if it does not, the config did not apply (re-check with ceph config get mds ... on the host). Second, under Rook, set the value through the CephFilesystem CRD’s metadataServer.resources.limits.memory, not only via ceph config set, otherwise the next reconcile may reset it.

Throttle the scan, do not throttle the MDS

If a find, rsync, or backup job is the trigger, the right fix is on the workload side. The variable that matters is the metadata-op rate, not CPU or data-IO priority, so nice/ionice alone usually do not help; you need to slow the directory-walk rate itself.

  • Add a small sleep between directory descents in scan scripts (find ... -exec sh -c '...' \; with a delay, or a rate-limited walker). This is the only reliable way to drop the cap acquisition rate below the recall rate.
  • Replace rsync -a /src/ /dst/ over the whole tree with targeted syncs, or use --update plus a manifest to avoid re-scanning unchanged subtrees.
  • For backup systems that walk CephFS, prefer CephFS snapshots (ceph fs snap) as the backup source instead of a live walk. Snapshots do not pin caps the same way.
  • Schedule scans during low client-count windows so per-client cap pressure is bounded by mds_max_caps_per_client.

If you must tune the MDS to absorb a known one-time scan, raise mds_recall_max_caps (per-session recall batch size, default 30,000). This lets the MDS recall more caps per session per second, which can close the gap on scan-driven incidents. The tradeoff is more recall traffic on the MDS-to-client link.

# Increase per-session recall batch (example: 50,000)
ceph config set mds mds_recall_max_caps 50000

Tune in small steps and watch ceph_mds_server_session_recall_throttle fall while ceph_mds_caps falls.

There is no definitive upstream guidance on tuning mds_recall_max_decay_threshold (default 128K) for very large directory trees; it is a decay threshold on the per-session recall throttle that operator mailing lists do tune, so treat it as an experimental lever and validate on a test workload.

Address a slow or stuck client

If session ls shows one or two clients with large recalled_caps counts and old recall timestamps, the MDS is waiting on them. Investigate the client host:

  • Kernel CephFS client: check dmesg for cap-related stalls, look at client memory pressure, verify the client is not itself under OOM.
  • FUSE client: check the client process CPU and memory.
  • Network: cap recall messages share the same path as data. Packet loss or saturation between the client and the MDS will manifest as slow recall.

If the client is genuinely wedged, manual eviction is a last resort. The client will see EBLACKLISTED; clear the blocklist entry once the client has remounted and the workload is no longer stuck.

Verified syntax: ceph tell mds.<id> session evict <client-id> (the session evict admin-socket command takes a client identifier).

Standby-replay false positive

If ceph health detail shows MDS_CACHE_OVERSIZED on a standby-replay daemon with “0 inodes in use by clients, 0 stray files,” this is the documented false positive. The standby-replay daemon replays the active MDS’s journal and accumulates metadata without having any clients to recall caps from, so it cannot trim. Restarting the standby-replay MDS clears the warning with no client impact:

Verified unit names: under classic deployment this is `ceph-mds@<daemon-id>.service`; under cephadm it is `ceph-<fsid>@mds.<host>.<name>`. Adjust to your deployment.

# Safe: only the standby-replay daemon is affected. Active MDS continues serving.
systemctl restart ceph-mds@<standby-replay-daemon-id>

If this fires repeatedly, the underlying issue is tracked at tracker.ceph.com/issues/48673 (“High memory usage on standby replay MDS”); consider reducing journal size or filing a bug against your Ceph version.

Prevention

  • Size mds_cache_memory_limit for the working set, not the average. Measure ceph_mds_inodes_with_caps during peak load and set the limit with at least 30% headroom. The 4 GiB default assumes a modest namespace; large CephFS deployments routinely run at 16 to 32 GiB.
  • Alert on ceph_mds_mem_rss trending toward host RAM, not just on MDS_CACHE_OVERSIZED. The warning fires at 150% of the configured cache limit; OOM fires at host RAM. The gap between those two is your runway.
  • Track ceph_mds_server_cap_revoke_eviction as a counter and alert on any increase. Evictions are user-visible incidents; they should never be a surprise.
  • Gate backup and indexing jobs that walk CephFS. Treat any process that does find / or rsync / against CephFS the way you would treat a full table scan in a database: require an approval, schedule it, and rate-limit the walk.
  • If you run multi-active MDS (max_mds > 1), monitor per-rank cache pressure separately. Subtree partitioning does not guarantee balanced cap distribution; one rank can hit MDS_CACHE_OVERSIZED while others are idle.
  • Under Rook, set both requests and limits for the MDS pod. Verify the effective mds_cache_memory_limit after deploy, not just the CRD value.
  • Review recall throttle settings when the cluster’s client count or namespace grows. Defaults tuned for a small cluster will starve recall on a large one.

How Netdata helps

  • The ceph_health_detail metric with the name label lets you alert on MDS_HEALTH_CACHE_OVERSIZED and MDS_CLIENT_RECALL independently. Both firing on the same MDS rank is the strongest signal that the cascade has started.
  • ceph_mds_caps and ceph_mds_inodes_with_caps per MDS daemon show the distributed cache size directly. A rising curve against a flat mds_cache_memory_limit is the leading indicator, before the health check fires.
  • ceph_mds_server_session_recall_throttle and ceph_mds_server_global_recall_throttle counters, viewed as rates, show whether the recall machinery is saturated. Flat ceph_mds_caps with rising throttle counters is the “recall is losing” signature.
  • ceph_mds_server_cap_revoke_eviction viewed as increase(...[5m]) is the eviction alarm. Any non-zero value means clients are already seeing EBLACKLISTED.
  • Correlating ceph_mds_reply_latency and ceph_mds_slow_reply against the cache-pressure metrics separates an MDS that is merely hot from one that is stuck behind its own recall throttle.
  • ceph_mds_mem_rss next to the host’s available memory gives the OOM runway directly. Per-second resolution catches the sharp ramp that precedes an OOM kill, which the 150%-of-limit health check does not see.