The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-too-many-pgs-per-osd

Operations Guides

Ceph too many PGs per OSD: mon_max_pg_per_osd, peering cost, and sizing

PG count per OSD is one of the few Ceph sizing decisions with direct cost in both directions. Too many PGs per OSD means more memory and CPU spent on peering, recovery scans, and PG log maintenance across more logical units. Too few means CRUSH cannot spread data evenly, producing hotspots on specific OSDs while others sit idle.

mon_max_pg_per_osd (default 250; it was 200 during most of Luminous and was raised to 250 in the 12.2.10 timeframe) is the cluster’s failsafe against runaway PG counts. When an OSD approaches this number, Ceph raises TOO_MANY_PGS and blocks new pool creation, pg_num increases, and replication factor changes. This is a guard rail, not a performance target. The existing cluster keeps serving I/O; only topology changes that would add PGs are blocked.

Each PG carries fixed overhead per OSD that hosts it: a PG log, peering state, recovery tracking structures, and BlueStore metadata. This overhead is roughly linear in the number of PGs the OSD participates in. An OSD with 250 PGs does more peering work, holds more PG logs in memory, and takes longer to recover than an OSD with 100 PGs holding the same amount of data.

There is also a hard limit: osd_max_pg_per_osd_hard_ratio (default 3.0). At 3.0 times mon_max_pg_per_osd, an individual OSD refuses to instantiate new PGs entirely. With defaults, that is 750 PGs per OSD. Reaching this limit means PGs cannot peer on that OSD, which can leave PGs stuck in creating or peering states.

The operational target is 100 to 200 PGs per OSD. The PG autoscaler (enabled by default on Nautilus and later) works toward this range, but its conservative threshold and the dynamics of OSD loss mean you can still trip the warning.

How PG count per OSD changes

CRUSH maps each PG to a set of OSDs based on the pool’s pg_num, the replication factor, and the available OSDs. The per-OSD PG count is not static. It changes when:

  • Pools are created or deleted. New pools add PGs distributed across OSDs by CRUSH.
  • pg_num is increased. Existing PGs split, increasing the count on every OSD in the acting sets.
  • OSDs are added or removed. CRUSH redistributes PGs. Adding OSDs lowers per-OSD counts; removing OSDs raises them.
  • The autoscaler splits or merges PGs. When the autoscaler decides a pool’s PG count is wrong for its data volume, it splits (too few PGs) or merges (too many PGs) automatically.

The PG autoscaler runs in the manager daemon and evaluates each pool against mon_target_pg_per_osd (default 100). It only acts when the target differs from the current pg_num by more than a factor of 3 (its configurable threshold, default 3.0). This conservative threshold means a pool can be significantly over- or under-provisioned before the autoscaler intervenes. When it does act, splitting PGs causes brief I/O stalls as the new PGs peer and activate.

New pools created with the autoscaler enabled start with a single PG by default. The autoscaler then evaluates the pool’s actual data volume and splits as needed. Pools created with the --bulk flag start with a full complement of PGs based on the target calculation and scale down only if usage is uneven.

flowchart TD
    TRIGGER["OSD loss, pool creation, autoscaler split"] --> CHECK{"PG count per OSD"}
    CHECK -->|under 50| FEW["hot spots, uneven I/O"]
    CHECK -->|100 to 200| TARGET["target range, balanced load"]
    CHECK -->|near 250| WARN["TOO_MANY_PGS, new pools blocked"]
    CHECK -->|near 750| HARD["hard limit, new PGs refused"]
    AUTO["PG autoscaler"] -.->|merges or splits toward target| TARGET

Where this shows up in production

OSD loss is the most common trigger. A cluster running at 220 PGs per OSD across 10 OSDs can jump to roughly 244 PGs per OSD on survivors after one OSD fails. The total PG count has not changed, but the denominator (number of OSDs) has shrunk. Lose two OSDs on a small cluster and you can cross the 250 threshold without creating any new pools.

Small clusters are disproportionately affected. With only 3 OSDs and a replication factor of 3, every PG lands on every OSD. A pool with 256 PGs means 256 PGs per OSD on all three. Add a second pool with the same count and you are at 512 PGs per OSD, well past the warning threshold. The autoscaler helps here by starting pools at 1 PG, but clusters provisioned before autoscaling was default, or pools created with explicit pg_num, can carry legacy counts that are too high for the cluster size.

Pool proliferation in multi-tenant clusters. CephFS deployments with separate metadata and data pools, RGW deployments with separate index, data, and log pools, and erasure-coded pools alongside replicated pools can accumulate dozens of pools. Each pool adds its PG count to the per-OSD total. Even if each pool is modestly sized, the aggregate can cross the threshold.

Autoscaler splits during load. When the autoscaler splits PGs, the new PGs must peer and activate. During this window, affected PGs are temporarily unavailable for I/O. On a busy cluster, a large autoscaler-triggered split can cause visible latency spikes. The autoscaler steps pgp_num gradually to amortize the cost, but expect brief remapping and backfill during splits.

Sizing: targets, tradeoffs, and calculation

The target range is 100 to 200 PGs per OSD. This balances two competing costs:

  • Below 100 PGs per OSD: CRUSH has fewer buckets to distribute objects across. With small pools, data concentrates on fewer OSDs. Individual OSDs become hotspots for specific workloads, and rebalancing after OSD loss is coarser because each PG move shifts more data.
  • Above 200 PGs per OSD: Per-PG overhead starts to dominate. Peering after topology changes takes longer because more PGs need to negotiate state. Memory consumption from PG logs and metadata increases. Recovery after OSD failure involves more PG-level coordination. Above approximately 500 PGs per OSD, peering and RAM usage become excessive.

The standard sizing formula for a single pool:

PGs = (target_per_OSD * num_OSDs * data_fraction) / pool_size

Round up to the nearest power of 2. target_per_OSD is typically 200 for stable clusters. data_fraction is the proportion of cluster data this pool will hold (0 to 1). pool_size is the replication factor or (k+m) for erasure-coded pools.

For example, a 3x replicated pool expected to hold 50% of cluster data on a 30-OSD cluster:

PGs = (200 * 30 * 0.5) / 3 = 1000 -> round up to 1024

The official documentation recommends a target of 200 PGs per OSD for all but the very smallest deployments, and warns that values above 500 produce excessive peering traffic and RAM usage. The autoscaler’s own target (mon_target_pg_per_osd) defaults to 100, which is more conservative than that recommendation. This difference is intentional: the autoscaler starts small and splits upward as data arrives, while a fixed target assumes you want a static count sized for expected growth.

Checking and adjusting

Check the current per-OSD PG distribution and autoscaler state:

# Per-OSD PG count (look at the PGS column)
ceph osd df

# Autoscaler status and recommended PG counts for all pools
ceph osd pool autoscale-status

# Current mon_max_pg_per_osd value
ceph config get mon mon_max_pg_per_osd

If the warning fires because of OSD loss (not because you created too many pools), the fix is to restore the missing OSDs or add new ones. Do not raise mon_max_pg_per_osd to suppress the warning. If you must create a pool or increase pg_num while the warning is active, you can set mon_max_pg_per_osd to 0 temporarily to disable the check. This removes the guard rail entirely and can allow PG counts that cause pathological peering and memory consumption. Use it only as a deliberate, short-term override during an active incident.

If a pool has too many PGs for its actual data volume, ensure the autoscaler is enabled so it can reduce pg_num through merging:

# Enable autoscaler for a pool
ceph osd pool set <pool> pg_autoscale_mode on

For pools with overlapping CRUSH roots (for example, a pool spanning both SSD and HDD device classes), the autoscaler refuses to scale and issues a warning in the manager log. These pools require manual PG count management.

Signals to watch

SignalWhy it mattersWarning sign
Per-OSD PG count (ceph osd df)Directly measures the resource this article is aboutAny OSD approaching 250, or the max-to-min ratio across OSDs exceeding 1.5x
ceph_health_detail{name="TOO_MANY_PGS"}Fires when an OSD crosses mon_max_pg_per_osdActive value of 1
ceph_health_detail{name="POOL_TOO_MANY_PGS"}Per-pool advisory when the autoscaler in warn mode thinks a pool is over-provisionedActive on pools with stale pg_num
ceph_pg_peering countPeering is the direct cost of high PG counts; more PGs means slower peering after topology changesPeering count not decreasing after OSD recovery
ceph_pg_creating countIndicates active PG splitting (autoscaler or manual)Unexpected nonzero values during production hours
OSD memory (RSS per daemon)PG logs and metadata consume RAM proportional to PG countRSS climbing without workload change
OSD count (ceph_osd_up, ceph_osd_in)Losing OSDs raises per-OSD PG count on survivorsOSDs transitioning down

How Netdata helps

  • Per-second PG state metrics from the Ceph collector let you watch peering and creating counts change in real time. When the autoscaler splits PGs, you see ceph_pg_creating spike and then settle, correlating with any brief I/O stalls.
  • Per-OSD metrics (up/down, latency, memory) let you detect when OSD loss is driving the per-OSD PG count upward. Correlating an OSD down event with a subsequent TOO_MANY_PGS health check clarifies cause and effect without manual ceph osd df polling.
  • Health detail metrics with labels surface TOO_MANY_PGS and POOL_TOO_MANY_PGS as individual signals, so you can distinguish the cluster-wide guard rail from per-pool autoscaler advisories.
  • Anomaly detection on PG state counts flags unexpected peering activity outside of known maintenance windows, which is often the first sign that the autoscaler is splitting PGs or that OSD loss has triggered redistribution.
  • Recovery rate metrics (ceph_pool_recovering_bytes_per_sec) alongside PG state counts help you assess whether high PG counts are slowing recovery after an OSD failure.