The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-cpu-ready-time-high

Operations Guides

vSphere CPU ready time high (%RDY): VMs starved while the guest looks idle

A database VM takes twice as long to run its nightly batch. Application latency pings fire. You SSH into the guest, run top, and CPU utilization sits at 25%. Memory is fine. Disk I/O looks normal. Nothing inside the VM explains the slowdown.

This is the classic signature of CPU ready time in vSphere. The guest OS has no visibility into hypervisor scheduling decisions. When the ESXi CPU scheduler cannot find a free physical CPU for a runnable vCPU, the vCPU waits in the READY state. The guest never learns it was descheduled, so from inside the VM everything looks idle while the hypervisor sees a starved VM.

What this means

CPU ready time (%RDY in esxtop, cpu.ready.summation in vCenter) measures the time a vCPU was runnable (it had work to do) but waited for a physical CPU. It is the direct tax of CPU overcommitment.

The guest OS cannot see this. The guest CPU utilization counter measures work completed, not work demanded. A VM reporting 25% guest CPU with 15% ready time is running at roughly 85% of its requested speed, yet the guest reports a comfortable 25%.

Per-vCPU conversion (vCenter and PowerCLI)

The cpu.ready.summation counter is in milliseconds, summed across all vCPUs of the VM. Convert to a per-vCPU percentage:

per_vcpu_rdy_pct = (ready_ms / (interval_ms * vCPUs)) * 100

For the realtime interval (20 seconds = 20000 ms):

per_vcpu_rdy_pct = ready_ms / (20000 * vCPUs) * 100
             = ready_ms / (200 * vCPUs)

The interval changes with the statistics rollup level. 500 ms ready in a realtime (20s) interval is a different percentage than 500 ms in a 5-minute or daily rollup. Confirm which interval you are reading before applying the formula.

The aggregate trap

vCenter’s cpu.ready.summation (and the value returned by Get-Stat -Stat cpu.ready.summation) sums ready time across all vCPUs of the VM. A 4-vCPU VM reporting 8000 ms ready in a 20s interval is not 8000 / 20000 = 40% per vCPU; it is 8000 / (20000 * 4) = 10% per vCPU. Skipping this division is the most common operator mistake.

esxtop’s default per-world %RDY is already a per-vCPU percentage (one world per vCPU, vmmX:vmname). The trap applies to vCenter charts, PowerCLI Get-Stat, and any VM-grouped esxtop view that rolls worlds up into a single line. This aggregate behavior is consistent across ESXi 6.5 through 8.0 — the metric name, esxtop field, and vCenter counter are unchanged.

Per-vCPU thresholds

Per-vCPU %RDYSeverityAction
< 2%NormalRoutine operation. Target for latency-sensitive workloads (databases, VDI, real-time).
2 to 5%Early contentionInvestigate VM placement and sizing.
5 to 10%Performance impactLatency-sensitive workloads degraded. Right-size or rebalance.
> 10%Production incidentApplication-visible latency and timeouts likely. Combine with host CPU > 85% to confirm overcommitment.

False positives

  • Brief spikes during vMotion (1-2 minutes) are expected.
  • Mass VM boot events cause transient spikes that resolve in minutes.
  • Idle servers do not produce false positives because ready time is near-zero when nothing is queued.

Rolled-up data hides spikes

At 5-minute and 30-minute rollups, averaging masks short spikes. A VM at 0% ready for 4.5 minutes and 50% for 30 seconds reports roughly 5% at the 5-minute rollup. Realtime data (20s interval, retained for 1 hour in vCenter) is the only reliable source for spike analysis. After that window, the detail is gone.

flowchart TD
    A["High %RDY on VM"] --> B{"Host CPU utilization?"}
    B -->|"Above 85%"| C["Genuine overcommitment"]
    B -->|"60 to 80%"| D{"%CSTP above 3%?"}
    B -->|"Below 60%"| E{"%MLMTD non-zero?"}
    D -->|"Yes"| F["vCPU oversizing"]
    D -->|"No"| G["NUMA imbalance"]
    E -->|"Yes"| H["CPU limit throttling"]
    E -->|"No"| I["Scheduler anomaly or affinity"]
    C --> J["Add capacity or migrate VMs"]
    F --> K["Reduce vCPU count"]
    G --> L["Check NUMA locality"]
    H --> M["Remove CPU limit"]

Common causes

CauseWhat it looks likeFirst thing to check
Host CPU overcommitmentHigh %RDY on multiple VMs. Host CPU above 85%. All VMs on the host affected.esxtop PCPU %USED section
vCPU oversizingHigh %RDY on specific large VMs. Host CPU moderate (60-80%). %CSTP elevated on the same VMs.%CSTP column in esxtop
NUMA imbalanceHigh %RDY on VMs spanning NUMA nodes. One NUMA node saturated, the other idle.esxtop, press ’m’, N%L column
CPU limit misconfigurationVM is slow but host CPU is low and %RDY looks fine. %MLMTD is non-zero.%MLMTD column in esxtop
Co-scheduling overhead%RDY and %CSTP both elevated on multi-vCPU VMs. Worse on VMs with more vCPUs.%CSTP vs vCPU count

Quick checks

Run esxtop on the ESXi host (SSH or DCUI) or use PowerCLI against vCenter.

# CPU scheduler view. %RDY per world is already per-vCPU.
# Press 'c' for CPU view, then look at %RDY, %CSTP, %MLMTD per world.
esxtop

# Collect cpu.ready.summation via PowerCLI (realtime, 20s interval).
# This value is an aggregate in ms across all vCPUs. Divide by vCPU count.
Get-Stat -Entity (Get-VM "myvm") -Stat cpu.ready.summation -Realtime
# per-vCPU %RDY = value_ms / (20000 * vCPU_count) * 100

# Host CPU utilization. OverallCpuUsage is in MHz.
# Note: NumCpuCores counts physical cores; on HT-enabled hosts this
# understates logical capacity. Substitute NumCpuThreads where available.
Get-VMHost |
  Select Name, @{N='CpuPct';E={[math]::Round(
    $_.Summary.QuickStats.OverallCpuUsage /
    ($_.Summary.Hardware.CpuMhz * $_.Summary.Hardware.NumCpuCores) * 100, 2)}}

# NUMA locality. Press 'm' for memory, 'f' to toggle NUMA fields.
# N%L is NUMA locality. Below 80% means significant cross-node access.
esxtop

How to diagnose it

  1. Confirm the symptom is ready time, not in-guest CPU saturation. Check guest-internal CPU first. If the guest shows near 100% CPU, the problem is inside the guest. Ready time only applies when the guest looks idle or under-utilized while the application is slow.

  2. Convert the counter correctly. Take cpu.ready.summation in milliseconds and apply (ready_ms / (interval_ms * vCPUs)) * 100. If reading esxtop per-world %RDY, the value is already per-vCPU.

  3. Correlate with host CPU utilization. This correlation is the key diagnostic step.

    • High ready (> 5% per vCPU) plus high host CPU (> 85%): genuine overcommitment. The host is out of physical cycles.
    • High ready plus moderate host CPU (60-80%): vCPU oversizing or NUMA imbalance. The host has capacity but the scheduler cannot find enough simultaneously-free pCPUs for the specific VM.
    • High ready plus low host CPU (< 60%): scheduler anomaly, CPU affinity misconfiguration, or a hidden CPU limit.
  4. Check co-stop (%CSTP) for multi-vCPU VMs. Co-stop measures the time vCPUs wait for sibling vCPUs to be co-scheduled. %CSTP above 3% means the VM likely has more vCPUs than it needs. This is the definitive oversizing spiral signal: an admin sees a slow VM, adds vCPUs, and the problem worsens because the scheduler must find even more simultaneously-free pCPUs.

  5. Check for CPU limits (%MLMTD). A VM with a configured CPU limit shows low ready time but high max-limited time. The limit is an artificial MHz ceiling. Any non-zero %MLMTD on a slow VM means a configured limit is the culprit. Limits are inherited from templates and resource pools; the “unlimited” default is -1 in the API, not 0.

  6. Check NUMA locality. A VM whose vCPUs and memory span NUMA nodes suffers cross-node memory access latency and harder scheduling. Check N%L in esxtop memory view (press ’m’, then ‘f’ to enable NUMA fields). Locality below 80% for a latency-sensitive workload is a configuration problem. Hot-add of CPU or memory can break vNUMA alignment.

  7. Rule out transient causes. Check for recent vMotion migrations, mass boot events, or DRS storms after a host reconnect or vCenter restart. Sustained ready time for more than 5 minutes is not transient.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
CPU ready time (cpu.ready.summation)Primary contention signal. Invisible to the guest OS.Per-vCPU above 5% sustained, or above 2% for latency-sensitive workloads.
Co-stop (cpu.costop.summation)Co-scheduling overhead for multi-vCPU VMs. The oversizing indicator.Per-vCPU above 3% sustained. Should be near zero when sized correctly.
Max-limited (cpu.maxlimited.summation)Artificial CPU limit throttling. Invisible to ready time.Any non-zero value where the VM owner reports slowness.
Host CPU utilization (cpu.usage.average)Overall compute pressure. Distinguishes host-wide from VM-specific contention.Sustained above 85% combined with ready time above 5%.
NUMA locality (numa.local / numa.remote)Cross-node memory access penalty.Locality below 80% for latency-sensitive workloads.
Guest CPU utilizationIn-guest CPU usage. Low guest CPU plus high ready time is the invisible starvation signature.Guest below 30% while the application is slow.

Fixes

Host CPU overcommitment

When host CPU is above 85% and ready time is elevated across multiple VMs, the host is out of physical capacity.

  • Increase DRS aggressiveness or move it to fully automated to balance VMs across hosts with headroom.
  • vMotion the most impacted VMs to less-loaded hosts. Expect a brief stun (typically under 1 second, longer for memory-dirty workloads).
  • Add host capacity to the cluster or reduce VM density on the saturated host.
  • Find runaway VMs consuming disproportionate CPU.

vCPU oversizing

When host CPU is moderate (60-80%), co-stop is elevated, and large VMs show worse ready time than small VMs on the same host, the problem is VM size, not host capacity.

  • Check in-guest CPU utilization. A 16-vCPU VM at 20% guest CPU needs roughly 3-4 vCPUs.
  • Reduce vCPU count to match actual utilization. This requires a VM power cycle.
  • Fewer vCPUs means the scheduler needs fewer simultaneously-free pCPUs, directly reducing both ready time and co-stop.
  • Applications that spawn worker threads based on detected vCPU count may need reconfiguration after reduction.

NUMA imbalance

When a VM’s vCPU count exceeds the cores per NUMA node, it spans nodes by definition. This causes harder scheduling and remote memory access.

  • Size VMs to fit within a single NUMA node when possible.
  • Verify CPU Hot Add is not breaking vNUMA alignment.
  • Check Cores per Socket against physical topology.

Large database VMs that genuinely need more cores than one NUMA node provides will always span nodes. Accept the penalty or split the workload across smaller VMs.

CPU limit misconfiguration

When host CPU is low, ready time looks fine, but %MLMTD is non-zero.

  • Remove the CPU limit (set to Unlimited) in VM settings.
  • Check resource pool limits that cascade to child VMs.
  • Check VM templates for inherited limits that propagate at deployment time.

Removing limits on a shared host increases noisy-neighbor risk. Use shares instead of limits for proportional priority control.

Prevention

  • Monitor per-vCPU ready time, not the aggregate. Alert on the converted per-vCPU value, not the raw summation counter or a VM-grouped esxtop line.
  • Right-size at provisioning. Start with fewer vCPUs and add only when in-guest utilization justifies it. Reducing vCPUs requires a power cycle; adding does not.
  • Monitor co-stop alongside ready time. Co-stop above 3% is the definitive oversizing signal. Catch it before someone adds more vCPUs to a slow VM.
  • Scan for CPU limits during VM audits. Limits inherit from templates and resource pools and are invisible to the guest. Check for non-zero %MLMTD across the inventory.
  • Keep latency-sensitive VMs on hosts below 60-70% CPU. The ready time curve is non-linear. Between 80-90% utilization, ready time climbs sharply.
  • Run DRS fully automated for clusters with variable workloads. Manual DRS lets imbalance persist until an operator intervenes.

How Netdata helps

  • Netdata’s vSphere collector can bring per-VM CPU scheduler counters into a single view, shortening the path from “this VM is slow” to the specific cause: overcommitment, oversizing, NUMA misalignment, or a hidden CPU limit. Check the collector documentation for the current list of exposed counters.
  • Applying per-vCPU conversion to collected ready time data eliminates the aggregate-misread path that triggers false escalations.
  • When the max-limited counter is collected, CPU limit throttling does not hide behind a healthy ready signal.
  • Anomaly baselines on ready time can surface slowly creeping contention before it trips a static threshold.
  • NUMA locality metrics, where exposed by the collector, complement CPU signals for large VMs where cross-node memory access is the silent throughput penalty.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.