The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-how-it-works-in-production

Operations Guides

How vSphere actually works in production: a mental model for operators

vSphere is a layered virtualization stack with three interdependent planes, and most production incidents cross plane boundaries. A guest that “feels slow” may be starved at the hypervisor scheduler, throttled by a forgotten CPU limit, fighting for IOPS at the storage layer, or sitting behind a vCenter whose database has bloated to the point that DRS stopped rebalancing the cluster.

The runbooks treat each failure mode individually. This article is the connective tissue: how to read a symptom in one plane and trace it to its origin in another.

The three planes are: the ESXi hypervisor plane (the VMkernel that owns the hardware), the VM plane (the sandbox and the VMware Tools channel that lets the hypervisor see into the guest), and the vCenter management plane (the orchestration layer that turns individual hosts into a cluster).

Why plane separation matters

The hypervisor plane is the only one with direct access to physical resources. CPU cycles, RAM, storage I/O, and network bandwidth are all arbitrated here by the VMkernel. The VM plane runs entirely inside the sandbox the hypervisor carves out and can only observe what the hypervisor chooses to expose. The management plane sits one layer further out and has no direct data-path role. When vCenter is down, running VMs keep running, DRS stops calculating, and HA still restarts VMs because the HA agents live on the ESXi hosts.

This separation is the source of both vSphere’s resilience and its debugging difficulty. Every metric you collect comes from a specific plane, and each plane has blind spots the others fill. Guest CPU utilization is reported by the guest and cannot see hypervisor descheduling. CPU ready time is reported by the hypervisor and cannot see guest-internal lock contention. vCenter task duration is reported by vpxd and tells you nothing about whether the host underneath is healthy.

How it works

flowchart TD
  subgraph Mgmt["vCenter management plane"]
    vpxd[vpxd inventory cache]
    vpg[vPostgres: inventory, stats, events, tasks]
    orch[DRS every 5 min, HA FDM agents, vMotion]
  end
  subgraph Hyp["ESXi hypervisor plane (VMkernel)"]
    sched[CPU scheduler: shares, limit, reservation]
    mem[Memory: TPS, balloon, compress, swap]
    stor[Storage I/O: pvscsi, VMkernel SCSI, VAAI]
    net[Network I/O: vmxnet3, vSwitch, NIOC]
    numa[NUMA scheduler: home node locality]
  end
  subgraph VM["VM plane (sandbox)"]
    tools[VMware Tools / open-vm-tools]
    balloon[vmmemctl balloon driver]
    hb[guest heartbeat]
  end
  vpxd -->|SOAP over HTTPS via hostd| Hyp
  orch -->|placement and migration decisions| Hyp
  VM -->|inflate, heartbeat, stats| Hyp
  Hyp -->|counters and connection state| vpxd

The ESXi hypervisor plane

The VMkernel is a purpose-built microkernel that owns the hardware. Five subsystems matter to operators.

CPU scheduler. Maps virtual CPUs (vCPUs) to physical CPUs (pCPUs). Each VM is a schedulable “world” using proportional-share allocation with shares, limits, and reservations. For multi-vCPU VMs the scheduler must co-schedule, finding N free pCPUs simultaneously, which makes large VMs harder to place. The world states that matter in esxtop are RUN (executing), READY (runnable but waiting for a pCPU), COSTOP (co-scheduling wait), and WAIT (idle or blocked on I/O). READY and COSTOP are the two invisible-from-the-guest taxes on every overcommitted host.

Memory management. ESXi overcommits memory through a four-tier reclamation hierarchy invoked in order of increasing desperation. Transparent Page Sharing (TPS) deduplicates identical pages and is largely disabled by default since vSphere 6.0 for security. Ballooning inflates the vmmemctl driver inside the guest, forcing the guest to page internally. Compression compresses 4KB pages in host memory before swapping to disk. Host-level swapping writes VM memory pages to .vswp files on the datastore. The tiers have radically different performance implications, from invisible to catastrophic, and the cascade can complete in minutes during a workload spike.

Storage I/O path. VM disk I/O traverses the guest OS, virtual SCSI adapter (pvscsi or LSI), VMkernel SCSI stack, storage driver, and physical storage. The VMkernel maintains I/O queues at each layer with configurable depth. VAAI (vStorage APIs for Array Integration) offloads operations like clone, zero, and hardware locking to the array. When VAAI is broken or unsupported, those operations fall back to the host and consume CPU and IOPS.

Network I/O path. VM network I/O flows through the guest OS, virtual NIC (vmxnet3), port group on a vSwitch or dvSwitch, and physical uplink (vmnic). The VMkernel handles its own traffic (vMotion, NFS, iSCSI, management, vSAN) on separate VMkernel adapters. Network I/O Control (NIOC) and traffic shaping throttle bandwidth per traffic class. Without NIOC, a single VM can starve others on a shared uplink.

NUMA. Modern multi-socket servers are non-uniform memory access architectures. Each CPU socket owns a region of local memory, and accessing remote memory costs roughly 1.5-2x latency. ESXi’s NUMA scheduler tries to keep each VM within a single NUMA node. VMs that span nodes (“wide VMs”) experience silent performance degradation, typically 10-30% memory throughput loss, that is invisible in any single counter except locality percentage.

The VM plane

Each VM runs in a sandbox with emulated hardware. VMware Tools (or open-vm-tools) inside the guest provides the balloon driver, the guest heartbeat, time synchronization, quiesced snapshots, and metrics reporting. Without VMware Tools, the hypervisor is partially blind to guest health: ballooning cannot run, so the host skips directly to compression and swapping, and HA VM Monitoring has no signal to act on.

A “Tools OK” status in the vSphere Client does not guarantee the balloon driver (vmmemctl) is actually running. Outdated VMware Tools can report healthy while the balloon driver is non-functional, which silently removes the gentlest tier of memory reclamation and forces the host into compression and swap earlier than the counters suggest.

The vCenter management plane

vCenter Server (typically deployed as VCSA, the vCenter Server Appliance running Photon OS) is a set of interdependent Java and C++ services orchestrated by vmware-vmon, sitting atop an embedded PostgreSQL instance (vPostgres), fronted by the rhttpproxy reverse proxy, and secured by the STS (Security Token Service) for token-based authentication. The exact number of services varies by vCenter version and enabled features; use service-control --status or vmon-cli -l to list them on a given appliance.

The core daemon is vpxd, a large multithreaded C++ process that maintains an in-memory cache of the entire managed inventory (hosts, VMs, networks, datastores, clusters, resource pools, permissions, alarms) and dispatches every management operation as a SOAP-over-HTTPS task routed to hostd on each ESXi host. vPostgres stores the persistent state: inventory, statistics, events, alarms, and tasks. If vpxd restarts, it must rebuild the inventory cache from the database, which takes minutes to tens of minutes in large environments.

Three orchestration features sit on top of vpxd and matter for incident response.

DRS (Distributed Resource Scheduler). Evaluates VM placement every 5 minutes by default, comparing imbalance improvement against vMotion cost. It considers CPU and memory, not storage. After a vCenter outage or a host reconnect, DRS evaluates all placements at once, which can produce a migration storm for 1-2 hours.

HA (High Availability). Monitors host liveness using network heartbeats and datastore heartbeats, with master-slave election among the ESXi hosts. The HA agents (FDM) live on the hosts, so HA continues to restart VMs even during a vCenter outage. The datastore heartbeat is what distinguishes “isolated” (network down, host still running, split-brain risk) from “dead” (both heartbeats lost, safe to restart elsewhere).

vMotion. Live-migrates VM memory and device state between hosts. Memory is pre-copied iteratively and the final switchover involves a brief stun, typically under a second but potentially seconds for memory-dirty workloads. A stuck vMotion can leave a VM in a degraded state between two hosts, and aggressive DRS levels can trigger migration storms.

Where it shows up in production

Symptoms surface in one plane and the cause lives in another.

  • Guest slowness with low guest CPU. Originates at the hypervisor scheduler. CPU ready time is the universal tax, but a forgotten CPU limit produces a different counter (cpu.maxlimited.summation) that ready time alone does not capture.
  • Memory pressure cascade. Ballooning (moderate) precedes compression (noticeable) precedes swapping (catastrophic). The guest’s own internal paging caused by ballooning is invisible to ESXi.
  • Storage latency cliff. Queue saturation or array-side slowdown pushes latency from under 5ms to 50-500ms. This affects every VM on a datastore simultaneously, unlike CPU or memory contention which is per-host.
  • NUMA penalty. Large VMs spanning NUMA nodes silently lose memory throughput. Often discovered only after a hot-add of CPU or memory breaks alignment.
  • Snapshot accumulation. Delta VMDKs grow until the datastore fills. The guest is completely unaware. Consolidation then stuns the VM, sometimes for hours.
  • vCenter database bloat. Statistics, events, and tasks accumulate in vPostgres. vpxd slows, DRS skips cycles, and eventually vCenter becomes unresponsive while running VMs are unaffected.
  • HA isolation or split-brain. A host loses the management network but VMs keep running. The isolation response policy determines whether VMs are shut down, restarted elsewhere, or left running and at risk of duplicates.
  • Certificate expiration. STS or machine SSL certificates expire on fixed dates and cause cascading TLS failures. The STS signing certificate is the most devastating and is not visible in a browser.

Tradeoffs and common misuses

  • Monitoring guest CPU instead of hypervisor CPU ready. A guest reporting 30% CPU with 15% ready time is running at roughly 85% of requested speed and the OS does not know it.
  • Treating host memory as a percentage. A host at 85% consumed memory might be fine (no balloon) or in crisis (actively swapping). The percentage alone tells you nothing. Monitor reclamation indicators independently.
  • Assuming vCenter being up means the cluster is healthy. vpxd can be running while DRS has stopped calculating because vPostgres is too slow to complete a rollup window.
  • Sizing VMs for “more power.” A VM with more vCPUs than its workload uses is harder to co-schedule, accumulates ready time and co-stop, and runs slower. Right-sizing beats adding vCPUs.
  • Forgetting that VCSA is a VM. If vCenter sits on an overcommitted host, hypervisor contention inside vCenter looks like vCenter slowness. A VCSA VM with low internal CPU but high ready time is severely degraded, and guest-only monitoring will miss it entirely.
  • Ignoring the STS signing certificate. It has a different lifecycle from the machine SSL certificate, it is not visible in a browser, and it is the single most common cause of total management plane outages.

Signals to watch in production

SignalWhy it mattersWarning sign
CPU ready time per VM (cpu.ready.summation)The invisible-from-guest scheduling tax. The single most important CPU signal.Sustained >5% per vCPU; >10% with host CPU >85% is an incident.
CPU co-stop per VM (cpu.costop.summation)Scheduling overhead for multi-vCPU VMs. The smoking gun for vCPU oversizing.Sustained >3% on any multi-vCPU VM.
CPU max-limited (cpu.maxlimited.summation)Captures forgotten CPU limits that ready time misses.Any sustained non-zero value where the VM owner reports slowness.
NUMA locality (N%L in esxtop memory view)Remote memory access costs 1.5-2x latency. N%L is the esxtop field for percentage of memory accessed from the NUMA home node; NLMEM and NRMEM show local and remote memory in MB.Below 80% for any workload; below 70% for latency-sensitive workloads.
Memory balloon (mem.vmmemctl.average)First tier of active reclamation.Sustained ballooning >5% of VM configured memory.
VMkernel swap rate (mem.swapinRate.average)Last-resort reclamation. Catastrophic.Any sustained swap-in >0 for >60 seconds is an emergency.
Datastore latency DAVG / KAVG / GAVGLocates latency in the I/O path.GAVG >30ms with QUED >0 sustained. DAVG >10ms on all-flash.
Outstanding I/Os (QUED)Confirms queue saturation vs transient load.QUED >0 sustained for >30 seconds.
Datastore free spaceA full VMFS datastore halts every VM on it.<15% free on any VMFS datastore.
Snapshot age and chain depthForgotten snapshots fill datastores and add read latency.Any snapshot >72 hours or chain depth >3.
ESXi host connection stateDisconnected means unmanageable; notResponding triggers HA.notResponding >10 minutes with another host connected.
vCenter service health (vpxd, vPostgres, STS)Management plane liveness.vpxd or vPostgres STOPPED after uptime >600s.
Certificate validity (especially STS)Expiry causes cascading TLS failures.Any certificate expired; STS within 30 days.
vPostgres partition usage (/storage/db, /storage/seat)Database bloat slows vpxd and eventually crashes it.>70% on /storage/db or /storage/seat.
Authentication failure rateBrute force or broken service account rotation.>10 failures from a single IP in an hour.

How Netdata helps

The value is correlation across planes. Deploy Netdata agents inside performance-critical guests and on the VCSA VM to collect per-second data that makes inter-plane failures visible without stitching together multiple 5-minute dashboards.

  • Per-second collection on CPU, memory, and disk I/O captures the reclamation cascade (balloon, compression, swap) and storage latency spikes that rollups average away.
  • Anomaly detection flags gradual degradation (rising memory pressure, creeping latency) before it becomes an incident.
  • Process and filesystem monitoring on VCSA tracks vpxd and vPostgres resource usage and catches /storage/db or /storage/seat filling before they take down vpxd.
  • Layered dashboards let you pivot from a guest symptom to the underlying host and datastore signals in one place.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.