The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-drs-not-balancing

Operations Guides

vSphere DRS not balancing: affinity rules and reservations blocking placement

DRS is enabled and Fully Automated, yet one host sits at 90% CPU while another idles at 30%. Vmotions are not happening, the recommendations queue is empty or full of unapplied entries, and the DRS score is poor. In most cases DRS is doing exactly what it was told: it cannot find a migration that satisfies every constraint.

DRS gates every placement decision behind a chain of compatibility and capacity checks before it compares hosts by load. A single VM-Host anti-affinity rule, a reservation that fully commits a host’s CPU or memory, or a missing vMotion network can each silently zero out the set of legal destinations. When every candidate fails one check, DRS emits no recommendation and the cluster drifts. Cluster averages hide this: a cluster “averaging 60% CPU” can mask one host at 90% and another at 30%. The per-host skew that DRS is supposed to fix is invisible in aggregate, and so is the constraint that prevents the fix.

What this means

DRS evaluates VM placement every 5 minutes by default and emits migration recommendations based on a cost-benefit tradeoff between imbalance improvement and vMotion cost. The migration threshold slider (levels 1 to 5, default 3) filters which priority recommendations are generated or applied. Before DRS compares hosts by load, it filters the destination set by compatibility. A host only becomes a candidate if:

  • It is connected and not in maintenance mode.
  • The VM’s storage is accessible from that host (shared datastore, accessible VMX and VMDK).
  • CPU compatibility (EVC baseline, vendor, feature set) permits the migration.
  • A vMotion network exists with adequate bandwidth.
  • All VM-VM and VM-Host affinity and anti-affinity rules are satisfied by the placement.
  • The host has enough unreserved CPU and memory to satisfy the VM’s reservation.
  • The host has enough free physical memory to receive the VM without breaching memory pressure targets.

If no host outside the current one passes all of those checks, DRS emits nothing. The VM stays put, the imbalance persists, and the cluster’s DRS score stays poor. The symptom looks like “DRS is not running.” The reality is “DRS ran and found zero legal moves.”

flowchart TD
    A[DRS scan every 5 min] --> B{Candidate hosts pass compat checks?}
    B -- No --> C[No recommendation emitted]
    B -- Yes --> D{Migration priority >= threshold?}
    D -- No --> E[Recommendation suppressed]
    D -- Yes --> F{Automation level?}
    F -- Manual --> G[Recommendation queued, not applied]
    F -- Fully Automated --> H[vMotion initiated]
    C --> I[Imbalance persists, DRS score stays poor]
    E --> I
    G --> I

Common causes

CauseWhat it looks likeFirst thing to check
VM-Host affinity or anti-affinity ruleDRS never moves a specific VM, or a group of VMs is pinned to one hostCluster > Configure > VM/Host Rules and VM/Host Groups
VM-VM affinity or anti-affinity ruleVMs that should separate are stuck together, or VMs that should co-locate cannot migrate independentlySame rules page; check “Must” vs “Should” (mandatory vs preferential)
Reservation fully commits host capacity“No other hosts have sufficient CPU or memory resources” fault in DRS recommendation; one host is the only legal landing zoneCluster > Configure > Resource Allocation
Memory overcommit on target hostsHosts look balanced by CPU but DRS will not move a memory-heavy VM because no target has free physical RAMesxtop memory view (MCTLSZ, SWCUR) on candidate hosts
DRS in manual modeRecommendations accumulate in the cluster Monitor tab but no vMotion runsCluster > Configure > vSphere DRS > Automation Level
Per-VM DRS overrideCluster is Fully Automated but specific VMs never moveVM > Configure > vSphere DRS > Automation Level
Migration threshold too conservativeCluster is roughly balanced but small improvements never happen; slider at 1 or 2Cluster > Configure > vSphere DRS > Migration Threshold
vMotion prerequisites unmetNo DRS-triggered vMotions happen at all; manual vMotion failsvMotion VMkernel adapter enabled, shared storage accessible, EVC baseline satisfied
Compatibility check failure“Current host is incompatible” or transient incompatibility blocks migrationCompatCheckTransientFailureTimeSeconds advanced option (see below)
Host HA agent errorsAffected host shows agent unreachable, isolated, or partitioned; DRS will not place VMs on itCluster > Monitor > vSphere HA; host connection state

Quick checks

# Cluster DRS configuration and current recommendations
Get-Cluster | Select Name, DrsEnabled, DrsAutomationLevel
Get-DrsRecommendation -Cluster <cluster>

# Recent DRS migrations in the last hour
Get-VIEvent -Entity (Get-Cluster <cluster>) -Start (Get-Date).AddHours(-1) |
  Where-Object { $_ -is [VMware.Vim.DrsVmMigratedEvent] }

# Host connection and maintenance state
Get-VMHost | Select Name, ConnectionState,
  @{N='InMaintenance';E={$_.ExtensionData.Runtime.InMaintenanceMode}}

# Failed vMotion events in the last 24 hours
Get-VIEvent -Entity (Get-Cluster <cluster>) -Start (Get-Date).AddHours(-24) |
  Where-Object { $_ -is [VMware.Vim.VmFailedMigrateEvent] }

Note: Get-VIEvent retrieves all matching events client-side and can be slow on large vCenter inventories. Scope it to a cluster with -Entity and a time window.

# On the busy host: confirm CPU contention (source of the problem)
# On candidate destination hosts: check for memory reclamation (why DRS rejects them)
esxtop
# Press 'c' for CPU view: %RDY, %CSTP, %MLMTD per VM world
# Press 'm' for memory view: MCTLSZ (balloon target), SWCUR (swap current)

For per-VM DRS overrides and reservation headroom, use the UI paths in the table above. Per-VM overrides are easy to miss because the cluster-level setting looks correct.

How to diagnose it

  1. Confirm DRS is actually enabled and automated. Open Cluster > Configure > vSphere DRS. Confirm DRS is enabled, the Automation Level is Fully Automated (or Partially Automated), and the Migration Threshold is at 3 or higher. A cluster in Manual mode still computes recommendations but never applies them. This is the single most common cause.
  2. Check the recommendations queue. Cluster > Monitor > vSphere DRS > Recommendations. If the queue is empty, DRS found no legal moves. If it is full and nothing is running, the cluster is in Manual mode. Each recommendation carries a fault reason when inspected through Get-DrsRecommendation.
  3. Inventory affinity and anti-affinity rules. Cluster > Configure > VM/Host Rules. For each rule, note whether it is mandatory (“Must”) or preferential (“Should”). Mandatory rules are hard gates. A “Must run on hosts in group X” rule with only one host in group X pins those VMs regardless of load. A VM-VM anti-affinity rule with more VMs than hosts is mathematically unsatisfiable.
  4. Check per-VM DRS overrides. A VM set to Manual or Disabled at the VM level overrides the cluster Fully Automated setting. Open VM > Configure > vSphere DRS for suspect VMs. Common sources: a VM pinned during a previous incident, or a template-derived override that was never cleared.
  5. Compare reservation capacity to host capacity. Cluster > Configure > Resource Allocation shows total CPU (MHz) and memory reserved capacity for the cluster. If total reserved CPU or memory approaches total cluster capacity, every host is effectively fully committed and no migration will find unreserved headroom. Reservations that fully commit the capacity of the cluster or of individual hosts can prevent DRS from migrating VMs between hosts.
  6. Check memory headroom on candidate hosts. DRS will not move a VM to a host that lacks the free physical memory to receive it, regardless of CPU balance. Run esxtop on candidate destination hosts and check MCTLSZ and SWCUR at the host level. If candidates are ballooning or swapping, they are not legal destinations for a memory-heavy VM.
  7. Verify vMotion prerequisites. A manual vMotion of any VM in the cluster should succeed. If it fails, the error identifies the blocker: no vMotion VMkernel adapter, no shared storage accessibility, EVC mismatch, or CPU vendor mismatch. DRS uses the same vMotion path, so a manual failure explains DRS silence.
  8. Check host connection and HA agent state. DRS will not place VMs on disconnected, notResponding, or maintenance hosts, nor on hosts with HA agent errors (agent unreachable, isolated, partitioned). Cluster > Monitor > vSphere HA lists host health. HA does not consult DRS affinity rules during power-on, so VMs placed by HA may later violate rules and DRS will migrate them after the fact.
  9. Check CompatCheckTransientFailureTimeSeconds. This advanced option controls whether DRS will move a VM off a host it is transiently incompatible with. In vSphere 7.0 U1 the default was 600 seconds (DRS tolerates transient incompatibility for up to 10 minutes before migrating). In vCenter 8.0 U3 the default changed to -1, which means DRS will not move a VM out due to incompatibility with its current host at all. If you are on 8.0 U3 or later and see “current host incompatible” faults blocking legitimate rebalancing, set the value to 600 or another positive number via the cluster’s DRS advanced options (Cluster > Configure > vSphere DRS > Edit > Advanced Options) to restore the prior behavior.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
DRS balance scoreTracks cluster balance; sustained poor score means DRS cannot find legal movesScore stays poor despite Fully Automated mode
Per-host CPU utilizationCluster averages hide skew; one host at 90% while another idles at 30% is invisible in aggregatePer-host spread greater than 20 percentage points sustained
Per-host memory consumed and activeMemory headroom on candidate destinations gates DRS placement of memory-heavy VMsCandidate destinations near physical RAM
Memory balloon (mem.vmmemctl) and host swapBallooning or swap on a destination makes it a poor or illegal placement targetMCTLSZ greater than 0 or SWCUR greater than 0 on candidate hosts
CPU ready time (%RDY)Reveals VMs starved on saturated hosts while the cluster average looks fineSustained above 5% on VMs on the busiest host
vMotion failure eventsFailed migrations indicate infrastructure issues (CPU compatibility, vMotion network, storage accessibility)VmFailedMigrateEvent count rising
ESXi host connection stateDisconnected, notResponding, or maintenance hosts are excluded from DRS placementAny host not “connected” in a production cluster
HA cluster health and host agent stateHA agent errors generate compatibility failures that block DRSHost marked isolated, partitioned, or agent unreachable
DRS migration rateSustained high rate indicates thrashing; near-zero with poor balance indicates constraintsRate outside the expected band for the workload

Fixes

Affinity and anti-affinity rules. Review every rule and convert mandatory (“Must”) rules to preferential (“Should”) where the workload tolerates it. Preferential rules generate fault reasons instead of blocking placement. Check VM-Host group membership: a rule that pins VMs to a one-host group is a hard pin. Check VM-VM anti-affinity rules against host count: anti-affinity across N VMs requires at least N hosts. Removing or relaxing the rule is usually preferable to fighting the scheduler.

Reservations. Compare reserved CPU (MHz) and memory across the cluster against physical capacity. Reduce reservations on VMs that do not need hard guarantees and use shares instead, which respond to actual demand. If reservations are required for licensing or SLA reasons, size the cluster so total reservations leave real headroom on every host, not just in aggregate.

Memory overcommit on candidate hosts. If candidate destinations are ballooning or swapping, DRS will not place memory-heavy VMs on them. Either reduce memory pressure on those hosts (right-size VMs, reduce overcommit) or add memory. The memory reclamation cascade is documented in the vSphere monitoring playbook: ballooning is the first tier, followed by compression, then host-level swap to .vswp files.

DRS automation level. If the cluster is in Manual mode because someone was cautious during a change, set it back to Fully Automated. If recommendations have been accumulating, apply or dismiss them so the queue is not masking the real state.

Per-VM DRS overrides. Find VMs with Manual or Disabled overrides and reset them to the cluster default unless there is a documented reason.

Migration threshold. If the slider is at 1 or 2, only the highest-priority recommendations are generated. Move to 3 (default) for general balancing. Going to 4 or 5 increases vMotion churn and is rarely justified.

vMotion prerequisites. Confirm each host has a VMkernel adapter tagged for vMotion, the vMotion network has adequate bandwidth (10GbE is the practical floor for production), all hosts see the VM’s datastores, and the EVC baseline matches across the cluster. A manual test vMotion should succeed before expecting DRS to use the path.

Compatibility transient option. On vSphere 8.0 U3 and later, if “current host incompatible” faults block migrations that previously worked, set CompatCheckTransientFailureTimeSeconds to a positive value (the prior default was 600) via the cluster’s advanced options. This restores the behavior where DRS tolerates transient incompatibility for a bounded window.

Prevention

  • Audit rules and reservations quarterly. Affinity rules and reservations accumulate from incidents and one-off changes. Review them against current host count and cluster capacity.
  • Monitor per-host, not cluster-average. Cluster CPU and memory averages hide the exact skew DRS is supposed to fix. Alert on per-host spread, not on the aggregate.
  • Track DRS score over time. A steadily declining DRS score with no topology change usually means a new constraint was added: a rule, a reservation, or a per-VM override.
  • Gate changes that add reservations or rules. Treat a new mandatory VM-Host rule or a large reservation like a capacity change. Verify DRS still has legal moves after it is applied.
  • Verify vMotion health continuously. A periodic test vMotion across host pairs catches EVC drift, network changes, and storage accessibility issues before DRS needs the path.

How Netdata helps

  • Netdata collects per-host CPU utilization, memory consumed, and memory active metrics at high resolution, making per-host skew visible without waiting for vCenter’s 5-minute rollups. The cluster average can look fine while one host saturates.
  • The vSphere integration surfaces CPU ready time, co-stop, and max-limited per VM, which reveal whether VMs on the busiest host are actually being starved even when cluster averages look acceptable.
  • Memory balloon and host swap metrics identify candidate destination hosts that DRS will reject for memory placement reasons, before you spend time wondering why a memory-heavy VM will not migrate.
  • DRS migration events and vMotion failure events correlate with host connection state and HA agent health, so a migration that stopped working can be traced to a host that became unreachable or partitioned.
  • Anomaly detection on per-host utilization spread flags the moment a cluster starts to skew, which is often the first visible symptom of a constraint blocking DRS.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.