The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-ha-host-isolation

Operations Guides

vSphere HA host isolation and split-brain: when isolation response goes wrong

vCenter shows one or more ESXi hosts as “Not Responding.” VMs on those hosts may have been restarted on surviving hosts by HA. Or they may still be running on the unreachable host with no way to manage them. In the worst case, the same VM is now running in two places at once, and nobody noticed until storage corruption or application errors surfaced.

vSphere HA uses two independent heartbeat mechanisms, network and datastore, to decide whether a host is alive. When those signals disagree, the isolation response policy determines what happens next. If the policy and the actual host state diverge, you get split-brain: two running copies of the same VM. VMFS file locking is the only guardrail against concurrent writes. On NFS, that guardrail is advisory at best.

This article covers how to tell isolation from actual host failure, what each isolation response option does, how to detect and recover from split-brain, and how to prevent the configuration mistakes that cause it.

What this means

HA classifies hosts into three states during a network event. The classification depends on which heartbeat mechanisms succeed.

flowchart TD
    A[Host loses mgmt network] --> B{Datastore heartbeat alive?}
    B -- yes --> C[Host ISOLATED]
    B -- no --> D[Host declared DEAD]
    C --> E{Isolation response?}
    E -- Leave powered on --> F[VMs stay running, unmanageable]
    E -- Shut down or Power off --> G[VMs stopped, HA restarts]
    D --> H{Host truly down?}
    H -- Yes --> I[Clean failover]
    H -- No --> J[SPLIT-BRAIN: dual VM instances]

Isolated means the host has lost network heartbeats to HA peers but is still exchanging datastore heartbeats. The host is running, its VMs are running, but it cannot be reached over the management network. HA knows the host is probably alive but cannot manage it.

Dead means both network and datastore heartbeats have stopped. HA assumes the host has failed and proceeds to restart its VMs on surviving hosts.

The isolation response policy applies only to hosts classified as isolated, not dead. The three options in vSphere 7.0 and 8.0:

  • Leave powered on: VMs on the isolated host keep running. HA does not attempt to restart them elsewhere. No duplicate VMs are created. The tradeoff is that VMs on the isolated host are unmanageable until the network is restored, and if the host is actually hung rather than network-isolated, those VMs are stuck with no recovery path.

  • Shut down: HA attempts a graceful guest shutdown of VMs on the isolated host, then restarts them on surviving hosts. The advanced option das.isolationshutdowntimeout (default 300 seconds) controls how long HA waits before forcing power off.

  • Power off: HA forcibly powers off VMs on the isolated host, then restarts them on surviving hosts. Faster than shut down but risks guest filesystem corruption.

Split-brain happens when HA declares a host dead (both heartbeats lost) and restarts its VMs, but the host is still running. This occurs when datastore connectivity is lost alongside network connectivity, or when heartbeat datastores are not properly configured. VMFS file locking prevents the restarted VM from powering on if the original instance still holds the lock on its .vmx file. NFS does not enforce equivalent mandatory locking, so both instances can write to the same VMDK simultaneously.

Common causes

CauseWhat it looks likeFirst thing to check
Management network switch failureMultiple hosts go Not Responding simultaneouslyPhysical switch health and uplink status
Management VLAN misconfigurationHosts isolated after VLAN or trunk reconfigurationVLAN tagging on management vmkernel adapter
Single management NIC, no redundancySingle host isolated, one vmnic link downesxcli network nic list for link state
Firewall blocking HA heartbeatsHosts appear isolated after firewall changeHA inter-host communication ports (TCP and UDP 8182 for agent-to-agent traffic)
NFS storage with “Leave powered on” responseVMs running on isolated host and restarted elsewhereStorage type and VMFS lock state
Missing or misconfigured heartbeat datastoresHost declared dead instead of isolatedHA datastore heartbeat configuration

Quick checks

# Check HA cluster configuration and failover level (PowerCLI)
Get-Cluster | Select Name, HAEnabled, HAAdmissionControlEnabled, @{N='CurrentFailover';E={$_.ExtensionData.Summary.CurrentFailoverLevel}}

# Check host connection states (PowerCLI)
Get-VMHost | Select Name, ConnectionState

# Check FDM (HA agent) status on a specific ESXi host (ESXi SSH)
esxcli software vib list | grep -i fdm  # verify vmware-fdm VIB is installed
ps -T | grep fdm  # verify the fdm process is running

# Search HA agent log for isolation or partition events (ESXi SSH)
grep -i "isolation\|partition\|split" /var/log/fdm.log | tail -50

# Check VMFS lock holder on a VM during suspected split-brain (ESXi SSH)
# Run on the .vmx file itself, not the lock file
vmkfstools -D /vmfs/volumes/<datastore>/<vm-folder>/<vm>.vmx

# List VM power state from ESXi local shell (ESXi SSH)
vim-cmd vmsvc/getallvms
vim-cmd vmsvc/power.getstate <vmid>

# Check physical NIC link state and error counters (ESXi SSH)
esxcli network nic list
esxcli network nic stats get -n vmnic0

vmkfstools -D and vim-cmd vmsvc/power.getstate are read-only and safe. Do not run vim-cmd vmsvc/power.off until you have confirmed which host holds the VMFS lock.

How to diagnose it

  1. Verify actual host state via BMC before trusting vCenter. Use the out-of-band controller (iLO, iDRAC, IMM, XCC) to check whether the isolated host is powered on and responsive at the hardware level. If the BMC shows the host running and healthy, the problem is network isolation, not host failure. Acting on vCenter’s view without BMC verification leads to wrong decisions.

  2. Distinguish isolation from failure using datastore heartbeats. If the HA heartbeat datastores show recent timestamp updates from the affected host, the host is isolated, not dead. If heartbeat files have stopped updating, the host may have genuinely failed or lost storage connectivity.

  3. Check the FDM agent log on the affected host. /var/log/fdm.log records the HA agent’s view of events. Look for explicit isolation declarations, partition events, and the reasoning behind host state transitions. This log is authoritative for what HA decided and why.

  4. Identify the network root cause. Check physical switches, VLAN configuration, and vmnic health. A single uplink failure on a host with redundant management vmnics should not cause isolation. If it does, the redundancy is misconfigured or the failover policy is wrong.

  5. If HA already restarted VMs, check for split-brain. Compare the VM power state reported by vCenter (VMs running on surviving hosts) against the actual power state on the isolated host, accessible via ESXi local shell through BMC console redirect or a separate management network. If VMs are powered on in both places, you have split-brain.

  6. Determine VMFS lock ownership. For each affected VM, check which host holds the lock on its configuration file using vmkfstools -D on the .vmx file. The host holding the lock is running the legitimate instance. The other host is running the stale duplicate.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ESXi host connection stateDistinguishes isolated from disconnected from downnotResponding for more than 10 minutes with other hosts connected
HA cluster health and host statusShows isolation and partition events directlyHost marked isolated or partitioned for more than 5 minutes
Datastore heartbeat freshnessProves host is alive via secondary pathHeartbeat file not updating, suggesting host is dead not isolated
Management network dropped packetsEarly warning of network degradationSustained non-zero drop rate on management vmnic
VM heartbeat statusIndicates guest OS responsivenessRed heartbeat on VM that vCenter shows as running
VMFS lock stateIdentifies lock holder during split-brain recoveryLock owner differs from where vCenter believes the VM runs

Fixes

Host is isolated but still running

If the BMC confirms the host is alive and the issue is network-only, fix the network before HA takes irreversible action. With isolation response set to “Leave powered on,” VMs continue running on the isolated host but are unmanageable through vCenter until connectivity is restored. No duplicate VMs, no unnecessary restarts, no data risk.

If the isolation response is “Shut down” or “Power off” and the network blip was transient, VMs may have been needlessly restarted. For most production environments, transient management network blips are more common than genuine host failures, making “Leave powered on” the better default.

HA already restarted VMs, potential split-brain

If HA declared the host dead and restarted its VMs on surviving hosts, but the host is actually still running, follow this recovery sequence:

  1. Do not reconnect the isolated host to the management network yet. Reconnecting while both VM instances are powered on can cause storage corruption, especially on NFS.

  2. Identify which instance to kill. On VMFS, use vmkfstools -D to find the lock holder. The host that does not hold the lock has the stale copy and should be powered off first. On NFS, there is no lock-based determination. Power off the instance on the host that was incorrectly declared dead (the one that lost network heartbeats).

  3. Power off the duplicate VMs on the original host. Access the host via BMC console redirect or SSH on a separate network path. Use vim-cmd vmsvc/power.off <vmid> to force power off each duplicate instance.

  4. Verify the surviving instances are healthy before reconnecting the original host to the management network. Check guest OS state, application health, and storage I/O.

  5. Reconnect the host and verify HA agent health. After reconnection, confirm the FDM agent returns to a healthy state and the host rejoins the HA cluster properly.

Adjusting isolation response configuration

Match the isolation response to your storage type and downtime tolerance:

  • VMFS with properly configured heartbeat datastores: “Leave powered on” is safe. VMFS file locking prevents split-brain even if HA wrongly declares the host dead and attempts restarts. The SCSI reservation on the .vmx file blocks the new instance from powering on.

  • NFS: “Leave powered on” is risky. NFS does not enforce the same mandatory file locking as VMFS. If HA declares the host dead and restarts VMs while the original host is still running, both instances can write to the same VMDK simultaneously. Consider “Shut down” for NFS-backed environments, accepting the restart downtime as the cost of split-brain prevention.

  • vSAN: In a vSAN cluster, the vSAN datastore is NOT used for HA heartbeat signals. If no non-vSAN shared datastores are available, HA has no datastore heartbeats, which means network isolation is harder to distinguish from host failure. If another non-vSAN datastore is available, it can be used for heartbeats. Consult VMware documentation for vSAN-specific HA configuration guidance.

Prevention

  • Redundant management network. Use at least two vmnics for the management vmkernel adapter, connected to separate physical switches configured with active/standby or LACP teaming. A single uplink failure should never cause isolation.

  • Configure multiple isolation addresses. By default, HA pings the management network default gateway to determine isolation. If the gateway is unreachable but the rest of the network is healthy, all hosts may declare themselves isolated simultaneously. Add peer host management IPs as additional isolation addresses using the advanced settings das.isolationaddress0 through das.isolationaddress9. If your default gateway is not a reliable reachability target, set das.usedefaultisolationaddress to false.

  • Configure heartbeat datastores. Ensure HA has at least two heartbeat datastores on separate storage arrays or LUNs. Without heartbeat datastores, any network loss looks like host death.

  • Match isolation response to storage type. Use “Leave powered on” for VMFS with proper heartbeat datastores. Consider “Shut down” for NFS-backed environments where split-brain risk is higher. Document the choice and the reasoning so future operators understand the tradeoff.

  • Test HA isolation behavior. Periodically simulate a management network failure in a non-production cluster by disconnecting the management uplink. Verify that isolation is detected correctly, the response matches expectations, and VMs behave as configured.

  • Keep HA admission control enabled. If admission control is disabled, the cluster may lack resources to restart VMs after a genuine failure. A cluster showing VMs as “protected” without admission control enabled is aspirational, not guaranteed.

How Netdata helps

  • Host connection state correlation at per-second resolution. Netdata surfaces individual host availability changes alongside HA events, letting you distinguish a single-host isolation event from a vCenter-side issue affecting all hosts at once. Whether all hosts went Not Responding simultaneously (vCenter problem) or sequentially (cascading network failure) changes the entire response.

  • Management network metrics before isolation triggers. Per-vmnic packet drops, error counters, and link state changes appear at second-level granularity. Sustained drops or error rate increases on the management vmnic precede isolation declarations. Correlating these with the actual HA event in the same timeline shortens root cause identification.

  • Datastore connectivity and latency signals. Storage path health and datastore latency help determine whether heartbeat datastores remain genuinely accessible during a network event, or whether the host is losing both heartbeat paths simultaneously. This distinction determines whether HA treats the host as isolated or dead.

  • VM power state and heartbeat tracking. Unexpected VM power-off events and heartbeat status changes, overlaid with host connection state transitions, reveal whether HA restarts are occurring and whether split-brain may be in progress. A VM showing red heartbeat on a host vCenter believes is disconnected is a strong split-brain indicator.

  • Composite event timeline. Overlaying host disconnection events, network error rates, storage path state changes, and VM restart operations on a single timeline turns a confusing “vCenter says hosts are down but VMs are still running” report into a diagnosable incident with a clear sequence of events.

The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.