The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-vcenter-vcha-replication

Operations Guides

vCenter HA (VCHA) replication broken: protection that is not protecting

VCHA makes vCenter look protected. The VAMI dashboard shows a green cluster, the active node serves traffic, and the passive node exists as a standby. When PostgreSQL streaming replication between the active and passive node stops, the passive node holds a stale database. A failover loses every transaction written since replication broke.

VCHA’s health surface is shallow. The VAMI summary can report healthy while pg_stat_replication shows NOT_REPLICATING. The passive VM is powered on. The witness is unreachable, so automated failover cannot reach quorum. None of this surfaces until an operator triggers failover and discovers the passive is hours behind, or until WAL accumulation on the active fills /storage/db and vpxd stops.

What this means

VCHA is an active/passive/witness topology for the vCenter Server Appliance. The active node runs all vCenter services and owns the management IP. The passive node is a hot standby that receives database changes through PostgreSQL streaming replication over a dedicated network. The witness is the third vote in a quorum that decides which node can become active during failover.

Healthy state:

  • pg_stat_replication on the active node shows state = streaming for the passive node’s IP
  • Replay lag is near zero (low single-digit MB at worst)
  • All three nodes are reachable across the VCHA network
  • Automated failover can reach quorum (2 of 3 nodes agree)

When any of these drift, VCHA stops protecting without making it obvious. The passive is powered on, the active is serving traffic, but a failover either does not happen (no quorum because the witness is down) or happens onto a stale database (replication broken).

Three operational constraints make silent VCHA breakage worse:

  • Failover is not instant. It restarts all vCenter services on the passive node. Expect 5-10+ minutes of management plane downtime, not seconds. ESXi-level HA continues independently of vCenter.
  • The management IP moves during failover. Client DNS caching extends apparent downtime beyond the failover window. Tools with long-lived connections to vCenter must reconnect.
  • Large file-backed artifacts such as local Content Library payloads or Update Manager patch repositories can create a failover gap if they are not shared or covered by the build-specific VCHA file-replication set. Verify their location and coverage for the deployed vCenter version.
flowchart TD
    ACT[Active VCSA
all services run here] -->|streaming replication| PAS[Passive VCSA
hot standby] ACT -.->|heartbeat| WIT[Witness
quorum vote] PAS -.->|heartbeat| WIT ACT -->|writes WAL| PG[(vPostgres
pg_wal)] PAS -->|replays WAL| PASDB[(vPostgres
database)] FAIL[Automated failover] -->|needs 2-of-3 quorum| WIT FAIL -->|restarts all services
5-10+ min| PAS IP[Management IP] -.->|moves during failover| DNSC[Client DNS cache
extends apparent downtime] WIT -.->|unreachable| NOQ[no quorum
no automated failover] ACT -.->|replication broken| STALE[passive DB is stale
failover loses data]

Common causes

CauseWhat it looks likeFirst thing to check
VCHA network degraded or partitionedpg_stat_replication lag climbs, state flips to NOT_REPLICATING; witness unreachable in VAMIip addr show on the dedicated VCHA NIC; ping between nodes
Passive node resource mismatch (vSphere 8.0)VCHA setup stuck at “PostgreSQL replication is not in progress”; replication never starts after cloneCompare vCPU and memory of active vs passive; check max_connections drift in postgresql.conf
WAL accumulation filling /storage/db/storage/db approaches 100%; vpxd stops; pg_wal directory is fulldu -sh /storage/db/vpostgres/pg_wal; check replication lag first
Witness node unreachableVAMI reports witness down; automated failover disabledCheck witness VM power state and VCHA network connectivity
Time drift between ESXi hosts hosting VCHA nodesRepeated failover every 5-10 minutestimedatectl status on each VCSA node; NTP on the ESXi hosts
Root password expired on passive or witnessFile replication fails, VCHA health degrades, vpxd may failCheck /etc/shadow expiry on passive and witness; verify patch level (fixed in 8.0u3e)
Snapshot taken on VCSA VMFailover under high load; failback from passive leaves VCHA inoperableGet-Snapshot on the vCenter VM; verify backup tool does not snapshot the VCSA

Quick checks

Run these on the active VCSA node. They are read-only and safe.

# Check VCHA cluster state and health. The shell command has varied by vCenter build; prefer the VCHA page in the vSphere Client
# or the vCenter Automation API operation com.vmware.vcenter.vcha.cluster.get.
vcha cluster get 2>/dev/null || /opt/vmware/share/vcha/vcha-util status 2>/dev/null || echo "use VAMI, vSphere Client VCHA page, or vCenter API"

# Check PostgreSQL replication state and lag (the ground truth)
/opt/vmware/vpostgres/current/bin/psql -U postgres -c "
SELECT client_addr, state, sync_state,
       pg_wal_lsn_diff(pg_current_wal_lsn(), replay_lsn) AS replay_lag_bytes,
       pg_wal_lsn_diff(pg_current_wal_lsn(), sent_lsn) AS sent_lag_bytes
FROM pg_stat_replication;"

# Check disk pressure on /storage/db (WAL accumulation indicator)
df -h /storage/db
du -sh /storage/db/vpostgres/pg_wal/

# Check VCHA network interface (dedicated NIC; name varies by Photon version)
ip addr show | grep -A2 "vcha\|eth1"

# Check VCHA logs for failover, heartbeat, or replication events
grep -i "failover\|replication\|heartbeat\|split.brain" /var/log/vmware/vcha/vcha.log | tail -50

# Check time synchronization
timedatectl status

When the psql -U postgres command prompts for a password, use the password in /etc/vmware-vpx/embedded_db.cfg on that node; local authentication can differ between VCSA builds.

How to diagnose it

  1. Verify replication state. On the active node, run the pg_stat_replication query above. If the result is empty or shows state other than streaming, replication is broken. This is the single most important check. A VAMI health summary that says “healthy” does not override a NOT_REPLICATING state here.

  2. Check the VCHA network. VCHA uses a dedicated NIC separate from the management network. Verify the interface is up, has the correct IP, and can reach both the passive and witness nodes. Packet loss, MTU mismatch, or bandwidth saturation on this network degrades replication silently before it breaks.

  3. Check replay lag in bytes. If replay_lag_bytes is growing, the passive cannot keep up with WAL replay. Sustained lag above 10MB warrants investigation. If it grows unboundedly, the active accumulates WAL that the passive has not consumed, and /storage/db on the active fills.

  4. Verify witness reachability. The witness is required for automated failover quorum. If the witness is unreachable, automated failover cannot happen regardless of replication health. Check the witness VM power state and network connectivity from both active and passive.

  5. Check /storage/db on the active node. Replication lag manifests as WAL accumulation. If /storage/db is filling and replication lag is high, the replication break is the root cause. Do not manually delete WAL files. Doing so breaks PostgreSQL crash recovery.

  6. Check time synchronization. Photon OS on the VCSA syncs time from the ESXi host. If NTP is misconfigured on the hosts, VCHA nodes drift and can trigger repeated failover every 5-10 minutes.

  7. Check root password expiry on passive and witness. In vSphere 8.0 versions before Update 3e, root password expiry on a passive or witness node caused file replication to fail, health to degrade, and vpxd to fail (PR 3456483). Verify the patch level and password expiry.

  8. Check for snapshots on the VCSA VM. Snapshots on a VCHA-enabled vCenter are not supported. They can trigger failover under high load and break VCHA operability on failback.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
pg_stat_replication.stateGround truth for replication healthAnything other than streaming
Replay lag (bytes)How far the passive is behind the activeSustained > 10MB or growing
/storage/db utilization on active nodeWAL accumulation from lag fills this partitionTrending above 70%; critical at 90%+
Witness node reachabilityRequired for automated failover quorumWitness down = no automated failover
VCHA network interface errors and dropsSilent replication degradation from packet lossAny sustained non-zero drop rate
Time sync (NTP offset) between VCHA nodesDrift causes spurious failover and cert issuesOffset > 5 seconds
VCSA service health on passive nodePassive must be ready to take overServices not running on passive
Root password expiry on all nodesExpired password breaks file replicationPassword within 7 days of expiry

Fixes

Replication broken from network partition

If pg_stat_replication shows NOT_REPLICATING and the VCHA network is the cause, fix the network first. Check the dedicated NIC, verify IPs, test connectivity between all three nodes. Do not attempt a manual failover while replication is broken unless the active node is permanently lost. If you must fail over onto a stale passive, treat it as a data-loss event and reconcile from file-based VAMI backups.

WAL accumulation filling /storage/db

If /storage/db is filling due to WAL from replication lag, the fix depends on the root cause. If replication is broken, fixing replication allows the passive to catch up and WAL to be consumed. If the passive cannot keep up due to a resource constraint, investigate the passive node’s resources. Do not manually delete WAL files, and do not turn off WAL archiving as an ad hoc VCHA workaround without Broadcom support guidance; PostgreSQL retains archived WAL when archiving is misconfigured.

Passive node resource mismatch (vSphere 8.0)

If VCHA setup stalls at “PostgreSQL replication is not in progress,” the passive node’s resources may not match the active. In vSphere 8.0, the max_connections setting in postgresql.conf is auto-configured based on total memory. If the passive has different memory or vCPU than the active, the active cannot open connections to the passive database and replication never starts. Do not change virtual hardware on VCHA nodes after cloning. They must be exact mirrors.

Witness node unreachable

If the witness is down, automated failover is disabled. The cluster still has a replication pair, but without the witness there is no quorum vote for automated takeover. Fix the witness: verify the VM is powered on, check the VCHA network, and re-establish connectivity. If the witness is permanently lost, VCHA must be destroyed and reconfigured.

Time drift causing repeated failover

If VCHA is failing over every 5-10 minutes, check NTP on the ESXi hosts hosting the VCHA nodes. Photon OS syncs time from the host, and drift between hosts causes the nodes to disagree on timing, triggering failover. Configure NTP correctly on all ESXi hosts before re-enabling automated failover.

Root password expired on passive or witness

If running a version before 8.0 Update 3e, root password expiry on passive or witness breaks file replication and can cause vpxd to fail. Update to 8.0u3e or later (PR 3456483) and reset the root password on all nodes.

Snapshots taken on the VCSA VM

Snapshots on a VCHA-enabled vCenter can cause failover under high load and leave VCHA inoperable after failback. Remove snapshots from the VCSA VM and reconfigure backup tooling to use file-based VAMI backups instead of VM-level snapshots.

Prevention

  • Monitor pg_stat_replication directly, not just VAMI health. VAMI can report healthy while replication is broken. The PostgreSQL view is the ground truth.
  • Alert on replay lag bytes, not just replication state. A state of streaming with growing lag is still a problem. Sustained lag above 10MB means the passive is falling behind and WAL is accumulating on the active.
  • Monitor /storage/db on the active node with VCHA context. WAL accumulation from replication lag is a VCHA-specific failure mode. A filling /storage/db on a VCHA active node should trigger a replication check immediately.
  • Monitor witness reachability independently. A healthy replication pair with a dead witness means no automated failover. This is a protection gap, not an active outage, but it should be a ticket-level alert.
  • Never snapshot a VCHA-enabled vCenter. Snapshots can cause failover under load and break VCHA on failback. Use file-based VAMI backups, which are the only VMware-supported backup method for VCSA.
  • Verify VCHA is supported in your topology. VCHA is not supported for vCenter Server instances managed by VMware Cloud Foundation (VCF). Network requirements include latency under 10ms and at least 1 Gbps between nodes, and multi-homed NICs are not supported.
  • Patch VCHA nodes together. Version drift between active, passive, and witness is not supported and causes replication failures.

How Netdata helps

  • PostgreSQL replication metrics surface pg_stat_replication state and WAL lag per-second, so a flip from streaming to NOT_REPLICATING is visible immediately rather than at the next manual check.
  • Per-partition disk utilization on /storage/db catches WAL accumulation before it fills the active node, with rate-of-change alerts that distinguish steady-state growth from replication-lag-driven spikes.
  • Network interface metrics on the dedicated VCHA NIC show packet drops, errors, and bandwidth saturation that silently degrade replication before it breaks entirely.
  • NTP offset and time sync metrics detect drift between VCHA nodes and their ESXi hosts, the precursor to repeated failover loops.
  • Correlating VCSA service health, disk pressure, and replication state in one timeline shortens diagnosis when VAMI reports healthy but the cluster is not actually protecting.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.