The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-vcenter-service-dependency-deadlock

Operations Guides

vCenter service will not start: vmon dependency order and the max-restart wall

When vCenter services refuse to start, the failure is rarely the service you see stuck STOPPED. vmon (the VMware service lifecycle manager) starts each child service in dependency order and supervises it with a per-service restart policy. If a low-level dependency like vPostgres or STS is down or slow, every service layered on top of it fails its health check, retries within vmon’s bounded budget, and then silently stops trying.

The operator-facing symptom is a vCenter that looks half-alive: service-control --status --all shows a cascade of STOPPED services, the vSphere Client times out or shows a white screen, and vmware-vpxd is the one everyone focuses on even though it is almost never the root cause. The real failure is one layer down, in vPostgres, STS, or vmon itself.

A freshly restarted VCSA takes 5 to 15 minutes to bring all services up in order. Transient STOPPED states during that window are expected.

What this means

vmon is the systemd unit vmware-vmon.service that spawns and supervises roughly 30 interdependent vCenter services. It reads per-service JSON configuration that declares startup order, health checks, timeouts, and recovery actions. Two properties drive most of the behavior you will debug.

Dependency order. Services have a declared dependency chain. The canonical chain for the core management daemon is:

vmware-vpostgres -> lookupsvc -> vpxd-svcs -> vmware-vpxd

with rhttpproxy also required for external access. If vPostgres is not accepting connections, every service above it in the chain will fail its health check. vmon will not skip ahead. It waits, retries, and eventually marks each dependent as failed in turn.

The max-restart wall. Each service has a recovery action profile that tells vmon what to do on crash and on health-check failure. A common default profile performs two RESTART_SERVICE actions and then switches to NO_ACTION; verify the JSON for the service and build. When the budget is exhausted, vmon leaves the service STOPPED. This is silent: there is no event, no alarm, and no further automatic recovery until an operator intervenes or the VCSA is rebooted.

flowchart TD
    A["vPostgres STOPPED"] --> B["vpxd-svcs cannot reach DB"]
    B --> C["vpxd health check fails"]
    C --> D["vmon retries per service policy"]
    D --> E{"vpxd healthy?"}
    E -- no --> F["retry budget exhausted"]
    F --> G["vmon: NO_ACTION"]
    G --> H["vpxd STOPPED silently"]
    H --> I["SDK returns errors to clients"]

Two things make this pattern confusing in practice:

  • The service reported as failed is usually not the broken one. vpxd STOPPED almost always means vPostgres or STS is the real problem. Chasing vpxd first wastes the time you need to spend on the layer below.
  • Some STOPPED services are intentional. Auto Deploy (vmware-autodeploy) is correctly STOPPED if you do not use it. A VCHA passive node runs a deliberately reduced service set. Knowing your expected service profile prevents you from chasing services that are stopped by design.

Common causes

CauseWhat it looks likeFirst thing to check
vPostgres down or unhealthyvpxd fails to start, everything above the DB fails in cascadevmon-cli --status vmware-vpostgres, then /storage/db disk usage
STS signing cert expired or corruptSTS STOPPED, all auth fails, vSphere Client login brokenvecs-cli entry list --store STS_INTERNAL_SSL_CERT and checksts.py
Partition full (/storage/log, /storage/db, /storage/core)Services crash on log or DB write, restart loop, partition at 100%df -h and df -i
vmon hit its internal restart limitService stuck STOPPED with no recent restart attempts in vmon logsvmon-cli --status <svc> and vmon logs
vmon itself masked or start-limited by systemdNo vCenter services start at all, vmon unit inactivesystemctl status vmware-vmon.service
Duplicate JSON config files in svcCfgfilesvmon PANIC at boot, VERIFY bora/vim/apps/vMon/src/ServiceManager.cpp:108ls /etc/vmware/vmware-vmon/svcCfgfiles/ for .orig, -modified, or backup files
Startup profile wrong (VCHA orphan)Only a minimal service set starts, most stay STOPPEDcat /storage/vmware-vmon/defaultStartProfile; should be ALL on standalone
StartTimeout too low for slow I/OService fails health check during boot, may recover on manual startPer-service JSON StartTimeout value, vmon logs for timeout messages

Quick checks

Run these read-only before changing anything. They give you the dependency-layer view.

# Overall service state from vmon
/usr/lib/vmware-vmon/vmon-cli --list

# Status of the four core services
/usr/lib/vmware-vmon/vmon-cli --status vpxd
/usr/lib/vmware-vmon/vmon-cli --status vmware-vpostgres
/usr/lib/vmware-vmon/vmon-cli --status sts
/usr/lib/vmware-vmon/vmon-cli --status rhttpproxy

# Legacy wrapper; cross-check, may lag vmon state in newer versions
service-control --status --all

# Is the supervisor itself healthy?
systemctl status vmware-vmon.service
systemctl list-unit-files | grep vmware-vmon.service

# Disk space on the partitions that kill services when full
df -h
df -i

# Startup profile; ALL on standalone, reduced set on VCHA passive
cat /storage/vmware-vmon/defaultStartProfile

# Stray backup JSON files that crash vmon at parse time
ls -la /etc/vmware/vmware-vmon/svcCfgfiles/

Add a functional probe so you know what consumers actually see:

# Unauthenticated /sdk probe; proves rhttpproxy and vpxd are listening.
# Connection refused = rhttpproxy down; 503 = rhttpproxy up but vpxd not responding.
curl -sk -o /dev/null -w "%{http_code}\n" https://localhost/sdk

# Authenticated VAMI health endpoint (requires valid root credentials)
curl -sk -u 'root:<password>' https://localhost:5480/rest/appliance/health/system

How to diagnose it

  1. Confirm you are outside the boot window. If the VCSA came up less than 15 minutes ago, wait. Services start in dependency order and the full set takes 5 to 15 minutes. Paging on transient STOPPED states during boot creates noise and masks real failures.

  2. Find the lowest-layer STOPPED service. Start with vmon-cli --list and walk the dependency chain from the bottom. If vmware-vpostgres is STOPPED or unhealthy, that is your root cause and vpxd cannot start until it is fixed. If vPostgres is STARTED but vpxd is STOPPED, the failure is at the vpxd or STS layer.

  3. Check the supervisor before the supervised. If no services are starting at all, vmon itself is the problem. Run systemctl status vmware-vmon.service and look for result: start-limit (systemd gave up restarting vmon) or a masked unit. A masked vmon unit is a known state after image-based backup restores.

  4. Read the vmon log for the failed service. The per-service recovery action and the restart attempts are logged. Repeated RESTART_SERVICE entries followed by NO_ACTION means vmon has hit the max-restart wall and will not try again without intervention. Check /var/log/vmware/vmon/vmon.log and, on builds that emit it, vmon-syslog.log; message format varies by release.

  5. Check disk space per partition. Use df -h, not the VAMI UI, which rounds aggressively. The relevant partitions are /storage/log, /storage/db, /storage/core, and /storage/seat. A single partition at 100% with the others healthy is the normal failure shape.

  6. Check certificates if STS is involved. STS STOPPED with no disk pressure and no vPostgres issue is almost always a certificate problem. Check the STS certificate with Broadcom KB 318968 and inspect the STS_INTERNAL_SSL_CERT store with vecs-cli.

  7. Check the startup profile if only some services start. Run cat /storage/vmware-vmon/defaultStartProfile. On a standalone vCenter it must contain ALL. If it contains HACore left over from a destroyed VCHA configuration, vmon starts only the minimal HA set and leaves the rest STOPPED.

  8. Check for duplicate config files. Run ls /etc/vmware/vmware-vmon/svcCfgfiles/. Any file that is not a legitimate service config (.orig, -modified, .bak, editor swap files) can crash vmon at startup with a PANIC and a ServiceManager.cpp verify message.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
vmon per-service restart countvmon stops trying silently after the retry budget is exhaustedRestart count climbs then holds steady while the service stays STOPPED
Core service state (vpxd, vPostgres, STS, rhttpproxy)These four are the load-bearing servicesAny in STOPPED or FAILED outside the boot window
/storage/db, /storage/log, /storage/core utilizationFull partitions kill services and create restart loopsAny partition above 85%, or sudden growth rate
STS signing certificate validityExpired STS cert breaks all auth and cascades into service startup failuresLess than 30 days to expiry
SDK API response (authenticated probe)Ground truth for whether vCenter is actually workingNon-200 response or latency above 5 seconds
vmon unit state from systemdIf vmon is masked or start-limited, no services startsystemctl status showing inactive, masked, or start-limit
NTP offsetClock skew causes SAML token validation failures that look like STS problemsOffset above 5 seconds

Fixes

Recovery order matters. Fix the lowest-layer failure first, then let vmon bring the dependent services up in order. Restarting vpxd while vPostgres is still down just burns another restart attempt.

vPostgres down or unhealthy

Start here when vpxd will not start. Confirm vPostgres is the failure with vmon-cli --status vmware-vpostgres.

  • If /storage/db is full, free space before restarting. Truncate oversized logs with > /path/to/logfile (do not delete the file), check WAL accumulation, and review stats and event retention. Do not run VACUUM FULL under pressure without free space equivalent to the table size.
  • If vPostgres is STARTED but slow, check /storage/db I/O latency from the ESXi host perspective. A vPostgres health check that times out looks identical to vPostgres being down from vmon’s point of view.
  • On vCenter 8.0 U2, a corrupted postgresql.conf (file near zero bytes) is a known failure mode. The health check fails with Failed to read health xml file: /dev/shm/vmware-postgres-health-status.xml. Broadcom KB 424252 documents replacing the corrupted postgresql.conf from a healthy same-version vCenter, or using the matching PostgreSQL sample file, then restoring permissions and restarting services.

Once vPostgres is healthy, restart the dependent services in order, or reboot the VCSA if the cascade is wide.

STS down (certificate or otherwise)

STS STOPPED with vPostgres healthy is almost always certificate-related.

  • Run checksts.py to confirm which certificate is expired or corrupt.
  • For an expired STS signing certificate, use the current Broadcom certificate procedure — KB 318968 for STS checks and KB 385107 vCert for scripted replacement. This is a documented but delicate procedure. Do not improvise.
  • After renewal, services that cache the old certificate may need explicit restarts.

vmon hit the max-restart wall

When vmon has given up, the service stays STOPPED until you clear the state. Fix the underlying cause first, then restart the service or vmon:

# Start the specific service through vmon after fixing the root cause
service-control --start vmware-vpxd

# Or restart vmon itself to reset retry budgets across all services.
# Disruptive: briefly interrupts all managed services.
systemctl restart vmware-vmon.service

There is no --restart flag on service-control; you stop then start. Avoid repeated service-control --start calls on vCenter versions before 7.0 U3c because of a known file descriptor leak in vmon that eventually prevents new services from starting.

vmon itself masked or start-limited

# Check whether systemd masked the unit
systemctl list-unit-files | grep vmware-vmon.service

# Unmask if masked
systemctl unmask vmware-vmon.service
systemctl start vmware-vmon.service

# If systemd hit its own start-limit (distinct from vmon's per-service limit)
systemctl reset-failed vmware-vmon.service
systemctl start vmware-vmon.service

Wrong startup profile (VCHA orphan)

# Overwrites the startup profile. Standalone vCenter must start ALL services.
# Do NOT run this on a VCHA passive node; the reduced profile there is intentional.
echo -n ALL > /storage/vmware-vmon/defaultStartProfile
systemctl restart vmware-vmon.service

Duplicate JSON config files

# Inventory the config directory
ls -la /etc/vmware/vmware-vmon/svcCfgfiles/

# Move suspicious files out; do not delete until you are sure
mkdir -p /tmp/vmon-cfg-backup
mv /etc/vmware/vmware-vmon/svcCfgfiles/*.orig /tmp/vmon-cfg-backup/ 2>/dev/null
mv /etc/vmware/vmware-vmon/svcCfgfiles/*-modified.json /tmp/vmon-cfg-backup/ 2>/dev/null
systemctl restart vmware-vmon.service

StartTimeout too low

For services that consistently fail their health check during boot on slow I/O (notably statsmonitor), increase StartTimeout in the per-service JSON. The vpxd JSON documents a 300-second default in Broadcom KB 321336; other services vary, so read the specific JSON rather than assuming a common default. Editing these files is a VMware-support-guided change. Keep a backup and restart vmon afterward.

Prevention

  • Monitor the four core services by name, not just the overall health roll-up. vpxd, vPostgres, STS, and rhttpproxy each warrant individual alerting. A green overall roll-up can hide a single failed core service during the boot window.
  • Track vmon restart counts per service. A service that vmon has stopped retrying is invisible to anything that only checks whether the process is alive. Alert when the restart count stops climbing but the service remains STOPPED.
  • Alert on per-partition disk usage, not root. VCSA splits data across /storage/log, /storage/db, /storage/core, /storage/seat, and others. Root can be healthy while one storage partition is at 100%.
  • Track STS signing certificate expiry separately from machine SSL. The STS cert is invisible in a browser and is the one that takes down all authentication when it expires. Alert at 30 days minimum; 60 days is safer given the renewal procedure.
  • Suppress service-state alerts during the boot window. A VCSA that came up in the last 15 minutes is expected to have transient STOPPED services. Page only on core services outside that window.
  • Document the expected service profile for each VCSA role. Standalone, VCHA active, and VCHA passive have intentionally different service sets. Alerting on Auto Deploy STOPPED on a vCenter that does not use Auto Deploy is noise.
  • Verify vmon state after any image-based restore. A masked vmon unit after restore is a known failure mode. Run systemctl list-unit-files | grep vmware-vmon before declaring a restore complete.

How Netdata helps

  • Correlate vmon per-service state with /storage/* partition utilization in one view, so a vPostgres failure caused by a full /storage/db is obvious in seconds rather than inferred from separate dashboards.
  • Per-second metrics on VCSA CPU, memory, and disk let you see the boot window as it happens and distinguish a normal startup ramp from a service stuck in a restart loop.
  • Anomaly detection on vpxd log error rate and vmon restart count surfaces the max-restart wall before an operator notices the service has gone silent.
  • Certificate expiry tracking, including the STS signing certificate, gives weeks of lead time on the most common root cause of STS-down cascades.
  • NTP offset monitoring catches the clock skew that turns into SAML token validation failures, which otherwise look like STS problems.
  • Synthetic SDK probes with latency tracking give ground truth on whether vCenter is actually serving consumers, separate from whether individual services report STARTED.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.