The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-vcenter-cannot-login

Operations Guides

vCenter 'Cannot complete login due to an incorrect user name or password': SSO failures

The “Cannot complete login due to an incorrect user name or password” string is the exact message operators see in the vSphere Client, in PowerCLI sessions, and in API responses when SSO authentication fails. The text is misleading: the cause is rarely a typo. For a single user it is usually a credential or permission problem. For every account at once it is an SSO/STS infrastructure failure.

The first triage question is scope: does the local SSO administrator account (administrator@vsphere.local) still work? If yes, the STS signing certificate and token service are healthy, and the problem is in an identity source (AD/LDAP) or a service account. If administrator@vsphere.local also fails, the STS infrastructure itself is broken: expired STS signing certificate, clock skew rejecting SAML tokens, or STS memory pressure.

This page covers the infrastructure-failure case. The failure is loud in logs and silent in the UI, which shows only the generic login string. The diagnostic work happens in /var/log/vmware/sso/, the certificate stores, and the NTP configuration.

What this means

vCenter authentication is a chain. The browser or API client hands credentials to the reverse proxy (rhttpproxy), which forwards them to the Security Token Service (vmware-stsd, STS). For local SSO users, STS validates against vmdir (the embedded LDAP directory) and issues a SAML token. For AD/LDAP users, STS forwards the bind to the configured identity source. Every link in that chain has its own failure mode, and almost every failure surfaces to the user as the same generic login string.

The web client emits that string whenever the underlying STS call returns an authentication error. It tells you nothing about which link failed. The job is to find the failing link before users escalate.

Two failure shapes dominate. The first is total SSO failure: every account fails, including administrator@vsphere.local. The STS signing certificate has expired, the STS service is in a crash loop, or the appliance clock is far enough off that SAML token validation rejects everything. The vSphere Client may also show “503 Service Unavailable”, “no healthy upstream”, or “[500] An error occurred while fetching identity providers”. VAMI on port 5480 may report “certificate verify failed: certificate has expired”. The second shape is identity source failure: administrator@vsphere.local still works, but AD/LDAP users fail. Domain controllers are unreachable, clock skew exists between VCSA and DCs, the LDAPS certificate on the identity source was rotated without re-adding the source, or a service account password has expired and is flooding the logs.

A third, noisier variant is worth separating. A single expired or locked service account can generate thousands of LOGIN_FAILED entries per minute in vmware-sts-idmd.log. The UI works for humans, but the log volume is alarming and can fill /storage/log if left uncorrected.

flowchart TD
    A[Login fails for many users] --> B{administrator@vsphere.local works?}
    B -- No --> C[STS infrastructure broken]
    B -- Yes --> D[Identity source broken]
    C --> C1{STS cert expired?}
    C1 -- Yes --> C2[Run vCert Option 6]
    C1 -- No --> C3{Clock skew on tokens?}
    C3 -- Yes --> C4[Fix NTP / chrony]
    C3 -- No --> C5[STS heap or GC pressure]
    D --> D1{Single principal flooding?}
    D1 -- Yes --> D2[Service account locked or expired]
    D1 -- No --> D3{DC reachable from VCSA?}
    D3 -- No --> D4[DNS / firewall / DC down]
    D3 -- Yes --> D5{LDAPS cert rotated?}
    D5 -- Yes --> D6[Remove and re-add source]
    D5 -- No --> D7[Clock skew vs DC]

Common causes

CauseWhat it looks likeFirst thing to check
STS signing certificate expiredAll accounts fail, including administrator@vsphere.local. VAMI reports cert expired.VECS CLI store STS_INTERNAL_SSL_CERT dates
Clock skew vs AD domain controllersLocal admin works, AD users fail. Log shows LW_ERROR_CLOCK_SKEW.chronyc tracking on VCSA and a DC
AD/LDAP identity source unreachableLocal admin works, AD users fail. Log shows LDAP timeout or bind failure.ldapsearch against a DC from the VCSA
STS memory pressure or GC pausesIntermittent failures across all users. STS Java heap near max.jstat -gc on the STS PID
Expired or locked service accountUI works, but thousands of failures per minute from one principal.vmware-sts-idmd.log grouped by principal
Post-cert-renewal extension mismatchAfter cert replacement, extensions (EAM, RBD, ImageBuilder) fail to log in.updateExtensionCertInVC.py run per extension
ADFS password grant broken (8.0 U3h+)API logins for AD service accounts via ADFS fail with 400 Bad Request.ADFS server supports the password grant type
Windows Server 2025 LDAP signingPlain LDAP identity sources fail with “Strong(er) authentication required”.DC LDAP signing and channel binding policy

Quick checks

Run from the VCSA shell as root. These are read-only.

# Service health
service-control --status --all

# STS signing certificate dates - the cert that breaks everything when expired
# First list entries in the STS store to find the alias on your version:
/usr/lib/vmware-vmafd/bin/vecs-cli entry list --store STS_INTERNAL_SSL_CERT
# Then get the cert using the alias shown above:
/usr/lib/vmware-vmafd/bin/vecs-cli entry getcert --store STS_INTERNAL_SSL_CERT \
  --alias <alias_from_list_output> 2>/dev/null | openssl x509 -noout -dates

# Machine SSL cert - the one the browser sees, different lifecycle
echo | openssl s_client -connect localhost:443 2>/dev/null | openssl x509 -noout -dates

# NTP state - clock skew breaks SAML token validation
chronyc tracking

# Failed logins in the last hour, grouped to spot the noisy principal
grep -E "LOGIN_FAILED|Authentication.*failed" /var/log/vmware/sso/vmware-sts-idmd.log \
  | awk '{print $1, $2}' | sort | uniq -c | sort -rn | head -20

# Clock skew errors against AD
grep -i "LW_ERROR_CLOCK_SKEW\|clock skew" /var/log/vmware/sso/vmware-sts-idmd.log | tail -20

# LDAP bind errors pointing at identity source trouble
grep -i "ldap.*error\|ldap.*timeout\|bind.*fail" /var/log/vmware/sso/vmware-sts-idmd.log | tail -20

# Disk space on the log partition - SSO failure loops fill it fast
df -h /storage/log

# STS Java heap pressure - identify the STS JVM by matching its process name
# (VCSA runs multiple JVMs; match on the STS process, not a generic java lookup)
jstat -gc $(ps -eo pid,args | grep '[s]ts' | grep java | awk '{print $1}' | head -1) 2>/dev/null

How to diagnose it

  1. Confirm scope first. Try administrator@vsphere.local. If it works, skip to step 4. If it fails, STS infrastructure is broken and you are in the expired-cert or clock-skew branch.

  2. Check the STS signing certificate. This is the single most common cause of total login failure. On fresh installs of vCenter 7.0 U1 and later, the STS cert is valid for 10 years; older deployments and upgraded appliances may have shorter certificates, so check the date. The machine SSL cert in the browser is a different certificate and may be fine while the STS cert is expired. vCenter 7.0 U1 and later sends weekly notifications starting 90 days before STS cert expiry.

  3. Check the appliance clock. SAML tokens carry time bounds. If the VCSA clock is more than a few minutes off the domain controllers, STS rejects tokens and AD binds fail with LW_ERROR_CLOCK_SKEW. Run chronyc tracking on the VCSA and compare against a DC. NTP misconfiguration between VCSA and DCs is the usual root cause.

  4. Check the identity source. If local admin works, test AD/LDAP directly from the VCSA shell with ldapsearch. A successful bind proves the network path; a failure isolates the problem to DNS, firewall, or the DC itself. If you recently rotated the SSL certificate on an LDAPS identity source, the source must be removed and re-added; vCenter does not pick up the new cert from a live update.

  5. Inspect the log for the specific principal. A single expired or locked service account can produce thousands of failures per minute. Group the LOGIN_FAILED lines by principal and by source IP. A single account generating hundreds of entries per minute is almost always a password rotation that did not propagate to every integration.

  6. Check STS heap and GC behavior. Intermittent failures across all users, with no certificate or NTP issue, often point at STS Java heap pressure. Frequent full GC pauses cause token validation timeouts that look random from the outside.

  7. Check for post-patch regressions. vCenter 8.0 U3h enforced ADFS policies (MFA, geofencing) on password-grant API logins that were previously bypassed, breaking AD service accounts using ADFS as an identity source. Windows Server 2025 enables LDAP server signing requirements and LDAP channel binding by default, which breaks plain LDAP identity sources with “Strong(er) authentication required”. If the failure started immediately after a patch or DC upgrade, treat this as the leading hypothesis.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
STS signing certificate days to expiryTotal login failure when it expires; renewal is complexLess than 30 days
Machine SSL certificate days to expiryUI and API TLS failuresLess than 14 days
NTP offset (VCSA vs authoritative source)SAML token and Kerberos rejectionMore than 5 seconds sustained
SSO authentication failure rateDistinguishes typo noise from infrastructure failureMore than 20% of attempts, or a sudden spike
Per-principal failure countCatches a single expired service account flooding logsOne principal over 100 per minute
STS Java heap utilizationGC pauses cause intermittent token validation timeoutsOver 85% of Xmx sustained
/storage/log free spaceSSO failure loops fill the partition and cascade into more service failuresUnder 40% free
vmdir replication state (ELM)Divergent SSO state between linked vCentersReplication lag greater than zero

Fixes

Expired STS signing certificate

This is the highest-impact fix and the one most likely to be needed outside business hours. The deprecated checkSTS.py script is no longer the supported path. Use the vCert script documented in VMware KB 385107 for all certificate replacement on vCenter 7.x, 8.x, and 9.x. For an already-expired STS cert, the documented recovery is vCert Option 6 (reset all certificates) run from an SSH session to the VCSA. This is disruptive. Coordinate with VMware support if you have not done it before.

After STS cert renewal, check solution user and extension certificates. Extensions such as EAM, RBD, and ImageBuilder can fail to log in with the same generic string after cert replacement until updateExtensionCertInVC.py is run for each one.

Clock skew

Fix NTP. On vSphere 7 and later the VCSA uses chrony. Confirm /etc/chrony.conf points at reachable, redundant sources, ideally the same sources the domain controllers use. Do not rely on VMware Tools time sync from the ESXi host for the VCSA; it can fight NTP. After correcting the configuration, allow chrony to slew the clock, or use makestep for an immediate correction if the offset is large.

Unreachable AD/LDAP identity source

Verify DNS resolution of the domain controllers from the VCSA, verify the firewall permits the relevant port (389 for LDAP, 636 for LDAPS), and verify the DCs themselves are healthy. For LDAPS, confirm the certificate on the DC is trusted by the VCSA. If you rotated the LDAPS certificate, remove and re-add the identity source; live updates are not sufficient.

If you are still using Integrated Windows Authentication (IWA, the “Join Domain” method), plan the migration now. IWA is deprecated as of vSphere 7.0 and is removed in vSphere 9.0. The VCSA must leave the AD domain before a 9.0 upgrade or the pre-upgrade check will block it. AD over LDAPS or Identity Federation (Okta, Entra ID, ADFS) is the supported replacement.

Windows Server 2025 LDAP signing

Server 2025 enables LDAP server signing requirements and LDAP channel binding by default. Plain LDAP identity sources fail with “Strong(er) authentication required”. Either relax those policies on the DC (a security tradeoff) or migrate the identity source to LDAPS.

ADFS password grant (8.0 U3h and later)

If API logins for AD service accounts started failing with 400 Bad Request after patching to 8.0 U3h, the ADFS server must support the password grant type. The security fix enforces ADFS policies that were previously bypassed, including MFA and geofencing. Either reconfigure ADFS or move the affected service accounts to a different identity source.

Expired or locked service account

Identify the principal from vmware-sts-idmd.log, correct the account state at the AD level (unlock, reset password, update expiry), and update the credential in every integration that uses it. Log volume drops within seconds of the account being restored.

STS memory pressure

If STS heap is the bottleneck, restarting the STS service via service-control can clear the immediate pressure but does not fix the underlying sizing problem. Engage VMware support before changing Java heap parameters; the appliance ships with tuned values per deployment size.

Prevention

  • Track every certificate, not just machine SSL. The STS signing certificate is invisible in the browser and has caused more total outages than any other cert.
  • Standardize on vCert, not checkSTS.py. The older script is deprecated across vCenter 7.x, 8.x, and 9.x.
  • Monitor NTP offset continuously. Five seconds is a minimum alert threshold; thirty seconds is the realistic danger zone for SAML token rejection.
  • Alert on SSO failure rate, not just absolute counts. A rate spike without a corresponding login-attempt spike points at infrastructure, not typos.
  • Watch per-principal failure counts. A single noisy service account is the most preventable log flood in vCenter.
  • Plan the IWA migration before vSphere 9.0. The pre-upgrade check will block you otherwise.
  • Monitor /storage/log headroom aggressively. SSO failure loops can take a 40%-full partition to 100% in hours.

How Netdata helps

  • Per-second metric collection on the VCSA VM surfaces CPU, memory, and disk pressure that precede STS degradation.
  • NTP offset and chrony source state are collected directly, so you can correlate a clock skew event with the exact minute SSO failures began.
  • Disk utilization per /storage/* partition is tracked independently, so a log bomb from an SSO failure loop shows up as a steep slope on /storage/log.
  • SSO log-derived counters such as LOGIN_FAILED rate and per-principal failure count can be piped into Netdata as custom metrics, turning the generic login error into a rate signal you can alert on.
  • Anomaly detection on STS Java heap and GC frequency catches the silent degradation pattern where intermittent token validation failures precede a total outage.
  • Correlation across the stack (VCSA guest metrics, ESXi host CPU ready and memory balloon on the VCSA VM, NTP, disk) collapses the “is it vCenter or the host it runs on” question into a single timeline.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.