The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-vcenter-ntp-clock-skew

Operations Guides

vCenter clock skew: the NTP offset that breaks tokens and disconnects hosts

Sudden SSO login failures. ESXi hosts flipping to “Not Responding” without a network cause. Certificate validation errors on certificates you know are valid. All three point to clock skew between vCenter, ESXi, and the identity infrastructure they depend on.

SAML token validation, certificate validation, Kerberos, HA heartbeats, and log correlation all assume clocks agree. When they diverge, the failures look like unrelated problems instead of one root cause. Nobody thinks to check the clock.

The vCenter Single Sign-On STS defines a hard clock tolerance of 10 minutes for SAML token validation (Broadcom vSphere SSO programming guide); symptoms below that limit are version- and service-dependent.

The thresholds are unforgiving. An offset greater than 30 seconds can cause intermittent SSO authentication failures. At 5 minutes, Kerberos breaks outright and ESXi hosts can disconnect from vCenter. Host-to-host skew above 5 minutes also blocks vMotion.: the vMotion compatibility alert is “Too large clock skew was detected” when host-to-host skew exceeds five minutes (Broadcom KB 397777).

What this means

Clock skew breaks vSphere through several interdependent mechanisms. The vCenter Single Sign-On Security Token Service (STS) validates SAML tokens within a clock tolerance window. When the VCSA clock drifts outside that window relative to the token issuer or the client requesting the token, validation fails. Kerberos has its own 5-minute tolerance by default. Certificate validation fails when a clock reads a time before the certificate’s “not before” date or after its “not after” date, producing “certificate not yet valid” errors on certificates that are fine.

ESXi hosts disconnect because their management agent communication with vCenter depends on timing-critical heartbeat and authentication exchanges. When the skew exceeds what the vpxa-to-vpxd channel can tolerate, the host appears “Not Responding” even though VMs on it continue running normally.

The failure cascade looks like this:

flowchart TD
  A[NTP offset grows] --> B{Magnitude}
  B -->|> 30s| C[Intermittent SSO token failures]
  B -->|> 5 min| D[Kerberos auth rejects]
  B -->|> 5 min host-to-host| E[vMotion compatibility alert]
  D --> F[ESXi hosts disconnect]
  C --> G[LOGIN_FAILED in STS logs]
  A --> H[Certificate not-yet-valid errors]

Common causes

CauseWhat it looks likeFirst thing to check
NTP source unreachableOffset grows steadily; chronyc sources shows no synced sourceFirewall on UDP 123, DNS resolution of NTP server names, NTP server availability
VMware Tools time sync fighting NTP on VCSAVCSA clock reverts to host time after each sync cycle; NTP appears configured but offset persistsVCSA VM settings: “Synchronize guest time with host” checkbox
Snapshot revert set the clock backOffset appears suddenly after a revert operation; clock reads the snapshot’s creation timeRecent snapshot revert events; compare date -u to known-good time
Circular NTP dependency on ADWhen AD DCs are unreachable, both auth and time break simultaneously; recovery requires fixing AD firstNTP server config: are all sources AD domain controllers?
NTP not configured at alltimedatectl shows NTP service inactive; offset grows without boundVAMI Time tab or /etc/chrony.conf (/etc/ntp.conf on older versions)

Quick checks

All commands below are read-only and safe on production systems.

# VCSA: overall time sync state
timedatectl status

# VCSA: NTP tracking (vSphere 7+ uses chrony)
chronyc tracking
chronyc sources -v

# VCSA: older versions (pre-7) use ntpd
ntpq -p

# VCSA: compare current time to known-good external time
date -u

# VCSA: NTP servers configured in chrony
grep ^server /etc/chrony.conf

# VCSA: check whether VMware Tools time sync is overriding NTP
vmware-toolbox-cmd timesync status

# VCSA: check NTP config via VAMI API (password will appear in shell history)
# The VAMI API authenticates with basic auth (root)
curl -sk -u 'root:' https://localhost:5480/rest/appliance/ntp

# ESXi host: system time and hardware clock
esxcli system time get
esxcli hardware clock get

# ESXi host: NTP peer status
ntpq -p

# VCSA: clock-skew-related SSO auth failures
grep -i "LOGIN_FAILED\|LW_ERROR_CLOCK_SKEW" /var/log/vmware/sso/vmware-sts-idmd.log | tail -20

# VCSA: certificate validation errors caused by time skew
grep -i "not.*yet.*valid\|certificate.*invalid" /var/log/vmware/vpxd/vpxd.log | tail -20

How to diagnose it

  1. Measure the offset on the VCSA. Run chronyc tracking (vSphere 7+) or ntpq -p (older). Look at the “System time” or “offset” field. Above 5 seconds, investigate. Above 30 seconds, you are likely seeing intermittent auth failures. Above 300 seconds, Kerberos is broken and hosts may be disconnecting.

  2. Check NTP source reachability. Run chronyc sources -v. A healthy config shows at least one source marked * (current sync source). If all sources show ? (unreachable) or the reach column is 0, check firewall rules for UDP 123, DNS resolution of NTP server hostnames, and connectivity to the NTP servers.

  3. Compare clocks across the stack. Check the VCSA clock (date -u), the ESXi host clock (esxcli system time get on each host), and your AD domain controller clocks. The skew that matters is between these components. A VCSA and its ESXi hosts could all agree on the wrong time and still function, as long as the skew between them stays within tolerance.

  4. Correlate with SSO failures. Check /var/log/vmware/sso/vmware-sts-idmd.log for LOGIN_FAILED entries and LW_ERROR_CLOCK_SKEW errors. If these spike when the offset grows, clock skew is the root cause, not a credential or certificate problem.

  5. Check for VMware Tools time sync conflict. Run vmware-toolbox-cmd timesync status inside the VCSA. If it reports enabled, the ESXi host clock periodically overrides the VCSA’s NTP-synced time. This is a common cause of persistent, unexplained drift.

  6. Look for recent snapshot reverts. A snapshot revert sets the guest clock back to the snapshot’s creation time. NTP then slews the clock gradually rather than stepping it, which can take minutes to hours for a large offset. Check vCenter Tasks for recent revert operations.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
NTP offset (VCSA and ESXi)Silent prerequisite for SAML, Kerberos, certificate validation> 30 seconds sustained warrants investigation; > 5 minutes is an emergency
NTP source reachabilityUnreachable sources mean the clock drifts without correctionNo source marked * in chronyc sources; reach column at 0
SSO authentication failure rateCorrelates with offset growth; confirms auth impactSpike in LOGIN_FAILED entries in vmware-sts-idmd.log
ESXi host connection stateLarge skew disconnects hosts from vCenterHosts showing notResponding without network or hardware cause
Certificate validation errors“Not yet valid” means the clock is behind real timeTLS errors in vpxd.log; integration failures from external systems
vMotion compatibility alertsHost-to-host skew above 5 minutes blocks migrations“Too large clock skew” alert between source and destination hosts
Log timestamp consistencySkewed clocks make incident investigation impossibleLog entries from different components with impossible timestamp ordering

Fixes

NTP source unreachable

Verify UDP 123 is open on firewalls between the VCSA, ESXi hosts, and NTP sources. Verify DNS resolution of NTP server names from the VCSA shell. If configured NTP servers are themselves down, reconfigure to reachable, authoritative sources. Use at least three sources for redundancy.

Configure NTP through the VAMI interface at https://<vcenter>:5480 under the Time tab for the VCSA. For ESXi hosts, configure NTP through vCenter host settings or the DCUI.

VMware Tools time sync conflicting with NTP

When the VCSA VM’s VMware Tools settings include “Synchronize guest time with host,” the ESXi host clock periodically overrides the VCSA’s NTP-synced time. If the host clock is wrong, the VCSA inherits the error, and NTP cannot stabilize.

Disable VMware Tools time synchronization on the VCSA VM. In the vSphere Client, edit the VCSA VM settings, expand VMware Tools, and uncheck the time synchronization options.

On older vSphere versions, a single “Synchronize guest time with host” checkbox controls all Tools time sync behavior. The VCSA should rely on its own NTP configuration, not the host clock.

Snapshot revert set the clock back

After a snapshot revert, the guest clock may be set to the snapshot’s creation time. NTP slews the correction gradually for moderate offsets, which can leave the VCSA with significant skew for an extended period.

For the VCSA’s chrony configuration (vSphere 7+), the makestep directive controls whether chrony steps the clock immediately instead of slewing. Check /etc/chrony.conf for the makestep setting. A configuration like makestep 1.0 3 tells chrony to step the clock if the offset exceeds 1 second, up to 3 times.

If the offset is very large and NTP is slewing too slowly, you can force an immediate step after confirming NTP sources are reachable and correct:

# WARNING: a large time jump can affect running transactions and active sessions.
# Confirm NTP sources are correct before running this.
chronyc sources    # verify at least one source is reachable
chronyc makestep   # force immediate clock step

On older VCSA versions using ntpd, restarting the service forces a resync:

systemctl restart ntpd

Circular NTP dependency

If the VCSA points exclusively at AD domain controllers for NTP, and those same DCs are the identity source for SSO, you have a circular dependency. When AD is down, both authentication and time synchronization fail simultaneously. Recovery requires fixing AD first, which itself may need time sync to be working.

Configure at least one NTP source that is not an AD domain controller. A dedicated NTP appliance, a network device that serves NTP, or a public stratum-1 or stratum-2 source provides a fallback when AD is unavailable. AD DCs can remain as additional sources, but should not be the only sources.

NTP not configured at all

Some environments deploy VCSA without configuring NTP, relying on the ESXi host clock via VMware Tools time sync. This works until the host clock drifts or the VCSA is vMotioned to a host with a different clock.

Configure NTP on the VCSA through VAMI, and on all ESXi hosts through vCenter. Verify after every patch, upgrade, or redeployment that NTP is still configured and functioning.

Prevention

  • Monitor NTP offset continuously. Alert at > 30 seconds offset and page at > 5 minutes. Do not wait for an auth failure to discover the clock is wrong.
  • Disable VMware Tools time sync on the VCSA VM. One-time fix that prevents the most common cause of persistent drift.
  • Point NTP at non-AD sources. Use at least one dedicated NTP source outside the AD dependency chain so time and auth do not fail together.
  • Verify NTP after maintenance. Patching, upgrading, and snapshot operations can disrupt NTP configuration or reset the clock.
  • Check ESXi host NTP alongside VCSA NTP. Host-to-host skew breaks vMotion and can disconnect hosts even when the VCSA clock is correct.
  • Audit NTP configuration after every VCSA redeployment or restore. File-based restores and new deployments start with default time settings.

How Netdata helps

  • Correlate NTP offset with SSO failure spikes. The chrony collector reports per-second NTP offset. When the VCSA offset crosses 30 seconds, overlaying it with LOGIN_FAILED rates from the STS logs (via the systemd journal collector) makes the causal relationship visible in seconds instead of hours of log diving.
  • Track ESXi host clocks alongside connection state. A host transitioning to notResponding at the same moment its NTP offset spikes points directly at clock skew rather than a network or hardware fault.
  • Catch drift before it breaks auth. Per-second offset collection means you see the drift trend building, not just the moment it crosses the Kerberos or SAML threshold.
  • Distinguish certificate errors from clock errors. “Certificate not yet valid” in vpxd.log correlated with a growing NTP offset confirms the clock is the problem, not the certificate. This prevents unnecessary certificate renewal.
  • Monitor post-maintenance time sync recovery. After a VCSA reboot, patch, or snapshot revert, watching the NTP offset return to baseline confirms recovery without manual verification.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.