The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-vcenter-slow-unresponsive

Operations Guides

vCenter slow and unresponsive: vpxd overload and the task queue backlog

vCenter is sluggish. The vSphere Client hangs on login, PowerCLI calls time out, and tasks that normally finish in seconds sit in Running for minutes. ESXi hosts start flipping to Not Responding even though the hosts themselves are healthy and VMs keep running. This is the vpxd overload cascade, and it is one of the most commonly misdiagnosed vCenter incidents.

The reflex is to restart vpxd. That is usually wrong. Restarting clears the queue for 5 to 15 minutes (the time it takes vpxd to rebuild its inventory cache from PostgreSQL), but whatever saturated the task pool will refill it the moment vpxd is back. You end up in a slower, noisier version of the same incident, with less forensic evidence left in the logs.

The right move is to identify the source of concurrent work, throttle or pause it, and let vpxd drain the backlog.

What this means

vpxd is a large multithreaded C++ daemon that maintains an in-memory cache of the entire inventory and dispatches every management operation as a task. It runs an internal Long-Running Operation (LRO) job queue with a finite worker pool. When too much concurrent work arrives at once, the pool saturates: new tasks queue, the SDK becomes slow, and host heartbeat processing falls behind. After the heartbeat timeout (default 120 seconds, configurable via config.vpxd.heartbeat.notRespondingTimeout), vCenter marks hosts Not Responding. The hosts are fine; vpxd just cannot service the heartbeat channel.

Two overload outcomes have very different signatures:

  • Thread pool exhaustion manifests as tasks stuck in Queued. vpxd is not picking up new work. The offending load is usually a flood of API calls from a single client.
  • Slow execution manifests as tasks stuck in Running for far longer than baseline. vpxd is executing, but slowly. The cause is usually downstream: database latency, storage latency on the VCSA VM, or stats rollup contention.

In the terminal stage, vpxd exhausts its hard memory limit and panics with Memory exceeds hard limit. Panic, then restarts. On a full LRO queue you may also see vmodl.fault.SystemError with reason Too many outstanding operations returned to SDK callers (confirmed in Broadcom KB 318208).

flowchart TD
    A["Concurrent work spike
backup, scripts, DRS storm"] --> B["vpxd task pool saturates"] B --> C["Tasks queue"] C --> D["SDK response time climbs"] D --> E["Clients retry, adding load"] E --> B B --> F["Host heartbeat processing delayed"] F --> G["Hosts appear Not Responding"] B --> H{"Memory pressure?"} H -- yes --> I["Memory exceeds hard limit
vpxd panics, vmon restarts"] H -- no --> J["Stuck degraded until source throttled"]

Common causes

CauseWhat it looks likeFirst thing to check
Backup solution snapshot stormHundreds of createSnapshot / removeSnapshot tasks land in the same window. Often correlates with the nightly backup schedule.Get-Task -Status Running during the backup window
Runaway PowerCLI or SDK clientA script opens many concurrent SDK sessions or issues createContainerView without destroy. SDK session count climbs from one client IP.vpxd-profiler.log per-session ClientIP and Username
DRS stormAggressive DRS level combined with frequent workload changes drives a burst of vMotions. Queue is full of DrmMigrate tasks.Get-VIEvent -Types DrsVmMigratedEvent rate over the last hour
Stats rollup overlapping busy window5-minute rollup job takes longer than its interval. vpxd_hist_stat* tables are large. CPU spikes on a regular cadence.Rollup lag in vpxd.log, latest sample_time in vpxd_hist_stat1
ContainerView leak from monitoring toolvpxd memory grows monotonically. createContainerView count far exceeds vim.view.View.destroy. Ends in Memory exceeds hard limit. Panic.grep createContainerView vs destroy in vpxd.log
SMS/SPS thread pool starvationLogs show Active thread count is: 20, Core Pool size is: 20, Queue size: X, ThreadPool Starvation Alert. High Storage Profile Service activity.SMS/SPS log entries
Under-provisioned VCSAA Tiny or Small appliance running a Medium or Large inventory. Symptoms appear gradually as inventory grows.Deployment size vs VMware Configuration Maximums for your inventory
Database fragmentationDoHostSyncTime values spike in vpxd.log [VpxProfiler] entries. vpxd_hist_stat* tables bloated with dead tuples.[VpxProfiler] entries, autovacuum dead tuple ratio

Quick checks

Start with read-only checks. Do not restart anything yet.

# Confirm vpxd is actually running (not crash-looping)
/usr/lib/vmware-vmon/vmon-cli --status vpxd

# Overall partition health - the most common adjacent killer
df -h /storage/log /storage/db /storage/seat

# Check the authenticated SDK endpoint, not just the WSDL
time curl -sk -X POST https://<vcsa>/sdk \
  -H "Content-Type: text/xml" \
  -d '<soap:Envelope xmlns:soap="http://schemas.xmlsoap.org/soap/envelope/">
    <soap:Body>
      <RetrieveServiceContent xmlns="urn:vim25">
        <_this type="ServiceInstance">ServiceInstance</_this>
      </RetrieveServiceContent>
    </soap:Body>
  </soap:Envelope>'
# Count tasks in each state - distinguishes Queued vs Running overload
Get-Task | Group-Object State

# See what is actually running right now
Get-Task | Where-Object {$_.State -eq "Running"} | Select Name, StartTime, EntityName

# Identify which hosts vCenter currently considers unreachable
Get-VMHost | Where-Object {$_.ConnectionState -ne "Connected"} | Select Name, ConnectionState
# Database latency indicator from vpxd profiler entries
grep "\[VpxProfiler\]" /var/log/vmware/vpxd/vpxd.log | grep "DoHostSyncTime" | tail -20

# ContainerView create vs destroy balance - imbalance indicates a leak
# NOTE: field position is version-dependent; adjust the awk column as needed
grep "createContainerView" /var/log/vmware/vpxd/vpxd.log | grep "BEGIN" | wc -l
grep "vim.view.View.destroy" /var/log/vmware/vpxd/vpxd.log | grep "BEGIN" | wc -l

How to diagnose it

  1. Confirm overload, not a sibling failure. Certificate expiry, disk full on /storage/db, and vPostgres being down all produce similar symptoms. Run the quick checks above. If vpxd is up, partitions are not at 100%, and the SDK probe returns slowly rather than refusing connection, you are in overload territory.

  2. Distinguish Queued from Running. This single distinction tells you where to look next.

    • Mostly Queued with low vpxd CPU: thread pool exhausted by a flood of API calls. Look for a single offending SDK client.
    • Mostly Running with high vpxd CPU or high DB latency: vpxd is genuinely compute- or database-bound. Look at stats rollup, DRS, or storage.
  3. Trace the offending session. vpxd-profiler.log records per-session metrics including ClientIP and Username. Find a session ID that appears in the slow task entries, then trace it.

    find /var/log/vmware/vpxd/ -iname "vpxd-profiler*" -type f \
      -exec grep -H "<session-ID>" {} \; | grep "ClientIP" | head -n 5
    

    Correlate ClientIP with your backup server, monitoring tool, or automation host.

  4. Check the ContainerView balance. If createContainerView counts dwarf destroy counts, a third-party monitoring or backup integration is leaking views. This is a common cause of vpxd memory exhaustion crashes.

  5. Check the database side. Look for [VpxProfiler] entries with DoHostSyncTime values well above your environment’s baseline. Confirm from the ESXi host running the VCSA VM: DAVG/cmd in esxtop above 25 ms sustained means the underlying datastore is slow.

  6. Check the stats rollup cadence. If the 5-minute rollup job is taking longer than 5 minutes, it overlaps the next interval and competes with live operations.

    # Read-only query. Do not modify VCDB tables directly.
    /opt/vmware/vpostgres/current/bin/psql -U postgres -d VCDB -c \
      "SELECT sample_time FROM vc.vpxd_hist_stat1 ORDER BY sample_time DESC LIMIT 1;"
    

Metrics and signals to monitor

SignalWhy it mattersWarning sign
vpxd CPU utilizationCompute saturation is the direct indicator of overload.Sustained above 70% of allocated cores with growing queue
Task queue depthDistinguishes saturation from slow execution.Sustained > 0 outside planned bulk operations
Tasks in Queued vs RunningTells you whether the pool is exhausted or just slow.Queued climbing while vpxd CPU is moderate
SDK response timeThe user-facing latency signal.Above 5 seconds; healthy is under 1 second
Hosts Not Responding countIndicates vpxd has fallen behind on heartbeat processing.Multiple hosts flipping simultaneously while hosts are actually up
vpxd RSS memoryInventory cache growth or leak.Steady climb without inventory growth, or approaching the hard limit
vpxd error rateSpecific patterns (OOM, DB connect, cert) identify root cause.Any Memory exceeds hard limit or Too many outstanding operations
DoHostSyncTime profiler entriesDatabase latency from vpxd’s perspective.Sustained values well above your environment’s baseline
/storage/db and /storage/seat usageDatabase pressure kills vpxd.Above 80%
SDK sessions by client IPCatches misbehaving integrations early.Single IP opening dozens of concurrent sessions

Fixes

Throttle the source first

Before changing anything on vCenter, find the client identified in step 3 of diagnosis and throttle it.

  • Backup solution: pause the job. Most enterprise backup products have a per-vCenter concurrency setting. Veeam, Commvault, and similar tools can saturate vpxd with concurrent snapshot operations across hundreds of VMs. Reduce the parallelism or shift the window off the stats rollup cadence.
  • PowerCLI or custom automation: limit concurrent runspace count, add retries with backoff instead of tight loops, and ensure every createContainerView is paired with a destroy. A script that opens SDK sessions in a loop without closing them will exhaust vpxd within minutes in a large inventory.
  • DRS storm: temporarily switch the cluster to manual. This stops new migrations immediately while you investigate. See the DRS thrashing guide.
  • Stats rollup overlap: reduce the statistics level to 1 or 2. Level 3 and 4 generate more database write load and are a common slow-burn cause of vpxd pressure. Lower the level, let the backlog drain, and review whether the higher granularity was needed.

Memory pressure and the ContainerView leak

If vpxd.log shows Memory exceeds hard limit. Panic or the create/destroy check shows a wide imbalance, a third-party integration is leaking ContainerViews. Updating or patching the offending client is the durable fix. In the short term, throttling that client’s session count stabilizes vpxd without a restart.

Database fragmentation

When DoHostSyncTime is consistently high and vpxd_hist_stat* tables have a high dead tuple ratio, database fragmentation is the bottleneck. The remediation is a VACUUM (FULL, ANALYZE) on the affected statistics table. This locks the table and requires equivalent free space, so schedule it in a maintenance window. Confirm the table is the problem first by checking pg_stat_user_tables.n_dead_tup.

Under-provisioned VCSA

If inventory counts are approaching the deployment size maximums and overload symptoms appear under modest concurrent load, the appliance is undersized. Vertical scaling from Small to Large or X-Large resolves the chronic pressure. This is not a quick fix; plan it as a maintenance operation.

When to actually restart vpxd

Restart only when:

  • vpxd is in a crash loop and vmon has given up restarting it.
  • The queue is so deep that throttling the source will not drain it before the next business window.
  • You have already captured the diagnostic evidence (profiler logs, task states, ContainerView counts).

A restart drops all in-flight tasks and requires 5 to 15 minutes to rebuild the inventory cache from PostgreSQL. In large environments, all hosts will briefly appear Not Responding during reconnection. Expect a follow-on DRS burst as catch-up migrations fire after vpxd is healthy again.

# Destructive: drops all in-flight tasks and disconnects all hosts temporarily
service vpxd restart

Prevention

  • Right-size the VCSA for inventory and concurrency, not just host count. A heavily tagged environment with many distributed portgroups hits limits sooner than raw VM count suggests.
  • Rate-limit SDK sessions per client. Know which integrations hold persistent sessions and audit them quarterly. Backup, monitoring, and orchestration tools are the usual offenders.
  • Keep statistics at level 1 or 2 unless you have an active reason to raise it. Document the reason and the expected end date if you do.
  • Monitor snapshot age daily. Backup jobs that fail to clean up snapshots create both datastore pressure and vpxd task load.
  • Track vpxd-profiler.log session patterns proactively, not just during incidents. A baseline makes the offending client obvious during the next overload.
  • Apply vCenter patches that resolve overload bugs. Several specific vpxd crash modes have been fixed in recent releases; check the release notes for your target version against current Broadcom KBs.

How Netdata helps

  • Per-second vpxd CPU and memory metrics surface the saturation climb before the queue is fully exhausted, giving you a window to throttle the source rather than recover from a panic.
  • Task queue depth and SDK response time correlated with vpxd CPU distinguish thread pool exhaustion from slow execution. Two lines on one chart replace several minutes of log grepping.
  • Host connection state changes correlated with vpxd load make it obvious when Not Responding is a vCenter-side problem rather than a host failure.
  • VCSA per-partition disk usage catches the adjacent failure modes (/storage/log log bombs, /storage/db database pressure) that often co-occur with overload or trigger it.
  • Anomaly detection on vpxd error rate and SDK latency flags the slow-burn version of this incident, where task duration creeps upward over weeks as inventory outgrows the deployment size.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.