The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-thin-provisioning-growth

Operations Guides

vSphere thin-provisioned VMDK growth: space that never comes back without UNMAP

A thin-provisioned VMDK starts small and grows as the guest writes data. It never shrinks on its own. When a guest deletes a 100 GB database dump, the space inside the guest filesystem becomes free, but the VMDK file on the datastore stays at its high-water mark. Over months, the datastore fills with blocks the guest considers empty. The symptom is a datastore that creeps toward full while the guests report plenty of free space inside.

VMFS does not know which blocks the guest has freed. The guest filesystem marks blocks as free in its own metadata, but nothing automatically tells the hypervisor or the storage array. Reclaiming that space requires a deliberate signal chain: the guest issues a TRIM or UNMAP command, the hypervisor translates that into a VAAI UNMAP operation, and the array releases the underlying blocks. Break any link and the space stays stranded.

This article covers the reclamation chain end to end, VMFS-5 versus VMFS-6 differences, NFS complications, and the signals that reveal silent over-commitment before a datastore fills.

Why thin VMDKs only grow

A thin VMDK is a sparse file. VMFS allocates blocks on demand as the guest writes to previously unwritten regions. Once allocated, those blocks belong to the VMDK. The guest can later delete files, and its filesystem will mark those blocks as free internally, but VMFS has no view into the guest filesystem. From the datastore’s perspective, the block is still in use.

The high-water mark behavior is intentional. It avoids probing guest intent and keeps the write path fast. The tradeoff is that space consumption is monotonic unless something actively tells VMFS that blocks are reclaimable.

That something is the SCSI UNMAP command, also called TRIM in ATA terminology. The guest must issue it, and the hypervisor must propagate it. Without both, thin VMDKs grow but never shrink.

How the reclamation chain works

Space reclamation requires cooperation across three layers: the guest OS, the VMFS datastore, and the storage array. Each layer can silently drop the signal.

flowchart TD
  A[Guest deletes file] --> B{Guest issues TRIM/UNMAP?}
  B -- No --> C[VMDK stays at high-water mark]
  B -- Yes --> D{Datastore type?}
  D -- NFS --> E[No UNMAP path
array plugin only] D -- VMFS-5 --> F[EnableBlockDelete=1 for in-guest
or manual esxcli unmap] D -- VMFS-6 --> G[Automatic background UNMAP] F --> H[VAAI UNMAP to array] G --> H H --> I[Array releases blocks]

The guest side is the first gate. On Linux, the fstrim command, or the discard mount option, issues UNMAP for freed blocks. On Windows, Optimize-Volume -ReTrim does the equivalent. Many distributions do not run fstrim on a schedule by default, and Windows Server does not always detect the virtual disk as thin-provisioned. If the guest never issues the command, nothing downstream happens.

The VMFS side is the second gate. VMFS-5 and VMFS-6 handle reclamation very differently, and mixing them up is one of the most common reasons operators believe UNMAP is broken.

VMFS-5: manual UNMAP only

VMFS-5 has no automatic background reclamation. To reclaim space, you must run a manual UNMAP against the datastore.

# Reclaim dead blocks on a VMFS-5 datastore
# WARNING: Generates sustained UNMAP I/O that can saturate array ports.
# Schedule during a maintenance window and monitor array latency.
esxcli storage vmfs unmap -l <datastore-name>
# The reclaim-unit (default 200) controls batch size per iteration.
# Lower it on arrays with limited UNMAP throughput.

For in-guest UNMAP to propagate through VMFS-5 to the array, the VM must have the EnableBlockDelete advanced parameter set to 1. It is disabled (0) by default. Without it, fstrim or Optimize-Volume inside the guest appears to succeed from the guest’s perspective, but VMFS-5 does not forward the reclamation to the array.

Thick-provisioned VMDKs (lazy zeroed or eager zeroed) do not support UNMAP at all. If you need reclamation on a thick disk, the supported path is converting to thin via Storage vMotion. After that conversion, Windows Server may not detect the disk as thin until a guest reboot, so Optimize-Volume -ReTrim can fail with “not supported” even though the VMDK descriptor shows ddb.thinProvisioned = "1".

VMFS-6: automatic background UNMAP

VMFS-6 introduced asynchronous background space reclamation. By default it runs at low priority, typically reclaiming at 25 to 50 MB/s. No manual intervention is required for array-side reclamation of blocks VMFS already knows are dead.

On VMFS-6, EnableBlockDelete is ignored. The auto-UNMAP mechanism handles array reclamation directly. However, the guest still must issue TRIM or UNMAP for the background process to have anything to reclaim. VMFS-6 only learns about dead blocks when the guest tells it.

Two granularity limits govern whether VMFS-6 reclamation actually works:

  • Array granularity: VMFS-6 auto-UNMAP requires the array’s unmap granularity to be 1 MB or less. If the array page or chunk size exceeds 1 MB, automatic reclamation silently does nothing. This is a common root cause for “UNMAP is configured but space is not coming back.”
  • Guest UNMAP granularity: VMFS-6 processes guest UNMAP requests only when the space to reclaim equals 1 MB or a multiple of 1 MB. Sub-1 MB or misaligned requests are dropped silently.

Auto-UNMAP is also slow by design. After a VM is deleted or a large VMDK is removed, reclamation can take 12 to 24 hours to complete. In-guest UNMAP, running fstrim or Optimize-Volume, triggers reclamation within minutes because it pushes the signal directly rather than waiting for the background sweep.

A historical note: ESXi 6.7 had a bug where the unmap rate calculation was incorrect, causing excessive UNMAP throughput even when priority was set to Low. If you run 6.7 and your array reports UNMAP storms, patch before chasing array-side causes.

If automatic reclamation is set to None, Broadcom KB 323112 still lists esxcli storage vmfs unmap for a manually reclaimed VMFS-6 datastore.

Snapshots block reclamation

Snapshots change the reclamation story. When a VM has snapshots, writes go to a delta disk in SEsparse format. Automatic space reclamation on VMFS-6 for snapshot delta files works only on ESXi 6.7 and later, and only on the top snapshot while the VM is powered on. Older ESXi versions, or VMs with deep snapshot chains, may not reclaim space from the delta layer at all.

This matters because environments with long-lived snapshots, often from backup jobs that failed to clean up, accumulate space in deltas that cannot be reclaimed until the snapshots are consolidated. A VM with a 200 GB base disk and an active snapshot can grow the delta to 200 GB, and none of that space is reclaimable while the snapshot exists.

NFS: the second layer of thin provisioning

NFS datastores add a complication. SCSI UNMAP is not supported on NFS datastores. Guest TRIM and DEALLOCATE commands do not propagate through the NFS client to the array. If your VMDKs live on NFS, the only reclamation path is whatever the array vendor provides through a vSphere plugin or a separate management interface.

NFS also introduces a second layer of thin provisioning. The NFS server reports free space to ESXi based on its own view of the LUN or volume, which may itself be thin-provisioned on the array. The datastore can report plenty of free space while the underlying array volume is near its physical limit. The guest thinks it has space, the NFS client thinks it has space, and the array is the only one that knows the truth.

A known reporting quirk: thin-provisioned VMDKs on NFS may display as thick in the vSphere Client because the NFS server returns metadata that ESXi misinterprets. Do not rely on the GUI provisioning type column for NFS datastores.

Over-commitment: provisioned vs consumed

The deeper operational risk with thin provisioning is over-commitment. Thin VMDKs let you provision more capacity than the datastore physically holds, on the assumption that not every guest will write to its full allocation. This is reasonable until it is not.

Two numbers matter, and they diverge over time:

  • Provisioned capacity: the sum of all VMDK sizes as configured. This can be many times the datastore’s physical size.
  • Consumed capacity: the actual blocks written. This is what the datastore physically holds.

When consumed capacity approaches the physical limit, the datastore is full regardless of what provisioned capacity says. The standard failure is a VM that tries to grow its thin VMDK, create a swap file, or extend a snapshot delta, and finds no space. All VMs on that datastore halt writes, not just the one that triggered the allocation.

The trap is that consumed capacity grows monotonically when reclamation is not happening. A datastore can sit at 40% consumed for months, then climb to 95% in a week if a workload churns through large temporary files. Without reclamation, those temporary files leave permanent residue.

Where reclamation silently breaks

When a datastore keeps filling despite guests deleting data, the problem is almost always a broken link in the reclamation chain. Check each layer:

  • Guest not issuing TRIM: Run fstrim -v / on Linux or Optimize-Volume -DriveLetter C -ReTrim -Verbose on Windows. If the command is unsupported or reclaims nothing, the guest is the gap. Confirm the disk is seen as thin and that the filesystem supports discard.
  • VMFS-5 without EnableBlockDelete: In-guest UNMAP appears to work but nothing reaches the array. Set EnableBlockDelete=1 or rely on manual esxcli storage vmfs unmap.
  • VMFS-6 array granularity above 1 MB: Auto-UNMAP runs but reclaims nothing. This is invisible from ESXi. Verify the array’s unmap page or chunk size through the array management interface.
  • Snapshots in the chain: Deltas block reclamation below the top snapshot. Consolidate or delete snapshots before expecting space to return.
  • NFS datastore: There is no guest-driven UNMAP path. Use the array vendor’s vSphere plugin or accept that space reclamation requires array-side action.
  • Thick VMDKs: Thick disks, including those converted from VMFS-3 datastores or created manually, do not support UNMAP. Convert to thin via Storage vMotion.
  • VAAI UNMAP silently unsupported: If the LUN partition table was created manually rather than via the vSphere Client, or if the datastore was upgraded from VMFS-3, VAAI UNMAP may not function even when everything else looks correct.

Signals to watch in production

SignalWhy it mattersWarning sign
Datastore free space (% and absolute)A full VMFS datastore halts all VMs on it.Under 15% is a ticket; under 5% AND under 500 GB free is page-worthy. For small datastores, under 10 GB alone.
Provisioned vs consumed ratioHigh over-commitment means a workload shift can fill the datastore.Provisioned more than 2 to 3 times physical with no reclamation policy.
Thin vs thick breakdownThick disks are not reclaimable. Knowing the mix tells you how much can be recovered.Large thick VMDKs on a tight datastore.
Snapshot delta growthDeltas grow with writes and may block reclamation.Any snapshot older than 72 hours, or delta approaching the base disk size.
VMDK size vs guest-used spaceIf the VMDK is much larger than the guest’s used space, reclamation is not happening.VMDK at 200 GB while the guest reports 60 GB used.
Array-side thin pool utilizationThe physical truth, especially for NFS and thin LUNs.Array pool near its limit while the datastore reports free space.
UNMAP I/O rate (VMFS-6)Confirms auto-UNMAP is actually running and reclaiming.No reclamation activity despite known large deletes in guests.

The datastore free space thresholds use AND, not OR. A 50 TB datastore at 5% free still has 2.5 TB. A 200 GB datastore at 5% free has 10 GB, which is an emergency. Use both percentage and absolute numbers.

How Netdata helps

Netdata’s VMware vSphere integration surfaces the datastore signals that reveal thin-provisioning drift before it becomes an outage.

  • Per-second datastore free space: catch the moment a datastore crosses a percentage or absolute threshold, not a rolled-up average hours later.
  • Provisioned vs consumed capacity correlation: see over-commitment ratios change as workloads churn, which is the earliest sign that reclamation has stopped working.
  • Datastore latency alongside free space: when a datastore nears full, latency often rises first as the array struggles with allocation. Correlating the two distinguishes a space problem from a performance problem.
  • Snapshot age and delta growth tracking: snapshots are the most common accelerant for datastore pressure. Netdata exposes snapshot age per VM, letting you catch deltas that block reclamation.
  • Anomaly detection on consumption rate: a sudden change in the daily free-space slope, flagged by ML, is often the first indicator that a workload is leaking space into an un-reclaimed thin VMDK.
  • Cross-layer correlation: correlate datastore free space with host swap activity, VM disk latency, and backup job timing to identify whether a backup storm or memory pressure cascade is driving the growth.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.