The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-memory-compression

Operations Guides

vSphere memory compression: the reclamation tier between balloon and swap

Memory compression is the third tier in ESXi’s four-tier memory reclamation hierarchy. It sits between ballooning, which asks the guest OS to return pages, and host-level swapping, which writes pages to .vswp files on a datastore and destroys latency. Compression is ESXi buying time: keep the page in RAM but in a smaller form, at the cost of CPU cycles and added access latency.

The key point up front: compression is not a tuning exercise. Any sustained non-zero compression rate means the host has exhausted ballooning and is in the zone between degraded and catastrophic.

What it is and why it matters

ESXi overcommits memory through a four-tier reclamation hierarchy, invoked in order of increasing desperation:

  1. Transparent Page Sharing (TPS). Deduplicates identical pages across VMs. Largely disabled by default since vSphere 6.0 for security reasons, so in most modern deployments this tier is effectively absent.
  2. Ballooning. The vmmemctl driver inside the guest is inflated, forcing the guest OS to page internally using its own swap or pagefile. This is the first tier of active reclamation.
  3. Compression. Pages that ballooning could not reclaim are compressed into a per-VM cache that lives inside the VM’s own memory space.
  4. Host-level swapping. VM memory pages are written to .vswp files on the datastore. This is catastrophically slow because the .vswp file competes with VM disk I/O on shared storage.

Compression exists to delay tier 4. Each compressed page avoids a disk I/O but consumes CPU cycles on both compress and decompress, and adds roughly 2-10x DRAM access latency when the guest touches that page. That is a painful tradeoff, but it is orders of magnitude better than touching a .vswp file on shared storage.

The reason operators need to treat compression as a distinct tier rather than background noise is that it is the earliest hypervisor-visible signal that memory pressure has crossed from manageable to dangerous. Ballooning can run for long periods with modest guest impact. Compression cannot. By the time you see sustained ZIP/s in esxtop, the host is already in trouble.

How it works

When ballooning has reached its limit, or cannot run because VMware Tools is absent, the VMkernel begins compressing memory pages instead of swapping them immediately. Each 4 KB page is evaluated for compressibility. Pages that compress well are stored in the compression cache. Pages that can be compressed to 2 KB or smaller go to the compression cache; pages that cannot are swapped to disk directly.

flowchart LR
    A["Memory pressure"] --> B["Tier 1: TPS
dedup, off by default"] B --> C["Tier 2: Balloon
guest pages internally"] C --> D["Tier 3: Compression
per-VM cache"] D --> E["Tier 4: Swap
.vswp on datastore"] D -.-> F["CPU cost
2-10x DRAM latency"] E -.-> G["Disk-speed RAM
catastrophic"]

The compression cache is per-VM and lives inside the VM’s own allocated memory, not in separate host memory. The default maximum cache size is 10% of the VM’s configured memory, controlled by the advanced setting Mem.MemZipMaxPct. Current Broadcom documentation continues to specify 10% for ESX 9.0. Because the cache occupies guest-visible address space, raising Mem.MemZipMaxPct to very high values can paradoxically increase memory pressure by consuming more of the VM’s own RAM for compressed storage rather than freeing host capacity.

When the compression cache fills, ESXi does not grow it dynamically. It evicts the least-recently-used compressed page, decompresses it, and writes it to the .vswp file. At that point, the host has crossed from compression into active swapping. Compression was only a buffer, and the buffer is now full.

Compression can be toggled with the advanced setting Mem.MemZipEnable (1 = enabled, 0 = disabled). It is enabled by default on current ESXi hosts; current Broadcom documentation continues to specify the same 10% default cache setting. Disabling it on a production host is disruptive: it removes the buffer between ballooning and swap entirely, so any pressure that ballooning cannot absorb goes directly to .vswp I/O with no intermediate step.

Where it shows up in production

The compression rate signals are visible in two places:

  • esxtop: the ZIP/s column (compression rate) and UNZIP/s column (decompression rate), per VM in the memory view.
# On the ESXi host (via SSH or DCUI), check per-VM compression rates
esxtop
# press 'm' for memory view
# press 'f' to add ZIP/s and UNZIP/s fields if they are not visible by default
# ZIP/s = pages compressed per second, UNZIP/s = pages decompressed per second
  • vCenter API: mem.compressionRate.average (KBps), mem.decompressionRate.average (KBps), and mem.compressed.average (KB total currently held in cache).

A sustained non-zero ZIP/s means the host is actively compressing pages because ballooning was insufficient. This is a TICKET-level condition: any sustained rate indicates memory pressure beyond what ballooning can handle, and the host is in the zone between degraded and catastrophic.

UNZIP/s (decompression rate) is equally important. Active decompression means the guest is touching pages that were previously compressed. Every decompress costs CPU and adds latency to that memory access. High UNZIP/s alongside high ZIP/s is a particularly bad pattern: the host is compressing and decompressing the same working set repeatedly, burning CPU on both sides without making net progress. This is the compression equivalent of swap thrashing.

The balloon-without-compression trap

If you see swap or compression activity with zero or near-zero balloon, the balloon driver is not functioning. VMware Tools may be absent, the balloon driver may be disabled, or the guest may be unresponsive to inflation requests. The host skipped tier 2 and went straight to tier 3 or 4. This is worse than the raw numbers suggest because ballooning is the gentlest reclamation mechanism. Without it, every pressure spike goes directly to compression and swap.

Poor compression candidates

Compression effectiveness depends on the data itself. Encrypted or highly randomized data compresses poorly. The specific production pattern to watch for: databases with Transparent Data Encryption (TDE), encrypted VMs, or in-memory databases with random data. The VMkernel spends CPU cycles attempting to compress these pages, fails to meet the compression threshold, and swaps them anyway. The compression step was pure overhead.

In environments running TDE databases under memory pressure, you may see elevated host CPU with no corresponding reduction in swap rate. The compression cache is churning through incompressible pages. If your workload is predominantly encrypted or random data, compression provides little benefit and disabling it with Mem.MemZipEnable = 0 removes the wasted CPU, accepting that the host will swap directly when pressure exceeds what ballooning can absorb. VMware has not published a specific KB or advisory for this TDE/encrypted-workload interaction in the sources reviewed; treat this as workload-dependent field experience.

Tradeoffs and when this matters

Compression is a brake, not a solution. It delays swapping but does not prevent it. The cache has a fixed ceiling, and once it fills, swap begins regardless. The operational question is never “should we tune compression” but “why is the host under enough pressure to need it at all.”

FactorImplication
CPU costEvery compress and decompress consumes pCPU cycles. On an already CPU-constrained host, compression adds a secondary tax.
Access latencyCompressed pages cost roughly 2-10x DRAM latency on access due to decompression. Faster than disk swap, slower than uncompressed RAM.
Cache ceilingDefault 10% of VM memory. When full, LRU eviction decompresses and swaps. Compression only delays swap.
Incompressible dataEncrypted, random, or already-compressed data wastes CPU before swapping anyway.
Guest memory footprintThe cache lives inside the VM’s own memory space. Raising Mem.MemZipMaxPct consumes more guest-visible RAM.

Correlating with the rest of the cascade

Compression never appears in isolation. Its diagnostic value comes from correlating it with the adjacent tiers:

  • Balloon should be at or near maximum when compression activates. If balloon (mem.vmmemctl.average, or MCTLSZ in esxtop) is low or zero but compression is active, the balloon driver is not functioning. This is the Tools-absent scenario described above.
  • Swap follows if compression is insufficient. Watch mem.swapinRate.average and mem.swapoutRate.average alongside compression. The moment swap-in rate goes above zero, the compression cache has filled and pages are being recalled from disk. This is the cliff edge.
  • Host CPU rises with compression activity. Compression and decompression consume pCPU. If you see host CPU climbing alongside ZIP/s and UNZIP/s, some of that CPU is the compression tax, not application demand. This matters when deciding whether to migrate VMs or add host memory: you may be closer to CPU contention than the guest-level numbers suggest.

The most dangerous misread is treating compression as normal background activity that a well-overcommitted host just does. A host that sustains compression for minutes is a host that is one cache-fill away from swapping, and swapping is where the death spiral begins.

Signals to watch in production

SignalWhy it mattersWarning sign
mem.compressionRate.average (ZIP/s)Active compression means pressure beyond ballooningAny sustained non-zero rate
mem.decompressionRate.average (UNZIP/s)Guest touching compressed pages; CPU and latency costHigh UNZIP/s alongside high ZIP/s means churning working set
mem.compressed.averageTotal pages currently held in the compression cacheApproaching 10% of VM memory means cache near full, swap imminent
mem.vmmemctl.average (balloon)Must be near max when compression activatesLow or zero balloon with active compression means Tools absent
mem.swapinRate.averagePages recalled from .vswp means compression failed to absorbAny sustained non-zero rate is an active performance emergency
Host CPU utilizationCompression consumes pCPU cyclesCPU climbing with no application change may be the compression tax

How Netdata helps

  • Per-second ZIP/s and UNZIP/s catches the transition from transient spike to sustained pressure before the cache fills and swap begins. Five-minute vCenter rollups average away the exact moment the cascade escalates.
  • Single-view correlation of balloon, compression, and swap rates shows whether balloon is at max, whether compression is absorbing the overflow, and whether swap has started despite compression.
  • Anomaly detection on compression rate flags the first sustained departure from zero, the earliest hypervisor-visible signal that a host has crossed from manageable pressure into the danger zone.
  • Host CPU alongside compression metrics separates the compression tax from genuine application demand, which matters when deciding whether to migrate VMs or add host memory.
  • Composite alerting across the memory cascade, from balloon through compression to swap, avoids the common failure of alerting only on swap, which fires after the damage is already done.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.