The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-storage-queue-depth

Operations Guides

vSphere storage queue depth: QUED, ACTV, and DSNRO saturation

Storage latency on an otherwise healthy all-flash array is usually a queue problem, not an array problem. The VMkernel stacks I/O requests in queues at every layer of the path between the guest and the physical device. When those queues fill, requests wait, and the latency the guest sees climbs independently of what the array is actually doing.

The two numbers that tell you whether you are in this state are ACTV (in-flight I/O at the device) and QUED (I/O waiting in the VMkernel). Sustained QUED greater than zero with KAVG climbing is the signature of device queue saturation. The fix is rarely “make the array faster”; it is almost always “give the device more outstanding-I/O capacity, or stop funneling everything through a single path.”

What this means

The VMkernel storage path has a configurable queue at each layer. From the guest downward:

  1. Guest and virtual adapter queue. Per virtual disk, governed by the SCSI controller type. PVSCSI gives a much larger adapter queue than LSI Logic and is the right choice for I/O-heavy VMs.
  2. VMkernel device queue (DQLEN). Per LUN or device. This is the queue the HBA draws from.
  3. HBA or adapter queue (AQLEN). Per physical HBA. Bounded by the HBA driver’s per-LUN limit (for example, ql2xmaxqdepth on QLogic).

ACTV is the count of I/Os the VMkernel has handed to the device and is waiting on. QUED is the count of I/Os sitting in the VMkernel because the device queue is full and cannot accept more. The relationships that matter:

  • ACTV can never exceed DQLEN.
  • When ACTV equals DQLEN and there is more demand, additional I/O backs up into QUED.
  • QUED greater than zero sustained drives KAVG up, because requests are waiting in the kernel before they ever reach the device.

GAVG (what the VM sees) equals DAVG (device and array latency) plus KAVG (kernel latency, including queue wait time). If the array reports sub-millisecond latency but the VM sees tens of milliseconds, and QUED is non-zero, the bottleneck is in the ESXi stack, not below it. That split is the single most useful diagnostic when an all-flash array “feels slow.”

flowchart TD
  A[VM I/O] --> B[PVSCSI adapter queue]
  B --> C[VMkernel device queue - DQLEN]
  C --> D[ACTV: in-flight at device]
  C --> E[QUED: waiting in VMkernel]
  D --> F[HBA / physical path]
  F --> G[Storage array]
  E -. queued because ACTV == DQLEN .-> D
  E -. raises KAVG .-> H[VM sees GAVG = DAVG + KAVG]

Common causes

CauseWhat it looks likeFirst thing to check
DSNRO throttling shared LUNsQUED rises the moment a second VM issues I/O to the same LUN; per-world outstanding I/O capped at DSNRO (default 32); KAVG climbs; DAVG stays lowesxtop device view: ACTV versus QUED while adding load from a second VM on the same LUN
Dead path funnelingQUED spikes after a path failure; ACTV pegged on one path; AQLEN exhausted on the surviving adapteresxcli storage core path list for dead paths; vmkernel.log
Device queue too low for the arrayAll-flash array reports under 1 ms; VMs see 10 to 50 ms; KAVG dominant; QUED sustained above zeroCompare DQLEN to the array’s expected outstanding I/O; check HBA driver default
Slow array or controller failoverDAVG dominant; KAVG rises secondarily as the queue backs up behind slow device responsesDAVG versus KAVG split; array-side alerts
Snapshot chain read amplificationSingle VM slow, not the whole datastore; latency on reads onlySnapshot inventory on the affected VM

Quick checks

Read-only and safe during production.

# esxtop -> press 'u' -> press 'f' -> enable Queue Stats (DQLEN, ACTV, QUED, %USD)
esxtop

# DSNRO is only visible under competing-worlds load.
# Probe with at least two VMs issuing I/O to the same LUN.

# Path state - look for any path that is not 'active'
esxcli storage core path list | grep -E "Path|State"

# Per-path statistics and error counts
esxcli storage core path stats list

# Multipath policy per device (Fixed, MRU, Round Robin)
esxcli storage nmp device list

# vmkernel log for path, HBA, NMP, APD, and PDL events
grep -iE "path|NMP|HBA|APD|PDL" /var/log/vmkernel.log | tail -50

How to diagnose it

  1. Confirm the queue is the bottleneck, not the array. In esxtop device view, compare DAVG and KAVG. KAVG above 2 ms sustained with DAVG low points at the ESXi stack. KAVG near zero with high DAVG points below ESXi.
  2. Look at QUED and ACTV together. QUED greater than zero for more than about 30 seconds with ACTV pinned at DQLEN means the device queue is saturated. If QUED is zero and latency is high, the problem is elsewhere (array, IP storage network, snapshot chain).
  3. Check for DSNRO contention. DSNRO only kicks in when more than one world (VM or virtual disk) is issuing I/O to the same LUN. A single-VM LUN can use the full device queue depth; the moment a second VM touches it, each world’s outstanding I/O is capped at DSNRO, which defaults to 32. If a previously fast VM starts queueing when a second VM spins up on the same datastore, DSNRO is the throttle. Compare ACTV per world during the contention window.
  4. Check path state. esxcli storage core path list should show all expected paths as active. Any dead path means the surviving paths carry the full load, and their per-path queue depth can be exhausted even when the aggregate queue depth looks fine.
  5. Check the multipathing policy. Fixed and MRU pin a LUN to one path; Round Robin spreads I/O across active paths. A LUN on Fixed policy with a single active path is one path failure away from a queue blow-up.
  6. Confirm with the array. If DAVG is the dominant latency and KAVG is low, the queue is full because the device is slow, not because the queue is small. Raising queue depth will not help and can make the problem worse by stacking more load on a device that is already behind.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
QUED (per device)Direct measure of VMkernel queueingSustained above zero for more than 30 s
ACTV divided by DQLENHow close the device is to max concurrencyACTV regularly above 70% of DQLEN
KAVG (disk.kernelLatency.average)Latency added by the VMkernel stackAbove 2 ms indicates queueing
DAVG (disk.deviceLatency.average)Array plus transit latencyAll-flash above 5 ms is a problem
GAVG (disk.totalLatency.average)What the VM actually seesAbove 30 ms sustained, application timeouts
Path stateRedundancy and per-path saturationAny path in dead state
DQLEN valueEffective device queue depth in useDQLEN well below HBA default suggests SIOC capping or driver limits

Fixes

Raise DSNRO for shared LUNs

The most common fix on all-flash arrays with multiple VMs per datastore. DSNRO is per-LUN and online-changeable, unlike HBA queue depth, which needs a reboot.

# Set DSNRO per LUN (online, no reboot, takes effect immediately)
esxcli storage core device set -d naa.6005076300810186f00000000000001 -O 64

Caveats:

  • On vSphere 6.5 and later, DSNRO cannot exceed the device’s max queue depth; attempting to set it higher returns an error. Match the HBA default: 64 for QLogic on ESXi 5.0 through 8.0, 128 for software iSCSI.
  • The per-host Disk.SchedNumReqOutstanding advanced option still exists as a global default, but the per-LUN esxcli form above is the recommended method for granular control.
  • DSNRO has no effect when only one world is issuing I/O to the LUN. It only matters when there is contention, which is exactly when you want it higher.

Raise PVSCSI adapter queue depth for I/O-heavy VMs

For a single VM that is saturating its virtual adapter queue, the PVSCSI controller is the right lever. Default per-virtual-disk queue depth is 64 and can be raised to 254, with a corresponding increase to the request ring pages from the default of 8 to 32. Both must be set together; raising only the queue depth does not deliver the full benefit because the ring buffer becomes the limiter. This is a VM-level change that requires a guest power cycle to apply. It helps when the bottleneck is the per-virtual-disk queue, not the VMkernel device queue. If QUED is the problem, raise DSNRO first.

Fix the pathing

If dead paths are the cause, restoring them redistributes I/O and relieves the surviving paths immediately. Verify the multipathing policy: Round Robin spreads I/O across active paths; Fixed and MRU do not aggregate bandwidth on a single LUN. Pathing fixes often resolve queue saturation without any queue-depth tuning, and they should always be the first thing you check when QUED spikes correlate with a path event.

Do not raise the HBA queue depth casually

HBA driver parameters such as ql2xmaxqdepth, lpfc_lun_queue_depth, and fnic_max_qdepth require a host reboot to change, and the defaults are vendor- and driver-version-specific. Raise them only when you have measured that the HBA-level queue is the limiter and DSNRO plus pathing are already correct. Changing these values also changes behaviour for every LUN on that adapter, so the blast radius is wider than a per-LUN DSNRO change.

Prevention

  • Standardize DSNRO on shared all-flash datastores. The default 32 is a legacy spinning-disk value. For modern flash arrays shared across many VMs, set DSNRO to match the HBA default as part of the host build.
  • Monitor path state continuously, not just on alert. A single dead path on a Fixed-policy LUN is a latent queue-saturation incident.
  • Prefer Round Robin for arrays that support it. It uses available paths and reduces the blast radius of a path failure.
  • Track the ACTV-to-DQLEN ratio as a leading indicator. Sustained ACTV above 70% of DQLEN during peak is the early warning before QUED ever appears.
  • Right-size the PVSCSI ring for known I/O-heavy VMs (databases, build agents, search nodes) at provisioning time rather than reactively.
  • Watch for SIOC queue-depth changes. When Storage I/O Control is enabled it can dynamically lower the effective queue depth based on detected congestion. If you see DQLEN moving on its own, SIOC is why.

How Netdata helps

Per-second collection catches queue saturation before it becomes an incident. The signals worth correlating:

  • QUED and ACTV per device, sampled every second. Sustained QUED above zero is the direct signal; ACTV pinned at DQLEN is the precursor. Per-second resolution catches the 30-second bursts that 5-minute vCenter rollups average away.
  • KAVG versus DAVG split. Correlating kernel latency against device latency tells you whether to tune the ESXi stack (KAVG dominant) or engage the storage team (DAVG dominant).
  • Path state alongside latency. A latency spike that coincides with a path going dead is a different incident than a spike with all paths active, and the fix is different.
  • GAVG per virtual disk, joined with guest-side I/O wait. When the guest’s I/O wait tracks GAVG, you have end-to-end confirmation that the queue is the user-visible bottleneck.
  • DQLEN drift detection. If DQLEN changes without an operator action, SIOC or a path event is reshaping the effective queue. Trending DQLEN over time surfaces this before QUED does.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.