The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kubernetes / kubernetes-volume-mount-failures

Operations Guides

Kubernetes pod stuck on volume mount: CSI, permissions, and timeouts

You scale a StatefulSet and the new pods sit in ContainerCreating for ten minutes. The node is Ready. The CSI driver pods are running. kubectl describe shows no FailedMount events, yet the containers never start. The absence of volume events often misleads operators into checking image registries or resource quotas instead of the storage path. The kubelet volume manager is blocked somewhere between attach and mount, and Kubernetes will not retry fast enough to hide the problem.

Unlike image pull errors or scheduling failures, stuck volume mounts often leave the node looking healthy. Existing pods keep running. Only workloads that need storage fail to start, which makes the symptom easy to blame on the application instead of the storage path.

What this means

The control-plane attach/detach controller reconciles volume attachment, while the kubelet volume manager reconciles node-local volume state and mounts. For CSI volumes, the kubelet calls node RPCs to stage and publish the volume. Each operation holds a goroutine. If an operation hangs, that goroutine blocks. When enough hang, the volume manager saturates, and subsequent pods with volumes stall while the node stays Ready.

One hung NFS mount or a cloud volume still finishing force-detach can queue up and delay every new stateful pod on the node. Additionally, the kubelet applies fsGroup ownership changes recursively during mount; on large volumes this can stall pod startup without producing FailedMount events.

Common causes

CauseWhat it looks likeFirst thing to check
CSI driver unavailable on the nodeEvents mention RPC errors or driver name; mount setup failsCSI node plugin pod on the target node
Stale VolumeAttachment or multi-attachEvents mention multi-attach errors or attach waits indefinitelyVolumeAttachment objects and which node owns them
Permission or fsGroup mismatchContainer exits with permission denied on the volume path, or mount fails with access errorsPod securityContext fsGroup and storage backend ownership
NFS or network storage hangNo FailedMount events; pod ContainerCreating for more than 10 min; node ReadyManual mount test from the node to the storage endpoint
Cloud provider detach delayAttach operation exceeds 2 minutes after previous node lossCloud provider console or volume status for pending detach

Quick checks

# Confirm the pod phase is ContainerCreating
kubectl get pod $POD_NAME -o jsonpath='{.status.phase}'

# Inspect events for MountVolume, AttachVolume, or RPC errors
kubectl describe pod $POD_NAME | grep -A 20 Events

# Identify if the volume is attached to a different node
kubectl get volumeattachment -o wide

# Verify the CSI node plugin is present and running
kubectl get pods -n $CSI_NAMESPACE -l app=$CSI_DRIVER --field-selector spec.nodeName=$NODE_NAME

# Scan cluster-wide volume errors
kubectl get events --all-namespaces | grep -E "FailedMount|AttachVolume|MountVolume"

# Read kubelet logs for low-level volume and CSI errors
journalctl -u kubelet --since "30 minutes ago" | grep -i "mount\|volume\|csi"

# Check storage operation latency from kubelet metrics
kubectl get --raw /api/v1/nodes/$NODE_NAME/proxy/metrics | grep storage_operation_duration_seconds

# Verify fsGroup is configured when containers run as non-root
kubectl get pod $POD_NAME -o jsonpath='{.spec.securityContext.fsGroup}'

# Inspect currently mounted volumes under the kubelet directory
mount | grep kubelet

# Test NFS reachability directly from the node with a timeout to avoid indefinite hangs
timeout 10 mount -t nfs $SERVER:$PATH /tmp/testmnt
# Unmount after testing: umount /tmp/testmnt

# Check CSI node plugin logs for RPC timeouts or driver errors
kubectl logs -n $CSI_NAMESPACE -l app=$CSI_DRIVER --field-selector spec.nodeName=$NODE_NAME --tail=200 | grep -iE "error|timeout|rpc"

How to diagnose it

  1. Determine whether the delay is in attach or mount. Look at kubectl describe pod events and scroll past image pull and scheduling messages to find volume entries. If you see AttachVolume.Attach with a timestamp older than two minutes, the problem is attach. If you see MountVolume.SetUp without completion, the problem is mount.
  2. Inspect VolumeAttachment objects. Compare each VolumeAttachment’s .spec.nodeName against the pod’s .spec.nodeName. If the volume is attached to a different node, flag it as stale for cleanup after you verify the previous workload has stopped.
  3. Verify the CSI node driver pod on the target node. If the pod is not running, volume operations cannot proceed.
  4. Check kubelet storage operation metrics. As rough thresholds, local volumes should mount in under one second, block devices under ten seconds, network volumes under sixty seconds, and cloud attach operations under two minutes. Operations exceeding these indicate a backend hang or slowness.
  5. Examine permission errors. If the container starts but cannot write to the mount path, compare the pod’s securityContext.fsGroup and runAsUser against the filesystem ownership on the storage backend. Correct mismatches in the pod spec.
  6. For network storage, test reachability from the node directly. If a manual mount command hangs, the issue is outside Kubernetes. If it succeeds but the pod still stalls, the problem is likely in the kubelet volume manager or CSI driver.
  7. Review kubelet logs for RPC errors from CSI drivers or mount syscall failures. RPC errors with codes like Internal or DeadlineExceeded point to driver-specific bugs or resource exhaustion.
  8. Check for volume manager saturation. If multiple pods on the same node are stuck in ContainerCreating and the node is Ready, look for the first hung operation. Resolving that operation often unblocks the rest.
flowchart TD
    A[Pod stuck in ContainerCreating] --> B{Events show attach or mount?}
    B -->|Attach timeout / Multi-Attach| C[Check VolumeAttachment objects]
    B -->|MountVolume.SetUp failed / RPC error| D[Check CSI driver pod health and logs]
    B -->|No events / stale for >10 min| E[Test storage reachability from node]
    C -->|Attached to wrong node| F[Delete stale VolumeAttachment]
    C -->|Cloud attach >2 min| G[Check cloud provider detach status]
    D -->|CSI pod unhealthy| H[Restart CSI node plugin]
    D -->|RPC permission error| I[Check pod fsGroup and SELinux context]
    E -->|Mount hangs| J[Review NFS mount options or network path]
    E -->|Mount succeeds| K[Check kubelet storage operation metrics]
    F --> L[Wait for controller to reattach and verify pod startup]
    G --> L
    H --> L
    I --> L
    J --> L
    K --> L

Metrics and signals to monitor

SignalWhy it mattersWarning sign
storage_operation_duration_secondsReveals attach and mount latency before pods visibly stallp99 exceeds 30 s for network volumes or 2 min for cloud attach
storage_operation_duration_seconds{status="fail-unknown"}Tracks volume operations that completed with an unknown failureSustained rate increase over a 5-minute window
volume_manager_total_volumesCount of volumes the kubelet is managingSudden spike correlating with ContainerCreating pods
Pod phase ContainerCreating durationDirect symptom of stuck mountsDuration > 5 minutes
Node Ready conditionNodes can remain Ready while the volume manager is blockedReady=True but volume-mounted pods fail to start
kubelet_pod_worker_duration_secondsSlow reconciliation can indicate blocked volume operationsSustained > 30 s
kubelet_runtime_operations_duration_secondsDistinguishes runtime slowness from storage issueslist_containers p99 > 5 s
VolumeAttachment object ageStale attachments block reattachmentOlder than the cloud provider’s normal detach window

Fixes

If the cause is CSI driver failure

Restart the CSI node plugin pod on the affected node. Check the CSI controller and node plugin logs for RPC errors such as Internal, DeadlineExceeded, or Unavailable. These indicate the driver cannot reach the storage backend, the node plugin socket is unreachable, or the driver is exhausting its memory limits. If the driver is crashing due to resource pressure, increase its memory limit or move it to a node pool with headroom. Do not restart kubelet until the driver is healthy.

If the cause is stale VolumeAttachment or multi-attach

Delete the stale VolumeAttachment object:

kubectl delete volumeattachment <name>

Warning: Only delete a VolumeAttachment after you verify the volume is no longer in use on the previous node. Deleting it while the volume is still mounted can cause data corruption or unclean filesystem state; with a matching pod or node intent, the attach/detach controller will recreate the required attachment. For cloud-backed volumes after a node failure, allow the cloud provider’s force-detach window to complete before manually detaching the volume, or detach it through the provider’s console if the cluster is stuck.

If the cause is permission or ownership mismatch

Add or correct securityContext.fsGroup in the pod spec to align with the group owning the files on the backing storage. For SELinux enforcing environments, ensure the volume context is compatible or configure seLinuxOptions in the securityContext. Recreate the pod to apply changes.

If the cause is network storage hang

For NFS, prefer mount options that prevent indefinite hard hangs, such as soft and a reasonable timeo value, while accepting that soft can return I/O errors to the application if the server becomes unreachable. If a mount is already hung, identify the blocking process before acting. Do not kill kubelet or the mount helper without first assessing workload impact; an unclean termination can leave the volume in an inconsistent state and prevent clean unmount. Restarting the kubelet without fixing the storage endpoint rarely helps and can extend recovery time.

If the cause is cloud API throttling or detach delay

Reduce the rate of new volume attachments to stay within the cloud provider’s API limits and volume queue depths. Verify that the previous node has fully released the volume before rescheduling stateful pods. If a node was preempted or terminated, expect a forced-detach delay.

Prevention

  • Monitor storage_operation_duration_seconds and alert on p99 thresholds before pods stall.
  • Set fsGroup and expected runAsUser explicitly in pod specs for volumes written by non-root containers.
  • Configure NFS and network storage mounts with soft and timeo to avoid hard hangs that saturate the volume manager.
  • Keep CSI driver and sidecar versions current, and monitor their pod health independently from kubelet health.
  • Maintain node disk and inode headroom; disk pressure slows image pulls and can indirectly delay volume reconciliation.
  • Track the gap between running and desired pod counts on each node to detect volume manager saturation early.
  • Document your storage backend’s attach limits and enforce them through CSI driver flags or scheduler constraints.
  • Test storage endpoint fail-over and recovery procedures in staging so that hang behavior is understood before it happens in production.

How Netdata helps

  • Correlates storage_operation_duration_seconds with node disk I/O wait and network latency to distinguish a slow storage backend from a saturated kubelet volume manager.
  • Tracks kubelet sync loop and PLEG relist latency alongside pod start duration to identify whether volume mounts are the bottleneck during startup storms.
  • Alerts on node filesystem usage and inode consumption before disk pressure triggers evictions that complicate volume recovery.
The Netdata solution

Kubernetes monitoring with Netdata

Netdata monitors Kubernetes with per-second metrics across the control plane, nodes, and every pod, with ML anomaly detection and zero per-pod configuration. Correlate API-server and etcd latency, kubelet PLEG stalls, scheduling pressure, and OOMKills in one place.