The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kubernetes / kubernetes-monitoring-checklist

Operations Guides

Kubernetes Monitoring Checklist: The Signals Every Production Cluster Needs

This article is a reference checklist for senior engineers who are wiring up, auditing, or hardening monitoring for a production Kubernetes cluster. It assumes you already understand the control plane architecture and focuses on what to collect, where to find it, and which symptoms matter. Use it during greenfield instrumentation, post-incident gap analysis, or routine health audits.

The signals are grouped by domain. Each entry leads with a short noun phrase, followed by one sentence explaining why it matters, and a concrete warning sign to alert on. Thresholds are drawn from upstream SLOs, kubelet defaults, and etcd operational limits documented in the Kubernetes source and production playbooks. If you run a managed service such as EKS, GKE, or AKS, treat control-plane metrics as provider-mediated; many etcd and API server internals are opaque in those environments.

Control Plane & etcd

API server liveness. Confirms the kube-apiserver process is accepting connections. Warning sign: non-200 response or timeout greater than 5 seconds on /livez.

API server readiness. Verifies etcd connectivity, informer sync, and post-start hooks. Warning sign: /readyz fails while /livez passes, indicating initialization deadlock or etcd loss.

etcd leader stability. Raft leadership churn directly blocks writes and causes brief outages. Warning sign: etcd_server_leader_changes_seen_total increments more than once per hour without maintenance activity.

etcd WAL fsync latency. Every etcd write fsyncs to disk; slow storage cascades into API latency. Warning sign: etcd_disk_wal_fsync_duration_seconds p99 greater than 100 ms sustained.

etcd database size. Approaching the quota triggers a NOSPACE alarm and makes the cluster read-only. Warning sign: etcd_debugging_mvcc_db_total_size_in_bytes or etcd_mvcc_db_total_size_in_bytes greater than 80 percent of --quota-backend-bytes.

API request latency by verb. Elevated mutating latency stalls controllers, kubectl, and CI pipelines. Warning sign: apiserver_request_duration_seconds p99 greater than 1 s for POST/PUT/PATCH sustained, or LIST p99 greater than 30 s.

Admission webhook latency. Synchronous webhook calls add directly to mutating request latency. Warning sign: apiserver_admission_webhook_admission_duration_seconds p99 greater than 200 ms for any webhook with failurePolicy: Fail.

API Priority and Fairness queue depth. Queued requests indicate a priority level is saturated and critical traffic may be delayed. Warning sign: apiserver_flowcontrol_current_inqueue_requests greater than 0 for the system or leader-election priority levels.

API server error rate. 5xx errors indicate etcd, webhook, or internal failures; 429s indicate APF throttling. Warning sign: apiserver_request_total with code=~"5.." or code="429" sustained above baseline.

Inflight requests. Approaching the hard limit causes 429 rejections and client retry storms. Warning sign: apiserver_current_inflight_requests greater than 80 percent of --max-requests-inflight or --max-mutating-requests-inflight.

Watch event throughput. High event rates indicate rapid cluster churn that can overload informers. Warning sign: apiserver_watch_events_total spiking greater than 10 times baseline.

Node & Kubelet Health

Node Ready condition. The kubelet’s self-reported ability to run pods. Warning sign: Ready=False or Ready=Unknown for more than 1 minute.

PLEG relist latency. Slow container runtime queries precede NotReady transitions. Warning sign: kubelet_pleg_relist_duration_seconds p99 greater than 10 s or approaching the 3-minute unhealthy threshold.

Kubelet sync loop duration. Slow reconciliation delays pod creation, termination, and status updates. Warning sign: kubelet_pod_worker_duration_seconds{operation_type="sync"} p99 greater than 30 s sustained.

Container runtime connectivity. Runtime disconnection blocks all container operations on the node. Warning sign: crictl info hangs or fails, or the kubelet /healthz endpoint reports runtime unhealthy.

Image pull duration. Slow pulls extend pod startup time and scale-out response. Warning sign: kubelet_image_pull_duration_seconds p99 greater than 2 minutes for typical images.

Node memory pressure. Triggers pod eviction and kernel OOM kills. Warning sign: MemoryPressure=True or available memory below 100 Mi (default hard eviction threshold).

Node disk pressure. Blocks new image pulls and triggers pod eviction. Warning sign: DiskPressure=True, or nodefs utilization above 90 percent or imagefs utilization above 85 percent.

PID pressure. Exhaustion prevents fork and blocks container startup. Warning sign: PIDPressure=True, or node PID usage above 80 percent of /proc/sys/kernel/pid_max.

Kubelet certificate expiration. Expired client certificates break API server authentication. Warning sign: kubelet_certificate_manager_client_ttl_seconds below 7 days, or any rotation error counter incrementing.

Kubelet error rate. Logs surface runtime, API, or certificate problems before they become node failures. Warning sign: greater than 50 errors per hour, or any panic message.

Workloads & Scheduling

Pod phase distribution. Pending or Unknown pods indicate scheduling failures or lost nodes. Warning sign: pods in Pending longer than 10 minutes, or Unknown phase increasing.

CrashLoopBackOff. Persistent container crashes indicate application bugs, OOMs, or misconfiguration. Warning sign: any pod in CrashLoopBackOff for more than 5 minutes.

Container restart count. Rising restarts signal instability before CrashLoopBackOff. Warning sign: restart count increasing by more than 5 in 10 minutes for production workloads.

OOM kill events. Memory limits exceeded or node-level pressure kills containers. Warning sign: lastState.terminated.reason=OOMKilled, or kernel dmesg showing OOM kills.

CPU throttling. CFS quota exhaustion degrades application latency. Warning sign: container_cpu_cfs_throttled_periods_total ratio greater than 25 percent for latency-sensitive pods.

Scheduler pending pods. Growing unschedulable queue means capacity or constraint exhaustion. Warning sign: scheduler_pending_pods with queue="unschedulableQ" growing for more than 5 minutes.

Controller workqueue depth. Backlog indicates the control plane is falling behind on reconciliation. Warning sign: workqueue_depth greater than 100 sustained for core controllers such as deployment or node.

Deployment rollout health. Stuck rollouts leave workloads under-replicated. Warning sign: readyReplicas less than spec.replicas or ProgressDeadlineExceeded for more than 5 minutes.

Networking & DNS

kube-proxy sync duration. Long syncs stale rules and hold the iptables lock in iptables mode. Warning sign: kubeproxy_sync_proxy_rules_duration_seconds p99 greater than 10 s, or approaching the sync period.

Conntrack table utilization. A full table silently drops new connections across all node traffic. Warning sign: nf_conntrack_count greater than 90 percent of nf_conntrack_max, or the drop or early_drop counter incrementing in conntrack -S.

kube-proxy healthz. Failure indicates API server disconnect or rule programming failure. Warning sign: /healthz returning 503 for more than 1 minute.

Service endpoint readiness. Zero endpoints means the Service has no healthy backends. Warning sign: zero available endpoints for a production Service, or endpoint count dropping more than 50 percent in 5 minutes.

CoreDNS health. DNS failures cascade to all inter-service traffic. Warning sign: CoreDNS pods not Ready, coredns_dns_responses_total with rcode="SERVFAIL" greater than 1 percent, or p99 latency greater than 500 ms.

Network policy enforcement. Silent drops from failing CNI plugins or misconfigured policies. Warning sign: connection timeouts between pods that should be allowed, with CNI pod restarts or policy sync delays.

Storage & Volumes

PVC binding status. Pending PVCs block pod startup. Warning sign: any production PVC in Pending for more than 5 minutes.

PV utilization. Full volumes cause write failures and potential corruption. Warning sign: kubelet_volume_stats_used_bytes greater than 90 percent of capacity, or inode usage greater than 90 percent.

Volume mount latency. Stuck mounts block pods in ContainerCreating indefinitely. Warning sign: storage_operation_duration_seconds greater than 2 minutes, or pods stuck ContainerCreating with mount events.

Security & Certificates

Control plane certificate expiration. Expiry breaks all TLS communication between components. Warning sign: any control plane, etcd, or kubelet certificate less than 30 days until expiration.

API server authentication failures. Mass 401s indicate certificate rotation failure or attack. Warning sign: apiserver_request_total with code="401" spiking greater than 10 times baseline.

RBAC modification rate. Unexpected privilege grants indicate compromise or misconfiguration. Warning sign: new cluster-admin bindings or clusterrolebindings created outside change windows.

Anonymous API requests. Successful anonymous access indicates misconfiguration or exposure. Warning sign: system:anonymous requests returning 200 for non-health endpoints.

Privileged container creation. Host-namespace access increases attack surface and escape risk. Warning sign: pods with privileged: true, hostNetwork: true, or hostPID: true outside system namespaces.

Audit log gaps. Missing logs break forensics and may indicate backend failure. Warning sign: unexplained gaps greater than 5 minutes in audit output during active cluster use.

How Netdata Helps

  • Netdata collects kubelet, container, and node metrics in real time without central query latency.
  • It correlates API server latency with etcd disk latency and admission webhook latency on the same timeline.
  • It tracks cgroup-level CPU throttling, memory usage, and disk I/O per pod to pinpoint noisy neighbors.
  • It surfaces conntrack utilization, kube-proxy sync duration, and CoreDNS latency alongside workload health so you can move from symptom to cause without switching tools.
The Netdata solution

Kubernetes monitoring with Netdata

Netdata monitors Kubernetes with per-second metrics across the control plane, nodes, and every pod, with ML anomaly detection and zero per-pod configuration. Correlate API-server and etcd latency, kubelet PLEG stalls, scheduling pressure, and OOMKills in one place.