The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / elasticsearch / elasticsearch-monitoring-maturity-model

Operations Guides

Elasticsearch monitoring maturity model: from survival to expert

Incidents escalate when teams monitor only the survival layer and miss leading indicators that predict cascades. A rising heap floor, a growing segment count, or a master task backlog surface long before cluster health turns red.

This guide organizes Elasticsearch monitoring into four levels: survival, operational, mature, and expert. Each level adds signals that reduce mean time to detection and prevent composite failure patterns. Build level 1 before going live, level 2 before handling production traffic, and levels 3 and 4 after your first serious incident.

flowchart TD
    L1[Level 1 Survival] -->|add throughput, errors, and saturation| L2[Level 2 Operational]
    L2 -->|add leading indicators and per-index rates| L3[Level 3 Mature]
    L3 -->|add state churn, per-shard stats, and pressure| L4[Level 4 Expert]

Level 1: survival

Level 1 is the minimum viable monitoring. These signals answer one question: is the cluster alive and accepting work? A node that fails to respond on port 9200, a cluster that stays red for more than a few minutes, or a heap above 85 percent with old GC firing are page-worthy conditions. Reachability checks should exercise the HTTP API, not just TCP, because the process can accept connections while the HTTP layer is hung or returning 503. Never operate without these baselines. Even a single-node instance needs reachability and disk checks; a flood-stage watermark blocks writes regardless of cluster size.

SignalWhy it mattersWarning sign
Node reachabilityConfirms the process is alive and accepting HTTP connections.No response on port 9200 for more than 30 seconds.
Cluster health statusGreen, yellow, or red indicates whether all primaries and replicas are assigned.Red sustained longer than 2 minutes; yellow sustained longer than 30 minutes outside of rolling restarts.
JVM heap used percentHeap pressure precedes GC death spirals and OOM kills.Sustained above 75 percent; page if above 85 percent with old GC or breaker trips.
Node countA drop means a node left the cluster and shard reallocation will begin.Any unplanned decrease, or loss of master-eligible quorum.
Disk usage vs. watermarksLow (85%), high (90%), and flood-stage (95%) triggers control allocation and write availability.Above low watermark; immediate attention at flood stage when indices receive read-only blocks.
Indexing rateBaseline confirmation that the cluster is ingesting data.Drop to zero while upstream data sources remain active.
Search rateBaseline confirmation that the cluster is serving queries.Drop to zero while client applications remain active.

Level 2: operational

Level 2 adds the signals that explain why a surviving cluster is slow or failing. Thread pool rejections show which subsystem is saturated: the write pool for indexing, search for queries, and management for background coordination. Latency splits into query and fetch phases so you know whether shards are slow to scan or slow to retrieve source. Unassigned shards and pending tasks reveal coordination and allocation health. ILM execution status catches silent index accumulation that eventually triggers disk watermark cascades. If you cannot identify which node is hot-spotted or why bulk requests return HTTP 429, you are not yet at level 2.

SignalWhy it mattersWarning sign
Thread pool rejections (check /_nodes/stats/thread_pool)A full queue means the node is pushing back. Write rejections return HTTP 429; search rejections fail user queries.Sustained nonzero rate on write, search, or management pools for more than 5 minutes.
Old GC count and timeOld-generation pauses are stop-the-world events that drive latency and node removal. Long pauses trigger the master to mark the node as failed, initiating reallocation.Frequency increasing, or individual pauses exceeding 5 seconds.
Search latency (query and fetch)The slowest shard determines end-to-end latency in the scatter-gather model.Sustained elevation above 2x baseline; query-phase latency rising independently of fetch.
Indexing latencyRising latency under constant load signals disk I/O contention, merge storms, or pipeline overhead.Sustained increase above 2x baseline.
Unassigned shard count and reasonUnassigned primaries mean unavailable data; unassigned replicas mean degraded redundancy.Any unassigned primary; replicas unassigned longer than 30 minutes.
Pending cluster tasks (check /_cluster/pending_tasks)Backlog on the master delays allocation, mapping updates, and index creation. A single slow task can block all subsequent metadata changes.More than 20 tasks, or any task older than 30 seconds.
Merge activity and segment countMerges reclaim deleted documents and improve search speed; too many segments degrade performance and consume heap.Segment count per shard exceeding 100; total segment memory growing on nodes.
Snapshot status and durationBackups that are not completing compromise recoverability.Last successful snapshot older than 2x the scheduled interval; failures accumulating.
Circuit breaker trips (check /_nodes/stats/breaker)The system is rejecting operations to prevent OOM. Repeated request breaker trips under query load often mean aggregations are too expensive.Any trip of the parent breaker; repeated fielddata or request breaker trips.
ILM execution statusStuck policies cause index and shard accumulation that eventually drives disk and heap pressure.Indices stuck in a phase for longer than the expected transition window.

Level 3: mature

Level 3 shifts from reactive to predictive. The heap sawtooth floor, not the peak, is the true indicator of memory pressure. Watch the post-GC minimum over a 24-hour window; a floor climbing toward 50 percent of max heap means long-lived objects are leaking or caches are unbounded. Per-index rates expose hot-spotting that cluster-wide averages hide. Cluster state size and version churn warn that the master is approaching instability long before elections flap. Every mapping update or dynamic index creation increments the state version and forces a publish to all nodes. Replica lag measured by sequence number gaps shows redundancy degrading in real time. These signals require historical trending and lower alert thresholds.

SignalWhy it mattersWarning sign
Heap sawtooth floorThe post-GC minimum heap is the best leading indicator of long-lived object accumulation.Floor trending upward over days; approaching 50 percent of max heap.
Fielddata cache size and evictionsFielddata on text fields loads terms into heap and should be near zero in modern deployments.Size above 10 percent of heap or any evictions occurring.
Cluster state size and versionEvery mapping, index, and alias inflates the state every node holds in heap. Rapid version increments signal churn.State consuming more than 5 percent of master heap; version incrementing faster than 10 per second sustained.
Translog size and uncommitted operationsLarge translogs extend recovery time and indicate flush problems.Uncommitted size well above the configured flush threshold or growing monotonically.
Refresh and flush timesSlow refresh creates segments slowly; slow flush delays durability and truncates translog.Average refresh time above 1 second or flush time above 30 seconds sustained.
Shard recovery activityRecovery competes with production traffic for disk I/O and network bandwidth.Recovery stalled at the same percentage for longer than 30 minutes.
Per-index indexing and search ratesCluster-wide averages hide hot-spotted indices or shards.Asymmetry where one index dominates node load or one shard carries disproportionate traffic.
Per-node segment memorySegment metadata lives in heap and scales with segment count and field count.segments.memory consuming more than 10 percent of node heap.
File descriptor utilizationEach segment consists of multiple files; exhaustion causes I/O and connection failures.Above 80 percent of the configured maximum.
Replica lag (sequence number gap)A growing gap between primary and replica means reduced redundancy and potential sync failures.Global checkpoint trailing max sequence number by more than 10,000 operations and growing.

Level 4: expert

Level 4 is for teams that have debugged enough incidents to know that averages lie. Per-shard segment distributions reveal the single oversized shard behind a latency spike. Hot threads and task cancellation isolate the exact query or merge consuming CPU. Indexing pressure stats break down memory usage by coordinating, primary, and replica stages. Adaptive replica selection metrics show which nodes the cluster is already avoiding. These signals are verbose and expensive to collect continuously, so sample them during incidents or bake them into automated diagnostics that fire when lower-level thresholds breach.

SignalWhy it mattersWarning sign
Cluster state version churnRapid version increments indicate unstable routing or excessive metadata changes.Version rate exceeding 10 per second sustained without planned topology changes.
Per-shard segment count distributionAverages hide individual problem shards with hundreds of segments.Outlier shards above 100 segments while siblings remain low.
Merge throttle timemerges.total_throttled_time_in_millis indicates I/O pressure forced Lucene to slow merges.Throttle time growing while segment count also grows.
OS page cache effectivenessElasticsearch relies on the kernel page cache, not heap, for segment access.Available memory for page cache shrinking relative to total segment data size.
Adaptive replica selection statsARS routes searches away from struggling nodes; monitoring it reveals hidden hot-spotting.Specific nodes consistently marked as poor targets by selection heuristics.
Indexing pressure statsMemory-based backpressure at coordinating, primary, and replica stages warns before heap failure.Current bytes sustained above 80 percent of the 10 percent heap limit, or rejections increasing.
Hot threadsCPU consumption breakdown by thread during incidents identifies what the JVM is doing right now.Persistent hot threads in merge, search, or OTHER_CPU (GC) categories.
Long-running tasksSearches or bulk operations that exceed expected duration hold resources and block queues.Tasks running longer than 30 seconds on a low-latency cluster.
Total mapping field countUnbounded field growth from dynamic mapping inflates cluster state and heap.Field count growing without bound or approaching index.mapping.total_fields.limit.
Snapshot incremental sizeTracks backup storage growth and segment churn between snapshots.Incremental size spiking after force merges or indicating unexpected data growth.

Advancing through these levels is not about collecting more metrics for their own sake. It is about reducing the time between symptom and mechanism. Level 1 tells you that something broke. Level 2 tells you which subsystem broke. Level 3 tells you it is breaking before it fails. Level 4 tells you exactly which shard, segment, or query is responsible.

How Netdata helps

Netdata collects Elasticsearch metrics from the JSON stats APIs and correlates them with system-level data. Elasticsearch performance is inseparable from OS resources.

  • Per-node heap and GC correlation: Netdata surfaces JVM heap alongside OS memory, letting you distinguish heap pressure from page cache starvation without switching tools.
  • Thread pool saturation visibility: Netdata tracks queue depth and rejection rates per pool, so you can spot the transition from queuing to rejection before clients fail.
  • Disk and I/O context: Disk watermark alerts in Elasticsearch are clearer when paired with OS-level I/O wait and throughput, showing whether saturation is from merges, recovery, or co-located workloads.
  • Cluster state and indexing pressure: Netdata indexes metrics like cluster health and indexing pressure bytes, enabling dashboards that show master-level signals alongside data-node saturation.
The Netdata solution

Elasticsearch monitoring with Netdata

Netdata monitors Elasticsearch with per-second metrics and ML anomaly detection. Correlate JVM heap pressure, shard counts, disk watermarks, mapping growth, and merge activity with cluster and node health in one view.