The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / elasticsearch / elasticsearch-jvm-heap-high

Operations Guides

Elasticsearch JVM heap usage high: reading the sawtooth and the post-GC floor

Your Elasticsearch alert fires: jvm.mem.heap_used_percent has crossed 75 percent and is holding there. You pull up the graph and see a jagged sawtooth climbing toward the ceiling. The first instinct is to add heap or restart the node. Both are usually wrong.

The sawtooth is normal. Elasticsearch runs on the JVM with a young generation that fills with short-lived objects and empties on young garbage collections. The peak of the tooth is noise. The signal that matters is the post-GC floor: the minimum heap used immediately after a collection. In a healthy node, the floor stays between roughly 30 and 50 percent of max heap, and young GC dominates. When the floor trends upward, old generation objects are accumulating. Old GC pauses stop the world, and once a pause exceeds the cluster fault detection timeout, the master removes the node and triggers shard reallocation. That reallocation places more heap pressure on the survivors, beginning a death spiral.

What this means

Elasticsearch runs on the JVM and typically uses G1GC, which divides the heap into young and old regions. Short-lived objects created during indexing and search live in the young generation and are collected frequently and cheaply. Long-lived structures such as segment metadata, fielddata caches, cluster state, and large aggregation buffers live in the old generation and require stop-the-world collections to reclaim.

Sustained heap_used_percent above 75 percent is a ticket-level concern. Above 85 percent combined with rising old GC time or circuit breaker trips, it is a page. Above 90 percent, the node is at risk of being dropped from the cluster.

Configure Elasticsearch with -Xms equal to -Xmx, and keep the max at or below 26 GB to stay within compressed ordinary object pointers. More heap is not always better; Lucene relies on the operating system page cache, so leave roughly half of physical RAM for the OS.

flowchart TD
    A[Young GC reclaims ephemeral objects] --> B[Post-GC floor stable]
    C[Segment metadata fielddata cluster state] --> D[Post-GC floor rises]
    D --> E[Old GC fires frequently]
    E --> F[Pause exceeds fault detection timeout]
    F --> G[Master removes node]
    G --> H[Shard reallocation]
    H --> I[Survivor heap pressure increases]
    I --> D

Common causes

CauseWhat it looks likeFirst thing to check
Shard count and segment metadatasegments.memory grows linearly with shard count; node holds thousands of shards_cat/nodes?v&h=name,segments.count,segments.memory
Fielddata cache on text fieldsfielddata.memory_size is significant or evictions are nonzero; fielddata breaker trips_nodes/stats/indices/fielddata?fields=*
Cluster state or mapping explosionPending cluster tasks grow; master node heap is high_cluster/stats?filter_path=indices.mappings.field_types
High-cardinality aggregationsHeap spikes during queries; request breaker trips; slow queries in log_nodes/stats/breaker and /_tasks?detailed=true&actions=*search*
Oversized bulk batchesTransient heap spikes; indexing latency rises_nodes/stats/indices/indexing

Quick checks

Run these read-only commands against the cluster to characterize the pressure.

# Heap percent and GC counters per node
curl -s 'http://localhost:9200/_cat/nodes?v&h=name,heap.percent,heap.max,gc.young.count,gc.young.time,gc.old.count,gc.old.time'
# Detailed heap usage and cumulative GC time in millis
curl -s 'http://localhost:9200/_nodes/stats/jvm?filter_path=nodes.*.jvm.mem,nodes.*.jvm.gc'
# Segment metadata heap per node
curl -s 'http://localhost:9200/_cat/nodes?v&h=name,segments.count,segments.memory'
# Fielddata cache size and evictions
curl -s 'http://localhost:9200/_cat/nodes?v&h=name,fielddata.memory_size,fielddata.evictions'
# Circuit breaker estimated sizes and trip counts
curl -s 'http://localhost:9200/_nodes/stats/breaker?filter_path=nodes.*.breakers'
# Active search tasks that may be consuming heap
curl -s 'http://localhost:9200/_tasks?detailed=true&actions=*search*'
# Mapping breadth by field type; sum counts as a proxy for cluster state weight
curl -s 'http://localhost:9200/_cluster/stats?filter_path=indices.mappings.field_types'

How to diagnose it

  1. Confirm the floor is rising. Compare the minimum heap_used_percent observed over a 10-15 minute window, or sample _nodes/stats/jvm shortly after an old GC event. If the minimum is climbing, you have structural accumulation, not a transient spike.

  2. Identify the dominant heap consumer. Query _nodes/stats/indices/segments,fielddata,query_cache,request_cache and compare segments.memory, fielddata.memory_size, and cache sizes. One category will dominate.

  3. If segments.memory is high, check shard density. Use _cat/nodes?v&h=name,segments.count. If a node holds more than a few hundred shards or individual indices show more than 100 segments per shard, segment metadata is the driver. Significant metadata moved off-heap in recent releases, but per-segment-per-field overhead still accumulates.

  4. If fielddata is high, find the offending field. Query _nodes/stats/indices/fielddata?fields=*&filter_path=nodes.*.indices.fielddata.fields. This usually means a text field is being aggregated or sorted. Fielddata is disabled by default on text fields; if it is enabled, that is the problem. Add a keyword sub-field and change the query to use it.

  5. If cache sizes are high, evaluate the working set. Large query or request caches reduce headroom for in-flight operations. Consider whether the cache hit ratio justifies the memory cost, especially on actively written indices where refreshes invalidate cached segments.

  6. If cluster state is large, check mapping breadth. Sum count values from _cluster/stats?filter_path=indices.mappings.field_types. Master nodes need adequate heap to serialize and publish state. If the total field count is growing without bound, you have a mapping explosion.

  7. Correlate with old GC behavior. In _nodes/stats/jvm, check jvm.gc.collectors.old.collection_count and jvm.gc.collectors.old.collection_time_in_millis. If the cumulative time is climbing rapidly, old collections are getting longer or more frequent. For individual pause duration, inspect the node’s GC logs for multi-second collections.

  8. Check for breaker trips. Any delta greater than zero on breakers.parent.tripped means the node is already rejecting operations to protect itself from OOM. breakers.fielddata.tripped points directly to text-field misuse.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
heap_used_percent sustainedInstantaneous memory pressure>75% for more than 5 minutes
Post-GC heap floorStructural accumulation of long-lived objectsRising over hours or days
Old GC collection count and cumulative timeStop-the-world frequency and durationIncreasing trend; individual pauses >5 seconds
segments.memoryPersistent per-segment-per-field overheadGrowing linearly with shard count
fielddata.memory_size and evictionsExpensive text-field aggregations>10% of heap or any evictions
breakers.parent.trippedNode is near OOMAny delta > 0
Pending cluster tasksMaster coordination health>20 tasks or any task older than 30 seconds

Fixes

Shard and segment metadata reduction

Close or delete unused indices and enforce retention with ILM. Warning: closing an index makes it unavailable for search; deleting is irreversible.

Reduce shard count by shrinking existing indices with the _shrink API or reindexing into fewer shards. For read-only indices, force merge to one segment with POST /<index>/_forcemerge?max_num_segments=1. Warning: do not force merge indices that are still receiving writes; the operation is I/O intensive and blocks the index. Target shard sizes between 10 and 50 GB.

Fielddata and mapping fixes

Stop aggregating or sorting on text fields. Use keyword sub-fields instead. If fielddata is explicitly enabled in a mapping, remove it. You can set indices.fielddata.cache.size to place a hard cap, but evictions under that cap are a sign of a mapping problem, not a healthy state.

Query and aggregation tuning

Cancel heavy tasks identified via /_tasks using POST /_tasks/{task_id}/_cancel. Warning: this aborts in-flight user requests.

Reduce aggregation cardinality; high-cardinality terms aggregations consume disproportionate heap. For heavy terms aggregations, benchmark the execution_hint setting (map versus global_ordinals) for your specific workload. The default global_ordinals avoids an in-memory map but must compute ordinals at search time; map can be cheaper for low-cardinality fields with many documents per shard.

Reduce size parameters or replace deep paging with search_after.

Cluster state and master relief

Pause rapid index creation or ILM churn if pending tasks are backing up. Cap field growth with index.mapping.total_fields.limit and index.mapping.depth.limit. If cluster state size is large, deploy dedicated master nodes with sufficient heap; three master-eligible nodes is the standard minimum.

Emergency braking during a cascade

If nodes are already dropping and reallocating, set cluster.routing.allocation.enable: none to stop the rebalancing storm while you fix the root cause. Warning: this stops all shard allocation, including recovery of replicas, and can turn a yellow cluster red if nodes stay offline. Re-enable allocation only after the floor stabilizes.

Lower indices.breaker.fielddata.limit temporarily to force earlier rejection of abusive queries while you fix mappings. Warning: this can cause legitimate aggregations to trip the breaker immediately.

Prevention

  • Monitor the post-GC floor, not just the peak. Alert when the floor rises above 50 percent or trends upward over a week.
  • Track shard count per node and segment count per index. Keep shard counts well below the per-node default limit of 1000, and comfortable below 200 per node if possible.
  • Review mappings before index creation. Disable dynamic mapping or set strict limits to prevent mapping explosions.
  • Size the heap correctly. Set -Xms equal to -Xmx and keep the max at or below 26 GB to preserve compressed OOPs.
  • Leave half of physical RAM for the OS page cache. Do not starve Lucene by overallocating heap.
  • Schedule maintenance operations such as snapshots and force merges outside peak traffic windows.

How Netdata helps

Netdata surfaces jvm.mem.heap_used_percent and GC metrics per node, so you can read the floor without manual GC log inspection. Correlate heap usage with indexing rate, search rate, and thread pool rejections to distinguish ingest pressure from query pressure. Alert on composite conditions such as sustained heap above 85 percent combined with rising old GC cumulative time to suppress noise from normal young-GC oscillation. Long-term retention of per-node segments.memory and fielddata sizes exposes gradual floor rise that point-in-time API checks often miss.

The Netdata solution

Elasticsearch monitoring with Netdata

Netdata monitors Elasticsearch with per-second metrics and ML anomaly detection. Correlate JVM heap pressure, shard counts, disk watermarks, mapping growth, and merge activity with cluster and node health in one view.