The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / elasticsearch / elasticsearch-mapping-explosion-dynamic-mapping

Operations Guides

Elasticsearch mapping explosion: dynamic mapping, cluster state bloat, and master pressure

When an Elasticsearch cluster ingesting unstructured JSON starts accumulating thousands of mapped fields per index, every node pays for it in heap. Mappings are part of the cluster state, which the elected master serializes and publishes to every node on each change. An index mapping with tens of thousands of fields means a larger cluster state blob resident in every node’s JVM heap, longer publication times, and mounting pressure on the master.

The first visible symptom is usually not a mapping error. It is rising heap usage across all nodes with no corresponding increase in data volume or query load. The cluster state version counter climbs rapidly as dynamic mapping registers new fields for each unique key in incoming documents. Pending cluster tasks accumulate because the master cannot process metadata changes fast enough. Eventually, the master becomes unstable, triggering elections that stall all writes and administrative operations.

If you catch it early, the fix is a mapping configuration change and a data pipeline redesign. If you catch it late, you are already in a heap pressure cascade or a master instability incident, and the mapping explosion is the root cause hiding behind the symptoms.

What this means

Dynamic mapping is the Elasticsearch feature that automatically infers field types from incoming documents. When a document contains a key that does not exist in the index mapping, Elasticsearch creates a new field definition for it on the fly. This is convenient for exploration and prototyping but dangerous for production workloads that ingest unstructured or semi-structured JSON.

The canonical failure pattern: an index receives documents with high-cardinality key-value pairs such as Kubernetes labels, user-defined tags, or arbitrary JSON payloads from third-party webhooks. Each unique key becomes a new mapped field. An index that started with 50 fields can accumulate 5,000 or 50,000 fields over weeks or months of operation.

The safeguard that should catch this is index.mapping.total_fields.limit, which defaults to 1000. When a document would push the field count past this threshold, Elasticsearch rejects the indexing request with an error like Limit of total fields [1000] has been exceeded. But this limit only fires if it is in effect and has not been raised. Many teams raise it to 100,000 or higher when they hit it the first time, not understanding that the problem is not the limit but the data model.

A critical detail: the field limit counts all mappers, not just leaf fields. A dotted path like host.os.name creates three mappers: one for host (object), one for host.os (object), and one for host.os.name (leaf). Multi-fields, such as adding a .keyword sub-field to a text field, each count individually. This means the effective field count can be much higher than the number of unique keys in your source data.

The downstream damage follows a predictable cascade:

flowchart TD
    A["Unstructured JSON ingested
(K8s labels, user tags)"] --> B["Dynamic mapping
creates new fields"] B --> C["Mapping grows to
thousands of fields"] C --> D["Cluster state balloons
stored in heap on every node"] D --> E["Heap pressure rises
cluster-wide"] D --> F["Master overwhelmed by
state publication overhead"] E --> G["Old GC frequency increases
nodes at risk of removal"] F --> H["Pending tasks backlog
master elections, stalled ops"] G --> H

Every node holds the full cluster state in JVM heap. A 200MB cluster state on a 30GB heap means 200MB consumed on every single node, not just the master. When mappings are large, the master must serialize and publish that state on every mapping change, and every node must deserialize and apply it. Under sustained mapping churn from dynamic mapping, this creates a feedback loop: the master falls behind, pending tasks pile up, and if the master becomes unresponsive during a large state publication, a new election occurs.

Common causes

CauseWhat it looks likeFirst thing to check
Unstructured JSON with dynamic mapping enabledField count grows steadily; cluster stats show high total field countIndex settings for dynamic and total_fields.limit
Raised total_fields.limit without fixing the data modelLimit is set to 10,000 or higher; field count is close to the new ceilingWhether the root cause (uncontrolled keys) was ever addressed
Nested objects with deep pathsEach document path creates multiple mappers; field count far exceeds unique key countWhether the data uses deeply nested objects that could be flattened
Index templates with overly broad dynamic templatesNew indices immediately start with large mappings; field count spikes on rolloverDynamic template definitions in composable index templates
System indices with restrictions.watches or other system indices hit the limit and cannot be updatedWhether the affected index is a system index with restricted settings

Quick checks

# Check total field count across all indices (primary indicator of mapping explosion)
curl -s 'http://localhost:9200/_cluster/stats?filter_path=indices.mappings' | python3 -m json.tool

# Check cluster state version churn rate (high churn = frequent metadata changes)
curl -s 'http://localhost:9200/_cluster/state?filter_path=version'

# Check pending cluster tasks (backlog indicates master is overwhelmed)
curl -s 'http://localhost:9200/_cluster/pending_tasks?pretty'

# Check master stability (run multiple times; master node ID should not change)
curl -s 'http://localhost:9200/_cat/master?v'

# Check per-index segment counts (high segment count correlates with large mappings)
curl -s 'http://localhost:9200/_cat/indices?v&h=index,docs.count,pri.segments.count&s=docs.count:desc' | head -20

# Check heap usage across all nodes (uniform pressure suggests cluster state bloat)
curl -s 'http://localhost:9200/_cat/nodes?v&h=name,heap.percent,heap.max,node.role'

# Check segment memory (high segment memory relative to heap = too many fields per segment)
curl -s 'http://localhost:9200/_cat/nodes?v&h=name,segments.count,segments.memory,heap.percent'

# Check the current total_fields.limit setting on a specific index
curl -s 'http://localhost:9200/<index>/_settings?filter_path=*.index.mapping.total_fields.limit'

# Check the dynamic mapping setting on a specific index
curl -s 'http://localhost:9200/<index>/_mapping?pretty' | grep dynamic

How to diagnose it

  1. Confirm the field count is abnormal. Run GET /_cluster/stats?filter_path=indices.mappings and examine the total field count. A healthy cluster typically has field counts in the hundreds or low thousands per index. Anything above 5,000 fields in a single index warrants investigation. Field count growing without bound is a leading indicator of cluster state problems.

  2. Identify which indices have the largest mappings. Use GET /<index>/_mapping on suspected indices. Look for patterns of auto-generated field names: UUIDs, timestamps, or arbitrary strings in field names that indicate dynamic mapping on unstructured data.

  3. Check the cluster state version churn rate. Sample GET /_cluster/state?filter_path=version twice, 30 seconds apart. If the version is incrementing rapidly (more than 10 times per second sustained), the cluster state is being modified constantly, which is typical of active mapping explosion.

  4. Correlate heap pressure with field count growth. Check GET /_cat/nodes?v&h=name,heap.percent,segments.memory. If heap is elevated across all nodes (not just one or two), and segments.memory is high, the cluster state and segment metadata from field-heavy mappings are likely consuming heap.

  5. Check pending cluster tasks. Run GET /_cluster/pending_tasks. A healthy cluster has zero or near-zero pending tasks. More than 20 tasks, or any task older than 30 seconds, indicates the master is falling behind on metadata processing.

  6. Verify master stability. Run GET /_cat/master?v several times over a few minutes. The master node ID should not change outside planned maintenance. Frequent master elections mean the cluster is near a major incident.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Total field count (indices.mappings)Directly measures mapping explosion severityGrowing without bound, or exceeding 1000 per index
Cluster state version churnIndicates how frequently metadata changes are publishedMore than 10 increments per second sustained
Pending cluster tasksShows whether the master can keep up with state changesMore than 20 tasks or any task pending more than 30 seconds
JVM heap used percent on all nodesCluster state lives in heap on every nodeSustained above 75%, especially with uniform pressure across nodes
Master node CPU and heapMaster bears the cost of serializing and publishing stateSustained CPU above 80% or heap above 75% on master node
Segment memory per nodeField-heavy mappings increase per-segment metadata overheadSegment memory growing as a proportion of total heap
Indexing failures (index_failed)Documents rejected when total_fields.limit is reachedSudden spike from near-zero

Fixes

Set dynamic mapping to strict or false

The most direct fix is to stop the source of new field creation. Set dynamic: strict to reject documents with unknown fields (the indexing request fails and the client must handle the error), or dynamic: false to silently ignore unknown fields (they are stored in _source but not indexed or searchable).

# Set dynamic mapping to strict on an existing index
curl -X PUT 'http://localhost:9200/<index>/_mapping' \
  -H 'Content-Type: application/json' \
  -d '{"dynamic": "strict"}'

With dynamic: strict, any document containing a field not in the mapping is rejected with a mapper parsing exception. This prevents new fields from being created but requires your ingest pipeline to either drop or transform unknown keys before they reach Elasticsearch.

dynamic: false is less disruptive. Documents are accepted, but unknown fields are not added to the mapping. They remain in _source and can be retrieved, but they cannot be searched or aggregated. This is often the right choice for indices receiving unpredictable JSON payloads where you want to preserve the raw data without indexing every field.

Tradeoff: Both settings require you to explicitly define the fields you need. If your application depends on searching arbitrary keys, you need a different approach, such as the flattened type described below.

Cap field count with total_fields.limit

If the limit has been raised, lower it back to a sane value. The default of 1000 exists for a reason. If your legitimate field count exceeds 1000, you likely have a data modeling problem.

# Check current limit
curl -s 'http://localhost:9200/<index>/_settings?filter_path=*.index.mapping.total_fields.limit'

# Set the limit (does not retroactively remove existing fields)
curl -X PUT 'http://localhost:9200/<index>/_settings' \
  -H 'Content-Type: application/json' \
  -d '{"index.mapping.total_fields.limit": 1000}'

Important: Lowering the limit does not remove existing fields from the mapping. It only prevents new fields from being added beyond the limit. To reduce the existing field count, you must reindex into a new index with a controlled mapping.

Related limits worth reviewing:

  • index.mapping.depth.limit (default 20): maximum depth of nested objects in a single document.
  • index.mapping.nested_fields.limit (default 50): maximum number of distinct nested field types.

Another related setting: index.mapping.total_fields.ignore_dynamic_beyond_limit (default false). When set to true, dynamic fields beyond the limit are silently dropped and the document is accepted instead of rejected. For indices in the logsdb index mode, this defaults to true; the standard and time_series modes keep the false default. This prevents indexing failures but means data is silently lost from the indexed mapping.

Use the flattened type for dynamic payloads

The flattened field type stores an entire JSON object as a single field. All keys and values are indexed as keyword terms under one mapper, regardless of how many unique keys the JSON contains. This is the recommended approach for high-cardinality dynamic payloads where you need basic term query capability but do not need full per-field type mapping.

{
  "mappings": {
    "properties": {
      "labels": {
        "type": "flattened"
      }
    }
  }
}

With this mapping, a document containing "labels": {"env": "prod", "team": "platform", "service": "api"} creates one field mapper instead of three. You can query labels.env, labels.team, and similar paths, but all values are treated as keywords. No numeric or date inference occurs.

Tradeoff: Flattened fields support term, prefix, and wildcard queries on sub-fields, but do not support full-text search, range queries on numeric values, or sorting on individual sub-fields within the flattened object. All values are indexed as keywords regardless of their original type.

Use subobjects: false (ES 8.3+)

On Elasticsearch 8.3 and later, the subobjects: false mapping option collapses dotted field paths into a single literal field name rather than creating intermediate object mappers for each path segment. For data with deeply nested paths like host.os.name, this can reduce mapper count significantly because only the full leaf path counts as a mapper, not each intermediate segment.

This setting must be specified at index creation time and cannot be changed on an existing index.

Control keys with an ingest pipeline

If you cannot change the source data, use an ingest pipeline to transform documents before they reach the index mapping. Common approaches include using a script processor to rename, drop, or consolidate keys, a foreach processor to normalize tag-like fields into a known structure, or a rename processor to map source fields to controlled names.

This is the most flexible approach but adds processing overhead to the ingest path. Monitor pipeline processor timing with GET /_nodes/stats/ingest to ensure the pipeline is not becoming a bottleneck on its own.

Reindex into a controlled mapping

If an index already has tens of thousands of fields, the only way to reduce the mapping is to create a new index with a controlled mapping and reindex.

# Create a new index with controlled mapping
curl -X PUT 'http://localhost:9200/<new_index>' \
  -H 'Content-Type: application/json' \
  -d '{
    "mappings": {
      "dynamic": "strict",
      "properties": {
        "@timestamp": {"type": "date"},
        "message": {"type": "text"},
        "labels": {"type": "flattened"}
      }
    }
  }'

# Reindex from the old index
curl -X POST 'http://localhost:9200/_reindex' \
  -H 'Content-Type: application/json' \
  -d '{
    "source": {"index": "<old_index>"},
    "dest": {"index": "<new_index>"}
  }'

Documents with fields not in the strict mapping will fail during reindex unless the source data is cleaned first. Plan for partial failures and use a pipeline to transform or drop unknown fields during the reindex.

Prevention

Set dynamic to strict or false on all production indices by default. Use index templates to enforce this. Only enable dynamic mapping on indices where the data schema is controlled and trusted.

Define explicit mappings for fields you need to search or aggregate. Dynamic mapping is a convenience feature, not a production data model. Identify the fields your application actually uses and define them explicitly in index templates.

Use the flattened type for unpredictable payloads. Any field that receives arbitrary user-defined keys should be mapped as flattened rather than left to dynamic mapping.

Monitor total field count. Track indices.mappings from GET /_cluster/stats over time. Alert on unbounded growth. This is the single best leading indicator for mapping explosion.

Apply index templates consistently. Ensure every new index, including those created by ILM rollover, inherits the correct mapping settings. A common mistake is fixing the mapping on one index but leaving the template unchanged, so the next rollover creates another exploded index.

Audit system indices. Some system indices, such as .watches, may have restrictions that prevent updating total_fields.limit. If these indices accumulate too many fields from complex metadata, the only recovery path may be to export the data, delete the index, and recreate it. Monitor their field counts proactively.

How Netdata helps

  • Total field count tracking. Netdata collects cluster-level statistics including mapping field counts, making unbounded growth visible before it causes master instability. A per-second collection interval catches rapid field accumulation that hourly or daily checks would miss.
  • Heap usage correlation across nodes. When mapping explosion inflates cluster state, heap pressure rises uniformly across all nodes. Netdata’s per-node JVM metrics let you distinguish cluster-state-driven heap growth (uniform across nodes) from workload-driven growth (asymmetric).
  • Master node health. Netdata surfaces master node CPU, heap, and GC metrics separately from data nodes. Elevated master CPU or heap combined with rising field count is a strong signal of metadata overload.
  • Pending cluster tasks. Netdata tracks pending task count and age. A growing backlog on the master directly correlates with cluster state publication overhead from large mappings.
  • Cluster state version churn. Frequent cluster state updates, visible through version increment rates, indicate active mapping churn from dynamic mapping. Correlating this with field count growth confirms the diagnosis.

Netdata’s Elasticsearch monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Elasticsearch monitoring with Netdata

Netdata monitors Elasticsearch with per-second metrics and ML anomaly detection. Correlate JVM heap pressure, shard counts, disk watermarks, mapping growth, and merge activity with cluster and node health in one view.