The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cassandra / cassandra-repair-overload

Operations Guides

Cassandra repair overload: when anti-entropy repair causes the outage it prevents

A full nodetool repair started during peak traffic can spike P99 read latency from milliseconds to hundreds of milliseconds, trigger write timeouts, and cause nodes to flap between UP and DOWN in gossip. The repair job meant to prevent inconsistency becomes the cause of the outage.

Anti-entropy repair is a heavy distributed scan, not a background task. It reads all local data to build Merkle trees, exchanges hashes with replicas, and streams differing ranges. On a multi-terabyte node this means terabytes of sequential disk reads, heavy CPU hashing, and gigabits of network traffic. Without dedicated headroom, repair competes with the commitlog, memtable flushes, compaction, and client requests for disk bandwidth, CPU, and network. The result is thread pool backpressure, dropped messages, GC pressure from Merkle tree construction, and eventually gossip failure as the node becomes unresponsive.

The symptoms look like cascading failure, but they resolve when the repair load is removed.

flowchart TD
    A[Full repair starts] --> B[Merkle tree construction reads all local data]
    B --> C{Disk or network saturated?}
    C -->|Yes| D[Foreground reads and writes starved]
    D --> E[Thread pools back up]
    E --> F[Dropped messages and timeouts]
    F --> G[GC pressure from queued requests]
    G --> H[Gossip failure node marked DOWN]
    H --> I[Client unavailable exceptions]
    C -->|No| J[Repair completes normally]

What this means

Repair overload is resource contention, not hardware failure. Cassandra’s anti-entropy process treats the node’s entire dataset as a checksum source. Full repair generates I/O and network load comparable to a major compaction combined with a bootstrap stream. Every local SSTable is read to construct a Merkle tree, which is then exchanged with replicas and differenced. Missing data is streamed over the network. If the node lacks headroom, repair starves foreground reads and writes. The node is healthy but saturated.

Common causes

CauseWhat it looks likeFirst thing to check
Full repair during peak trafficP99 latency and timeout rates rise within minutes of repair start; dropped mutations appearnodetool netstats
Unthrottled streaming on dense nodesDisk %util near 100%, commitlog pending tasks > 0, mutation stage pending sustainediostat -x 1 and nodetool tpstats
Repair colliding with bootstrap or decommissionNetwork bandwidth saturated, streaming sessions from multiple sources, node marked DOWNnodetool netstats and nodetool status
Anti-compaction backlog after repairCompaction pending spikes after repair completes, SSTable count grows, read latency stays elevatednodetool compactionstats

Quick checks

Run these safe, read-only commands to confirm whether repair is the source of saturation.

# Check active repair and streaming sessions
nodetool netstats

# Check dropped messages and thread pool saturation
nodetool tpstats

# Check compaction backlog including anti-compaction
nodetool compactionstats

# Check disk I/O saturation on data and commitlog devices
iostat -x 1

# Check coordinator latency percentiles
nodetool proxyhistograms

# Check JVM heap usage
nodetool info | grep -i "Heap Memory"

# Check node liveness and schema agreement
nodetool status
nodetool describecluster

How to diagnose it

  1. Confirm repair is running. Run nodetool netstats to look for active streaming sessions labeled with repair ranges. On Cassandra 4.0+, run nodetool repair_admin list to see active repair sessions and their token ranges.
  2. Correlate the timeline. Compare the repair start time against the onset of client timeout and unavailable metrics. If the latency spike begins within minutes of repair initiation, the correlation is strong.
  3. Check for thread pool saturation. Run nodetool tpstats and look at MutationStage, ReadStage, and Native-Transport-Requests. Sustained Pending > 0 means requests are queuing. Blocked > 0 means the submitting thread is being backpressure-blocked because the queue is full.
  4. Inspect disk I/O. Run iostat -x 1 on the data directory device and the commitlog device. If %util is > 80% sustained or await is elevated beyond baseline, the disk is saturated. If commitlog and data share the same device, repair reads directly contend with commitlog writes.
  5. Check for dropped messages. In nodetool tpstats, any non-zero rate of dropped MUTATION or READ messages means the node is shedding load. Dropped mutations risk replica inconsistency; the write may have succeeded on some replicas before the coordinator timed out.
  6. Evaluate compaction state. Run nodetool compactionstats. If pending tasks are growing while repair is active, the node cannot keep up with both anti-compaction and normal compaction.
  7. Check JVM pressure. Run nodetool info for heap usage. If heap usage is high, check GC logs for pause duration. Pauses > 500ms degrade latency. Pauses > 2 seconds risk gossip failure and nodes being marked DOWN.
  8. Determine repair scope. Check whether the running repair is full or incremental. Full repair reads the entire dataset; incremental repair (4.0+) processes only unrepaired data and is significantly lighter. If you see anti-compaction creating a large number of new SSTables immediately after repair, expect a secondary compaction wave that can prolong the overload.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Client request latency (coordinator P99)Direct measure of user-visible degradationP99 > 3x rolling 1-hour baseline sustained > 5 min
Dropped messages (MUTATION/READ)Node is overloaded and shedding loadAny sustained non-zero rate > 60 seconds
Thread pool pending tasks (MutationStage/ReadStage)Backpressure before messages are droppedPending > 0 sustained > 60 seconds
Disk I/O utilization (%util and await)Repair saturates disk bandwidth%util > 80% or await > 10ms on SSD sustained
Active repair streaming sessionsRepair streams differences between replicasStreaming sessions coinciding with latency spikes
Compaction pending tasksAnti-compaction adds to background debtPending trending upward during or after repair
GC pause durationMerkle trees and streaming pressure heapPause > 500ms; > 2s risks gossip failure
Node liveness (gossip state)Extreme overload causes phi accrual failureNode marked DOWN or flapping > 2 transitions in 10 min
Repair completion statusPartial repairs create false safetyRepair duration exceeds expected window without completion

Fixes

Throttle or relocate the repair load

If repair is causing acute client impact, cap outbound streaming bandwidth. Run nodetool setstreamthroughput <value_in_Mbit_per_s> (use -m for MiB/s) to reduce it dynamically. For a persistent change, lower stream_throughput_outbound (legacy stream_throughput_outbound_megabits_per_sec) in cassandra.yaml and perform a rolling restart. If you use Cassandra Reaper, configure more conservative per-segment throughput. If compaction is contending for disk bandwidth, temporarily lower its throughput cap with nodetool setcompactionthroughput <mb_per_sec>. This throttles background compaction to favor foreground traffic, at the cost of slower compaction catch-up. Move full repairs to off-peak windows.

Switch to incremental repair on Cassandra 4.0+

If you are running full repairs on 4.0 or later, migrate to incremental repair. Incremental repair tracks repaired SSTables via metadata and processes only unrepaired data written since the last cycle. It is the default in 4.0+ and significantly lighter than full repair. The tradeoff is that anti-compaction still creates additional SSTables, so monitor compaction pending after each run. Do not use incremental repair on versions earlier than 4.0; pre-4.0 incremental repair had bugs that could cause data corruption.

Use subrange repair with Reaper

Instead of a single full-range repair per node, use Reaper to orchestrate subrange repair. This divides a node’s token range into smaller segments, limiting per-session memory overhead and isolating failures to a single segment. If one segment fails, Reaper retries it without re-scanning the entire node. Subrange repair is more efficient for large clusters, and Reaper provides per-segment success visibility that nodetool repair alone does not. The tradeoff is longer total repair duration, but each segment imposes a smaller peak load and can be scheduled independently.

Separate commitlog and data directories

If commitlog and data directories share a physical device, repair reads on the data directory contend directly with commitlog writes. Separate them onto dedicated volumes. This is not an immediate fix during an incident, but it eliminates a major contention path.

Prevention

  • Schedule repairs during low-traffic windows. Never run full repairs during peak traffic. Repair generates I/O and network traffic comparable to a major failure. For multi-DC clusters, schedule repairs sequentially by datacenter to avoid cross-DC streaming load.
  • Prefer incremental repair on Cassandra 4.0+. This bounds repair cost to recent write volume rather than total dataset size. Verify that incremental repair completes successfully; partial runs can leave unrepaired SSTables that accumulate debt.
  • Use Reaper with subrange segmentation and conservative throughput limits. This avoids monolithic repair sessions and spreads load over time. Reaper also provides per-segment success and failure visibility that nodetool repair lacks.
  • Monitor repair completion, not just start. Repairs can silently fail to complete all token ranges. Alert when repair duration exceeds the expected window or when last repair time approaches 80% of gc_grace_seconds. A repair that starts but does not finish every range is worse than no repair because it creates a false sense of safety.
  • Maintain disk I/O headroom. Keep sustained disk utilization below 70% to absorb repair bursts without starving foreground traffic. Major compaction can transiently need up to 100% additional disk space. High utilization increases the risk of space exhaustion during background operations.

How Netdata helps

  • Correlate per-device disk I/O utilization (%util, await) with active repair streaming to identify saturation immediately.
  • Track JVM GC pause duration and heap usage to catch pressure from Merkle tree construction before gossip fails.
  • Monitor thread pool pending tasks for MutationStage, ReadStage, and Native-Transport-Requests to detect backpressure before messages are dropped.
  • Alert on dropped mutation and read rates, which are the first signals of repair-induced overload.
  • Overlay node gossip state (UP/DOWN transitions) with latency metrics to distinguish repair saturation from true hardware failures.
The Netdata solution

Cassandra monitoring with Netdata

Netdata monitors Apache Cassandra with per-second metrics and automatic dashboards. Correlate GC pauses, compaction backlog, tombstone rates, pending hints, and disk usage across nodes to catch a creeping cluster before it tips over.