The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cassandra / cassandra-dropped-reads-and-messages

Operations Guides

Cassandra dropped reads and other messages: reading nodetool tpstats Dropped

When nodetool tpstats reports non-zero values in the Dropped section, the node discarded internal messages that exceeded their stage timeout. These counters are cumulative since JVM startup, not rates. A non-zero value warrants investigation: the timeout is defined in cassandra.yaml by settings such as read_request_timeout_in_ms and write_request_timeout_in_ms, so a drop means the message sat in the queue for seconds.

Dropped messages are a lagging indicator. Correlate the drop type with the matching thread pool pending count, disk I/O latency, and GC pause duration to find the root cause.

What it is and why it matters

Cassandra routes internal operations through dedicated thread pools: ReadStage, MutationStage, HintsDispatcher, and RequestResponseStage. Every message carries a timeout. If a pool cannot process the message before expiration, the stage discards it and increments the Dropped counter for that scope.

Because these counters are cumulative since JVM startup, a large absolute number on a long-running node may reflect a past incident. What matters is the rate of change. A healthy node in steady state drops zero messages. Any sustained non-zero rate indicates overload.

How Cassandra decides to drop a message

When a message arrives, Cassandra places it in its assigned stage queue: ReadStage for reads, MutationStage for standard writes and read repairs, HintsDispatcher for hints, and RequestResponseStage for internode responses. Worker threads pull from these queues.

If threads are blocked on disk I/O, stalled by a GC pause, or outpaced by arrival rate, queue depth grows. Once a message’s age exceeds its timeout, the stage drops it. The timeout is defined in cassandra.yaml by settings such as read_request_timeout_in_ms and write_request_timeout_in_ms; exceeding it means the queue was severely backed up or the JVM was frozen.

flowchart TD
    A[Message arrives at stage queue] --> B{Queued longer than timeout?}
    B -->|Yes| C[Increment Dropped counter]
    B -->|No| D[Thread processes request]
    C --> E[READ / RANGE_SLICE]
    C --> F[MUTATION / COUNTER / BATCH]
    C --> G[HINT]
    C --> H[REQUEST_RESPONSE]
    E --> I[ReadStage pending / Disk I/O / GC]
    F --> J[MutationStage pending / Commitlog]
    G --> K[Target node health / Repair]
    H --> L[Cross-node latency / CPU]

What each drop type means

READ and RANGE_SLICE

READ drops mean a point read request was abandoned after exceeding its timeout in ReadStage. RANGE_SLICE drops mean a range scan suffered the same fate. Both indicate that the read path on this replica could not keep up.

Common causes include GC pauses freezing the stage, disk I/O saturation preventing SSTable lookups, or thread pool saturation from large partition reads and tombstone scans. When these drop, the client has already received a timeout. Check ReadStage pending tasks with nodetool tpstats | grep -E "ReadStage|MutationStage", disk await with iostat -x 1 on the data volume, and GC logs for stop-the-world pauses.

MUTATION, COUNTER_MUTATION, BATCH_STORE, and BATCH_REMOVE

These are write-path drops routed through MutationStage. MUTATION drops indicate a standard write was discarded on this replica. If the consistency level was met by other replicas, the client may think the write succeeded. The data is now inconsistent on this node and can only be fixed by repair.

COUNTER_MUTATION drops point to counter writes, which are read-modify-write operations and more expensive than standard writes. BATCH_STORE and BATCH_REMOVE relate to logged batch processing. BATCH_STORE drops indicate the batchlog metadata write failed. BATCH_REMOVE drops indicate batch cleanup failed.

For all four, check MutationStage pending tasks with nodetool tpstats, and whether compaction backlog is stealing disk throughput from the write path using nodetool compactionstats.

READ_REPAIR

READ_REPAIR drops occur when a read repair mutation is discarded. Because read repair is implemented as a mutation sent to inconsistent replicas, it routes through MutationStage and is governed by the write request timeout. If read repair mutations are dropping, the write path is overloaded and replica inconsistencies may persist. Do not tune read repair parameters first; investigate MutationStage saturation, commitlog latency, and disk await.

HINT

HINT drops occur in HintsDispatcher. They are almost always symptomatic of a problem elsewhere, not a primary bottleneck on the coordinator. Either the target replica is down or unreachable, or the node is overwhelmed by accumulated hint volume. Sustained HINT drops usually mean a node has been down long enough that hints are backing up, or hint replay is overwhelming a recovering node. Verify target node liveness with nodetool status, check the configured hints directory size on disk, and schedule repair. Do not tune the hint dispatcher before confirming replica health.

REQUEST_RESPONSE

REQUEST_RESPONSE drops mean the local node completed the work but could not send the response before the originating timeout fired. This indicates the node was slow to respond to another node even after processing finished. Check cross-node latency with ping or mtr, run nodetool netstats for pending internode responses, and review GC logs and CPU saturation on the local node.

Common misreads and missteps

Cumulative counters mask rate. Sample the counter twice with a known interval to compute the current rate:

date && nodetool tpstats | grep -A 20 "Message type"

On a long-running node, thousands of drops may reflect one past incident.

Client success does not mean zero data loss. At QUORUM, a dropped mutation on one replica may be invisible to the client if other replicas acknowledged. The missed replica is now inconsistent. If you only monitor client errors, you will miss replica-side drops.

GC pauses cause multi-scope drops. A long stop-the-world pause freezes every stage simultaneously. If you see READ, MUTATION, and REQUEST_RESPONSE all increasing together, check GC logs before tuning thread pools or disk layout.

Do not ignore HINT drops. They are not harmless metadata. They indicate either a downed replica or a coordinator that cannot keep up with replay, both of which require operational action.

Signals to watch in production

SignalWhy it mattersWarning signHow to check
ReadStage / MutationStage pending tasksLeading indicator that a stage is backing up before drops occur.Sustained pending greater than 0 for more than 60 secondsnodetool tpstats
GC pause durationLong pauses freeze all stages and directly cause message timeouts.Pause exceeds 2 seconds, or old-gen collections exceed 1 per minuteGC logs (location depends on JVM flags)
Disk I/O awaitSaturated I/O prevents reads and writes from completing in time.await greater than 10 ms on SSD sustained for 5 minutesiostat -x 1 on the data volume
Compaction pending tasksCompactions steal disk I/O and CPU from the write path.Pending compactions sustained above baselinenodetool compactionstats
Client request timeoutsCoordinator-level view of the same pressure causing drops on replicas.Non-zero timeout rate sustained for more than 60 secondsApplication metrics

How Netdata helps

  • Plots per-scope dropped-message rates from JMX in real time, so you do not need to sample nodetool tpstats manually.
  • Correlates MUTATION or READ spikes with thread pool pending tasks and GC pause duration on the same timeline.
  • Tracks disk I/O latency per device to distinguish commitlog contention from data directory saturation.
  • Alerts on non-zero dropped message rates by scope, treating any sustained rate as an anomaly.
The Netdata solution

Cassandra monitoring with Netdata

Netdata monitors Apache Cassandra with per-second metrics and automatic dashboards. Correlate GC pauses, compaction backlog, tombstone rates, pending hints, and disk usage across nodes to catch a creeping cluster before it tips over.