The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cassandra / how-cassandra-works-in-production

Operations Guides

How Cassandra actually works in production: a mental model for operators

If you operate Cassandra in production, you are managing a distributed, partitioned, replicated log-structured merge tree where every node is a peer. There is no master to restart, no central query planner to tune, and no automatic load balancing that will save you from a hot partition. Every write is a sequential append to a commitlog and an in-memory update on multiple replicas. Every read is a merge of memtables and immutable SSTables, filtered by probabilistic bloom filters and reconciled by timestamp.

Cassandra delegates all hard decisions to the operator. Compaction strategy sets your I/O profile, consistency level sets availability boundaries, and the failure detector’s phi threshold sets the stall tolerance before the cluster reconfigures. If you do not understand how the write path, read path, gossip, compaction, and repair interact, you will misdiagnose GC pauses as network partitions, compaction backlog as slow queries, and disk space exhaustion as capacity growth when it is actually a strategy mismatch.

What it is and why it matters

Cassandra uses a peer-to-peer token ring. The partitioner hashes the partition key to a token, and the snitch places replicas clockwise on the ring across racks and datacenters. Any node can receive a CQL request and act as the coordinator, forwarding operations to the replica set and waiting for acknowledgements to satisfy the consistency level. This eliminates single points of failure but makes latency, availability, and consistency emergent properties of local storage behavior, network topology, and the phi accrual failure detector.

The storage engine is an LSM tree. Writes are append-only and sequential; reads merge live memtables with immutable SSTables, applying tombstones and resolving conflicts across replicas. Most production incidents start in the gap between fast writes and expensive reads. Understanding this asymmetry prevents you from adding disk when queries are slow instead of recognizing that compaction cannot keep up with ingestion.

How it works

Four subsystems run continuously and interact to produce the behavior you see in metrics and logs.

Write path. A write arrives at the coordinator, which hashes the partition key to determine the replica nodes that own the token range. On each replica, the mutation is appended to the commitlog and inserted into the memtable, a per-table sorted in-memory structure. The coordinator acknowledges the write once enough replicas have responded to satisfy the consistency level. By default, commitlog durability is periodic (commitlog_sync: periodic), acknowledging before fsync; use batch only if you accept the latency cost for lower durability risk. When a memtable reaches its size threshold, it flushes to disk as an immutable SSTable. The corresponding commitlog segments become eligible for recycling. SSTables accumulate until compaction merges them. On startup, the node replays unflushed commitlog segments to recover memtables.

Read path. The coordinator sends requests to enough replicas to satisfy the consistency level. For levels below ALL, it typically requests full data from one replica and digests from the others. If digests mismatch, it fetches full rows and resolves by timestamp. On each replica, Cassandra checks the memtable first, then consults per-SSTable bloom filters to rule out SSTables that cannot contain the partition. A false positive causes an unnecessary SSTable read; the false positive rate is configurable per table via bloom_filter_fp_chance. For candidate SSTables, the partition summary narrows the index lookup, and the partition index locates the exact byte offset. Data from the memtable and all relevant SSTables is merged, tombstones are applied, and the result is returned.

Background maintenance. Compaction runs continuously to merge SSTables, discard tombstones, and consolidate data. STCS groups SSTables by size and merges similar-sized files, favoring write throughput at the cost of bursty I/O and high space amplification. LCS organizes SSTables into levels with exponentially increasing size targets, bounding read amplification but increasing write I/O. TWCS compacts data within time windows and drops expired windows as units, making it ideal for TTL workloads. UCS (5.0+) uses density-based triggers.

Gossip runs every second. Each node exchanges state with up to three peers, propagating heartbeats, topology, schema versions, and load. The phi accrual failure detector evaluates heartbeat interarrival times; with a default threshold of 8 and a 1-second gossip interval, a node missing heartbeats for roughly 18 seconds is marked DOWN by the local observer.

Hinted handoff stores writes locally when a replica is unreachable, replaying them when the target recovers within max_hint_window_in_ms (default 3 hours). Anti-entropy repair compares Merkle trees across replicas and streams differences. It must complete within gc_grace_seconds (default 10 days) to prevent tombstones from expiring on unrepaired data and causing deleted data to reappear. Incremental repair is the default since Cassandra 4.0. It marks repaired SSTables separately, which changes compaction behavior; unrepaired data compacts only with unrepaired data. Full repair is still required periodically to guard against disk corruption and operator error. Streaming handles bulk transfers during bootstrap, decommission, rebuild, and full repair.

flowchart LR
    Client -->|CQL request| Coordinator
    Coordinator -->|hash partition key| TokenRing
    TokenRing -->|forward write| ReplicaA
    TokenRing -->|forward write| ReplicaB
    ReplicaA -->|append| CommitLog
    ReplicaA -->|update| Memtable
    Memtable -->|flush| SSTable
    SSTable -->|merge| Compaction
    ReplicaA -->|exchange state| Gossip
    Gossip -->|phi threshold| FailureDetector
    ReplicaA -->|store replay| HintedHandoff
    ReplicaA -->|compare trees| Repair

Where it shows up in production

The subsystems above compete for a fixed set of resources. Recognizing the competition patterns lets you distinguish root causes from symptoms.

ResourceCompetition pattern
Disk I/OCommitlog writes, memtable flushes, compaction, and reads all compete. Compaction is usually the dominant consumer. Plan for SSDs.
JVM heapMemtables, key cache, query results, compression metadata, and in-flight requests. GC pauses directly impact availability.
Off-heap memoryBloom filters, compression metadata, direct buffers, and chunk cache. Invisible to JVM GC but counts toward Linux OOM.
CPUCompaction, query processing, encryption, and GC.
NetworkClient traffic, inter-node replication, gossip, and streaming.
File descriptorsEach SSTable opens multiple file handles; hundreds of SSTables multiplied by several files per SSTable equals thousands of FDs.
Disk spaceSSTables, commitlog, hints, and snapshots. STCS can transiently need up to 100% additional space during major compaction.

Deployment choices change which signals matter most. Multi-DC deployments add cross-DC latency and streaming saturation during repair. Higher replication factors increase write coordination costs. Virtual nodes (vnodes) change compaction and repair dynamics. Lightweight Transactions (Paxos-based) add roughly four round-trips of latency that must be monitored separately from standard reads and writes.

You can observe these competitions directly. nodetool compactionstats shows whether compaction is keeping up. nodetool tpstats reveals thread pool saturation in the mutation, read, and native transport stages, including the HintedHandoff stage. nodetool info exposes heap usage. Check OS-level file descriptor counts because Cassandra opens thousands of handles across SSTables. Interpreting these outputs requires knowing which subsystem they belong to.

Operational tradeoffs and failure archetypes

Production incidents in Cassandra tend to follow repeatable archetypes that emerge from the interaction of the subsystems described above.

GC death spiral. Heap pressure triggers long GC pauses. During a pause, gossip cannot exchange heartbeats, so peers mark the node DOWN. Clients timeout and retry; other nodes store hints. When the GC finishes, the node is flooded with replayed hints and retried requests, increasing memory pressure and triggering longer pauses. The trigger is often a large partition read, tombstone-heavy scan, or memory leak that promotes objects to the old generation. Once old-gen occupancy exceeds what a full GC can reclaim, the JVM enters continuous stop-the-world collection.

Compaction death spiral. Write rate exceeds compaction throughput. SSTables accumulate, so every read must check more files, and read amplification increases. Latency degrades, but writes remain fast, creating a false sense of health. Disk I/O saturates under the combined load of reads and compaction. Eventually disk space runs out or reads become too slow to meet SLAs. The signature is write latency remaining healthy while read latency degrades.

Tombstone storm. Deletes and TTL expirations create tombstones that persist until compaction after all replicas have been repaired. If tombstones accumulate across many SSTables, reads must scan and merge enormous amounts of dead data. This consumes CPU, memory, and I/O. At tombstone_failure_threshold (default 100,000), queries abort entirely. The root cause is usually a data model mismatch: using Cassandra as a queue, issuing frequent range deletes, or failing to run repair so that tombstones cannot be purged.

Disk space exhaustion. Compaction needs temporary space to write merged SSTables before deleting old ones. With STCS, major compaction can temporarily require space equal to the largest table size. If snapshots, hints, or uncompacted data consume the remaining space, compaction stalls. Without compaction, old SSTables are never deleted, and writes eventually block when the commitlog cannot allocate new segments.

Hint overflow. When a replica is down, coordinators store hints. If the outage exceeds max_hint_window_in_ms (default 3 hours), hints stop being stored. Data written during the remainder of the outage is permanently missing from that replica unless anti-entropy repair is run after recovery.

Signals to watch in production

SignalWhy it mattersWarning sign
Node liveness (gossip/phi)A DOWN node reduces effective replication and can cause quorum loss for affected token ranges.DownEndpointCount > 0 sustained > 5 min; flapping > 3 transitions in 30 min.
Coordinator read/write latency (P99)Direct measure of user experience and end-to-end path cost across replicas and merging.P99 > 3x rolling baseline or approaching half of the configured request timeout.
Pending compactionsLeading indicator of compaction debt. Rising pending means read amplification and disk space consumption will follow.Trending upward over 4+ hours; > 50 sustained in STCS or above expected per-level count in LCS.
GC pause durationPauses exceeding the failure detector threshold mark the node DOWN. Shorter pauses still degrade latency and can cascade.Max pause > 2s or GC time > 5% of wall clock over 5 minutes.
Dropped messages (MUTATION/READ)The node is shedding load. Dropped mutations create silent inconsistency that only repair can fix.Any sustained non-zero rate for > 60 seconds.
SSTable count per tableEach SSTable adds bloom filter checks and merge work to every read.> 50 in STCS, L0 > 32 in LCS, or trending upward over days.
Large partition sizeWide rows cause heap pressure on reads and can stall compaction.nodetool tablestats reports max partition size > 100 MB.
Repair completion timeUnrepaired data beyond gc_grace_seconds allows tombstones to expire and deleted data to resurrect.Last successful repair > 80% of gc_grace_seconds (default > 8 days).
Disk space availableCompaction cannot run without headroom. Exhaustion blocks flushes and writes.< 30% free; with STCS, < 50% free is dangerous due to major compaction amplification.

How Netdata helps

  • Correlate coordinator latency with replica GC pauses. Collect JVM GarbageCollector CollectionTime and ClientRequest latency percentiles to confirm whether a P99 spike originates from a replica death spiral or a network issue.
  • Spot compaction debt before reads degrade. Track Compaction PendingTasks alongside per-table LiveSSTableCount to detect the write-amplification gap that precedes read latency cliffs.
  • Monitor off-heap RSS, not just JVM heap. Track process RSS and OS memory metrics to surface off-heap growth from bloom filters and compression metadata that JMX heap graphs hide.
  • Distinguish phi convict flapping from hardware failure. Combine FailureDetector states with OS-level CPU wait and disk latency to determine if a node is marked DOWN due to GC pressure or actual network loss.
  • Track repair cadence against gc_grace_seconds. Alert on repair completion timestamps to ensure anti-entropy runs before tombstones resurrect deleted data.
The Netdata solution

Cassandra monitoring with Netdata

Netdata monitors Apache Cassandra with per-second metrics and automatic dashboards. Correlate GC pauses, compaction backlog, tombstone rates, pending hints, and disk usage across nodes to catch a creeping cluster before it tips over.