The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cassandra / cassandra-choosing-compaction-strategy

Operations Guides

Cassandra compaction strategies: STCS vs LCS vs TWCS vs UCS

Compaction merges immutable SSTables, discards tombstones, and reclaims disk space. The strategy assigned to a table controls the tradeoff between write amplification and read amplification, and it determines how much temporary disk headroom you must preserve. A fit strategy keeps SSTable counts low and latency predictable. A mismatch creates compaction debt: creeping P99 read latency first, then disk space exhaustion, and finally write rejections when compaction cannot reclaim space fast enough to keep up with flushes.

Operators usually discover mismatches during an incident: a node runs out of disk during major compaction, read latency spikes because L0 has accumulated hundreds of SSTables, or a time-series table recompacts closed windows after a client backfilled historical data. When that happens, you need to know which strategy is in use, why it is behaving that way, and whether the fix is headroom or a strategy migration.

It determines disk provisioning math, I/O baseline, and which signals you watch during an incident. SizeTieredCompactionStrategy (STCS), LeveledCompactionStrategy (LCS), TimeWindowCompactionStrategy (TWCS), and UnifiedCompactionStrategy (UCS) organize SSTables differently, produce different I/O patterns, and fail differently. Understanding the differences lets you distinguish normal background noise from a compaction death spiral before it saturates disks.

What compaction is and why the strategy matters

Cassandra is a log-structured merge-tree database. Writes append to commitlog and memtable; memtables flush to immutable SSTables. Over time, a partition’s data may split across many SSTables, and tombstones persist until compaction removes them. Compaction merges SSTables to reduce the files a read must touch and to evict tombstones.

Every table has a compaction strategy that decides which SSTables to merge and when. This controls three levers: read amplification (how many SSTables a read must consult), write amplification (how many times data is rewritten during compaction), and space amplification (temporary disk space used while merging files).

Compaction creates new SSTables before deleting the old ones, so every strategy consumes temporary disk space. The amount varies significantly between strategies. Provision disk based on raw data size alone and you will eventually lose a node to space exhaustion. The strategy also determines whether compaction I/O is bursty or steady, which affects repair scheduling and whether you can share the disk with other services.

How the strategies work

STCS

STCS is the default. It groups SSTables into buckets by size and compacts similarly-sized files into one larger SSTable when a bucket reaches a threshold. This produces exponential size tiers: many small files, fewer medium files, and rare large files. Compaction I/O is bursty because the system waits for a tier to fill. Reads may need to check every SSTable in a tier when bloom filters produce false positives, so read amplification grows with the number of uncompacted tiers.

LCS

LCS organizes SSTables into levels. Level 0 contains freshly flushed SSTables. Higher levels contain non-overlapping SSTables of roughly uniform size, each level an order of magnitude larger than the previous. Promoting an SSTable from L0 to L1 requires reading all overlapping L1 SSTables and rewriting them. For levels above L0, a read touching a given partition key will find at most one SSTable per level, but L0 may contain many overlapping files. This structure lowers read amplification at the cost of higher write amplification, as data is rewritten multiple times while climbing levels. Compaction I/O is a steady drumbeat rather than a spike.

TWCS

TWCS divides SSTables into time windows based on data timestamps. Within the active window, compaction behaves like STCS. When a window closes and all data inside it expires, Cassandra can drop the entire SSTable instead of processing individual tombstones. This avoids tombstone accumulation, but it depends on time-ordered ingestion. Out-of-order writes that land in old windows, or read repair that pulls historical data into a current window, force those windows to recompact and lose the whole-file deletion optimization.

UCS

UCS is available starting with Cassandra 5.0. It uses density-based triggers and sharded compaction instead of fixed size tiers or levels. Multiple compaction shards run in parallel, and tunable scaling parameters let you bias behavior toward write-heavy or read-heavy profiles without changing the strategy class. Pending tasks tend to be more evenly distributed, producing a smoother I/O pattern than STCS sawtooth spikes.

flowchart TD
    Flush[Memtable flush] --> SST[SSTable created]
    SST --> STCS[STCS bucket by size]
    SST --> LCS[LCS level zero buffer]
    SST --> TWCS[TWCS time window]
    SST --> UCS[UCS density shard]
    STCS --> Tier[Merge into larger tiers]
    LCS --> Level[Promote to fixed levels]
    TWCS --> Window[Compact then delete whole window]
    UCS --> Tune[Tunable read or write bias]

How each strategy behaves in production

STCS appears on tables where the default was never changed. It produces a sawtooth I/O pattern: long quiet periods followed by sudden spikes when a size tier reaches its threshold. During spikes, disk utilization can jump from 20 percent to 90 percent in minutes. If the disk is already above 50 percent full, the node risks entering a disk exhaustion pattern where compaction cannot allocate temporary space and flushes back up. STCS can transiently need up to 100 percent additional disk space during major compaction.

LCS appears on tables that serve latency-sensitive reads or range scans. The I/O pattern is continuous because the strategy must constantly promote files through levels. Compaction activity persists even during low write volume. The danger signal is L0 growth. If nodetool compactionstats shows L0 with dozens of SSTables and the count is rising, the node cannot promote data fast enough and read amplification climbs silently until latency degrades. As a practical rule, keep more than 30 percent of the disk free.

TWCS appears on metrics, logs, and sensor data. In healthy operation, old windows should show zero compaction activity. Compaction running on windows that closed days ago indicates out-of-order writes or read repair pulling historical data into current windows. This is the primary TWCS failure mode. Because expired windows can be dropped whole, TWCS typically needs only 20 percent disk headroom.

UCS appears on Cassandra 5.0 clusters or migrated tables. Pending compaction tasks are more evenly distributed, so dashboards look smoother than the STCS sawtooth. The main operational consideration is ensuring you are on Cassandra 5.0 or later; attempting to use UCS on earlier versions will fail. Plan for disk headroom comparable to LCS.

Tradeoffs and when to use each

StrategyBest forRead amplificationWrite amplificationDisk headroomPrimary risk
STCSWrite-heavy, bursty ingestionHighLowGreater than 50 percent freeDisk exhaustion during major compaction
LCSRead-heavy, steady workloadLow (excluding L0)HighGreater than 30 percent freeL0 backlog under write spikes
TWCSTime-series with uniform TTLMediumLow after window closeGreater than 20 percent freeOut-of-order writes forcing window recompaction
UCSGeneral purpose, Cassandra 5.0+TunableTunableSimilar to LCSVersion-locked to 5.0 and later

If you run Cassandra 5.0 or later and have no strong reason to use a legacy strategy, start with UCS and tune its scaling parameters. On earlier versions, use STCS for ingestion-heavy tables and LCS for latency-sensitive reads. Reserve TWCS for tables where every row has a TTL and writes are time-ordered.

Strategy migration is possible with an ALTER TABLE statement. The change applies to future compactions; existing SSTables are not rewritten immediately. However, the new strategy may still initiate large compaction jobs as it restructures existing data, which can saturate disk I/O for hours or days. Do not change strategies during an incident. If you must migrate from STCS to LCS, plan the operation during a maintenance window when you can tolerate the write amplification cost of the initial level buildup.

Signals to watch in production

SignalWhy it mattersWarning sign
Pending compactions trendLeading indicator of compaction debtIncreasing over 4 or more hours; STCS can spike to hundreds transiently, but LCS should stay low
SSTable count per tableDirect measure of read amplificationSTCS sustained high count; LCS growing L0 count or total count rising steadily
Disk space freeCompaction needs temporary space to write merged filesBelow strategy-specific headroom: STCS below 50 percent, LCS below 30 percent, TWCS below 20 percent
Disk I/O utilizationCompaction competes with reads and flushes for the same devicesAbove 80 percent sustained
Tombstone scan warningsIndicates TWCS window corruption or repair gapsSustained log entries or aborted reads

Watch the trend, not the absolute value. A pending compaction count that rises from 10 to 30 over six hours is more significant than a stable count of 50. For LCS, watch L0 SSTable count. For STCS, watch the size of the largest tier relative to free disk. In TWCS, check nodetool compactionstats for compaction activity on closed windows that should be idle.

How Netdata helps

  • Correlate pending compactions with per-device disk I/O utilization and P99 read latency to spot a compaction death spiral before it triggers client timeouts.
  • Apply strategy-specific thresholds to SSTable count and disk space alerts. An STCS node at 60 percent full is an emergency; an LCS node at 60 percent is not.
  • Track JVM heap usage and GC pause duration alongside compaction metrics. Compaction backlog increases heap pressure from SSTable metadata, and long GC pauses can masquerade as compaction stalls.
  • Use per-node comparison charts to detect asymmetric compaction backlog, where one node falls behind peers due to a hot partition or local disk degradation.
  • Annotate maintenance windows so post-restart compaction bursts and planned strategy migrations do not trigger false-positive alerts.
The Netdata solution

Cassandra monitoring with Netdata

Netdata monitors Apache Cassandra with per-second metrics and automatic dashboards. Correlate GC pauses, compaction backlog, tombstone rates, pending hints, and disk usage across nodes to catch a creeping cluster before it tips over.