The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / clickhouse / clickhouse-long-running-queries

Operations Guides

ClickHouse long-running queries: finding and killing the resource hog

A query that should finish in seconds is still running after twenty minutes. Memory on the ClickHouse node is climbing, query latency has doubled, and you suspect a single query is holding resources it will never release. In ClickHouse, a long-running query can be a legitimate analytical job crunching terabytes, a Cartesian JOIN exploding in memory, or a GROUP BY that has spilled to disk and slowed to a crawl. Telling the difference determines whether you kill it or let it finish.

The live view is system.processes. It shows wall-clock elapsed time, current memory_usage, read_rows, and the query text. High elapsed alone does not mean the query is broken; some workloads expect hours-long execution. The danger is the resource hog that is stalled, spilling, or multiplying rows without meaningful progress. Left alone, it can exhaust memory, saturate disk I/O, or trigger OOM kills that destabilize the node.

What this means

ClickHouse queries execute as a pipeline of processors that hold memory for hash tables, sort buffers, and decompressed blocks. A long-running query keeps those allocations until it finishes or is cancelled. It also holds file descriptors for the parts it reads. A large JOIN or GROUP BY may allocate until it hits the per-query memory limit, then either fail with exception code 241 (MEMORY_LIMIT_EXCEEDED) or spill temporary data to /var/lib/clickhouse/tmp/. Spill-to-disk turns a memory problem into an I/O problem. The query slows by orders of magnitude but continues holding resources.

system.processes.elapsed measures wall-clock time, not CPU time. A query can show high elapsed with almost no CPU use if it is waiting on disk I/O after spilling, waiting on a lock, or blocked by a distributed subquery that is itself stalled. Correlate elapsed with read_rows and memory_usage to separate real work from a runaway process.

flowchart TD
    A[Query elapsed high] --> B{read_rows increasing?}
    B -->|Yes| C[Legitimate heavy ETL]
    B -->|No| D{memory_usage high?}
    D -->|Yes| E[Runaway query]
    D -->|No| F{tmp directory growing?}
    F -->|Yes| G[Spill to disk]
    F -->|No| H[Stuck or blocked]

Common causes

CauseWhat it looks likeFirst thing to check
Cartesian-product JOINmemory_usage climbs rapidly; query text shows missing or broad JOIN keyssystem.processes.memory_usage and the query text for missing ON clauses
OOM-induced spill-to-diskelapsed high, read_rows flat or advancing slowly; temp directory grows/var/lib/clickhouse/tmp/ size and read_rows delta over 30-60 seconds
Full table scan from partition pruning failureread_rows increases linearly across the entire table, far exceeding expected scopesystem.query_log for read_rows vs expected rows from the same query pattern
Runaway GROUP BY or DISTINCTpeak_memory_usage is very high; may precede exception code 241system.processes.peak_memory_usage vs the user max_memory_usage limit
Legitimate heavy ETLread_rows advances steadily, memory is stable, and the query runs in a known batch windowsystem.processes.read_rows sampled twice over a minute

Quick checks

Run these read-only checks to size up the situation.

-- List queries running longer than 60 seconds
SELECT
    query_id,
    user,
    elapsed,
    formatReadableSize(memory_usage) AS mem_used,
    read_rows,
    substring(query, 1, 200) AS query_prefix
FROM system.processes
WHERE elapsed > 60
ORDER BY elapsed DESC
LIMIT 20;
-- Total memory held by all running queries
SELECT
    formatReadableSize(sum(memory_usage)) AS total_query_memory,
    count(*) AS running_queries
FROM system.processes;
-- Server-level tracked memory
SELECT metric, value
FROM system.metrics
WHERE metric = 'MemoryTracking';
-- Top queries by current memory usage
SELECT
    query_id,
    user,
    elapsed,
    formatReadableSize(memory_usage) AS mem_used,
    formatReadableSize(peak_memory_usage) AS peak_mem,
    read_rows,
    substring(query, 1, 200) AS query_prefix
FROM system.processes
ORDER BY memory_usage DESC
LIMIT 10;
-- Sample read_rows for a specific suspect query
-- Run this twice, 30 to 60 seconds apart, to see progress
SELECT query_id, read_rows
FROM system.processes
WHERE query_id = '<query_id>';
# Check temp directory for spill-to-disk activity
du -sh /var/lib/clickhouse/tmp/
-- Recent query failures by exception code
SELECT
    exception_code,
    count() AS cnt,
    any(exception) AS sample
FROM system.query_log
WHERE type = 'ExceptionWhileProcessing'
  AND event_time > now() - INTERVAL 10 MINUTE
GROUP BY exception_code
ORDER BY cnt DESC;

How to diagnose it

  1. Identify candidates. Query system.processes with elapsed > 60 (or your workload threshold). Sort by elapsed DESC to see the oldest queries, and by memory_usage DESC to see the heaviest consumers.

  2. Check progress. Run SELECT read_rows FROM system.processes WHERE query_id = '...' twice, 30 to 60 seconds apart. If read_rows increases, the query is actively scanning and may be legitimate heavy ETL. If read_rows is static, the query is stalled.

  3. Inspect memory footprint. High memory_usage or peak_memory_usage with flat read_rows suggests a runaway query, such as a Cartesian product or an unbounded hash table.

  4. Look for spill-to-disk. Check /var/lib/clickhouse/tmp/ with du. If it grows while the query runs and CPU is low, the query exceeded memory limits and is writing temporary aggregation or sort data to disk. This happens when max_bytes_before_external_group_by is reached; external sort has the analogous max_bytes_before_external_sort threshold. Exceeding a server or query memory limit throws an error rather than triggering this controlled spill.

  5. Examine the query text. Use substring(query, 1, 500) from system.processes. Look for JOINs without equality conditions, missing WHERE clauses on the partition key, or SELECT * across huge tables.

  6. Correlate with failures. Query system.query_log for ExceptionWhileProcessing entries with code 241. A cluster of memory errors around the same time suggests the long query is part of a broader memory pressure event.

  7. Assess server-level impact. Compare MemoryTracking from system.metrics against the server max_server_memory_usage setting. If the server is above 80% and the long query is the top consumer, it is starving other queries.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
system.processes.elapsedIdentifies queries exceeding expected durationQuery running > 10x expected duration or > 30 minutes (workload-dependent)
system.processes.memory_usageReveals resource hoarding per querySingle query using > 50% of max_server_memory_usage
system.processes.read_rowsDistinguishes progress from stagnationread_rows unchanged over 60+ seconds on a supposedly active query
system.events.FailedQuery / FailedSelectQueryReveals systemic query failuresSustained error rate > 1% of total queries over 5 minutes
MemoryTracking vs max_server_memory_usageShows server-level memory pressureMemoryTracking > 80% of limit
Temp directory size (/var/lib/clickhouse/tmp/)Indicates spill-to-disk from memory limitsGrowth correlating with a long-running query
system.query_log.query_duration_ms (P99)Tail latency degradationSustained P99 > 2x baseline for > 15 minutes

Fixes

Kill a runaway or stuck query

Use KILL QUERY WHERE query_id = '<query_id>' on the node where the query is executing. This sends a cancellation signal; the query stops at the next cancellation point. Read-only SELECT queries can usually be killed safely, though the client receives an error.

Warning: Be cautious with INSERT or ALTER operations. Killing an INSERT may leave partial data. Background mutations do not appear in system.processes; use KILL MUTATION WHERE ... in system.mutations instead.

Address spill-to-disk

If the query is slow because it is spilling, choose whether to let it finish or kill it and rewrite. To prevent recurrence, reduce GROUP BY cardinality, add filters, or tune max_bytes_before_external_group_by so spill is controlled rather than catastrophic. Tradeoff: more spill means more disk I/O and slower execution.

Fix the query plan

For full table scans caused by partition pruning failure, rewrite the WHERE clause to include the partition key. For Cartesian JOINs, add explicit ON conditions. See the related guide on full table scans.

Set guardrails with max_execution_time

Apply max_execution_time at the user profile level to cap interactive queries. Legitimate ETL can run under a separate profile with a higher limit or zero. Tradeoff: an overly aggressive limit kills valid long jobs.

Throttle heavy legitimate queries

If the query is valid but resource-intensive, run it under a dedicated user profile with lower max_threads to reduce CPU contention, or schedule it outside peak hours.

Prevention

  • Set per-user max_execution_time and max_memory_usage. Interactive users should have a low ceiling; ETL service accounts can have higher limits. This prevents accidental runaways from ever starting.
  • Review system.query_log weekly. Look for queries with high read_rows or memory_usage that are new or trending upward. Catch pattern changes before they become incidents.
  • Validate partition key filters in application queries. Missing partition keys are the most common cause of unexpected full scans that turn into long-running resource hogs.
  • Monitor system.processes as a leading indicator. An alert on elapsed > 300 catches hogs before they dominate node resources.
  • Size max_bytes_before_external_group_by deliberately. If you rely on spill-to-disk, set it so spills are predictable and do not fill the data disk.

How Netdata helps

  • ClickHouse query latency and memory charts shown alongside host CPU, memory, and disk I/O confirm whether a long query is saturating the node.
  • Alerts on rising MemoryTracking alongside per-query memory spikes detect memory pressure cascades early.
  • Disk latency charts for the data volume confirm spill-to-disk I/O contention.
  • FailedQuery and exception code tracking show when long queries start failing across the cluster.
The Netdata solution

ClickHouse monitoring with Netdata

Netdata monitors ClickHouse with per-second metrics and ML anomaly detection. Track merge debt, memory usage, replication lag, Keeper/ZooKeeper saturation, and disk headroom against the host signals that drive them.