The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-health-ready-endpoint

Operations Guides

CockroachDB /health?ready=1: load balancer checks, draining, and impaired nodes

CockroachDB exposes a readiness endpoint at GET /health?ready=1 on its HTTP port (default 8080). Load balancers and Kubernetes readiness probes use it to decide whether to route SQL traffic to a node. The node returns HTTP 200 when ready and HTTP 503 when not. This signal separates “process is alive” from “this node should receive client connections.”

The most common operator mistake is using a plain TCP check against the SQL port (26257) or the plain /health endpoint instead of /health?ready=1. A TCP check succeeds as long as the port is listening, telling you nothing about whether the node is draining, write-stalled, or GC-thrashing. The plain /health endpoint returns 200 whenever the process is running, regardless of draining state. Neither is safe for routing decisions.

What the endpoint checks

The ?ready=1 query parameter transforms /health from a liveness check into a readiness check. The endpoint returns:

  • 200 OK (empty body) when the node is operational and ready to accept SQL connections.
  • 503 Service Unavailable (JSON error body) when the node is not ready to serve traffic.

Per the official documentation, 503 is returned when:

  1. The node is in the wait phase of the shutdown sequence (draining).
  2. The node is decommissioning or decommissioned.
  3. The node cannot communicate with a majority of other nodes (cluster unavailability).

The “cannot communicate with a majority” condition is implemented through node liveness: since v23.2 the readiness check fails when the node’s own liveness record is expired, which happens when the node can no longer write heartbeats to the liveness range, i.e. it lost quorum connectivity. Issue #116194 (still open) asks for a stronger check that verifies whether the SQL gateway can reach a majority quorum of ranges beyond its own liveness, because a node with marginal liveness connectivity can still report 200 while some of its ranges are unreachable.

How it works

The readiness check is a lightweight HTTP request served by the node’s HTTP server. It does not test the full SQL path (TCP connect, TLS handshake, pgwire protocol, query execution). It checks whether the node’s internal state machine has reached a point where it believes it can serve clients.

flowchart TD
    A["LB / probe checks /health?ready=1"] --> B{Node draining?}
    B -->|yes| F["503 Service Unavailable"]
    B -->|no| C{Decommissioning?}
    C -->|yes| F
    C -->|no| D{SQL server started?}
    D -->|no| F
    D -->|yes| E["200 OK"]
    F --> G["LB removes node after N failed checks"]
    E --> H["LB routes SQL traffic"]

What the endpoint does not check:

  • Whether ranges on this node have quorum connectivity.
  • Storage engine health (write stalls, L0 sublevel overflow, compaction backlog).
  • Inter-node network latency or partition status.

A node can return 200 while experiencing write stalls, admission control throttling, or elevated GC pauses. The endpoint tells you the node believes it is ready, not that it is performing well. A synthetic SELECT 1 probe exercises the full client path and catches failures the readiness endpoint misses.

Where it shows up in production

Load balancer configuration

When using /health?ready=1, the check interval and failure threshold must be coordinated with CockroachDB’s shutdown timing.

During graceful shutdown, the node enters a drain phase. It stops accepting new client connections and transfers range leases. The server.shutdown.initial_wait cluster setting (previously server.shutdown.drain_wait, kept as an alias) controls how long the node waits before draining begins. The default is 0s, meaning the node starts draining immediately without giving the load balancer time to detect the 503 and stop routing new connections.

For HAProxy with default inter 2000 fall 3, the load balancer checks every 2 seconds and removes a backend after 3 consecutive failures. The node must return 503 for at least 6 seconds before the load balancer stops routing. With server.shutdown.initial_wait at 0s, new connections arrive at a draining node.

Set server.shutdown.initial_wait to at least your LB’s check interval multiplied by its fall threshold, plus margin. For the HAProxy example above, 8 seconds covers the 6-second detection window with a safety buffer.

Kubernetes readiness probes

The official CockroachDB StatefulSet manifest uses /health?ready=1 as the readiness probe path with failureThreshold: 2 and periodSeconds: 5. With these defaults, a pod is marked not-ready after roughly 10 seconds of 503 responses (2 failed checks at 5-second intervals). Ensure server.shutdown.initial_wait covers that window. The liveness probe is explicitly commented out in the manifest with the note: “We recommend that you do not configure a liveness probe on a production environment, as this can impact the availability of production databases.”

Set terminationGracePeriodSeconds to the sum of all server.shutdown.* timeouts, per the official recommendation; in most cases a value above 300 seconds is not required. If the grace period is too short, Kubernetes SIGKILLs the pod before CockroachDB finishes draining, causing connection resets and brief range unavailability.

CPU starvation and health check timeouts

Under severe CPU load, requests to the health endpoint can hang or timeout. This was reported as early as 2020 (issue #44832): a node near 100% CPU could take 20+ seconds to respond to /health?ready=1, causing probes to fail and mark the pod not-ready. It was partially addressed in v20.1 by making /health require no authentication and perform no KV operations (PR #45119), and further mitigated by the admission control work introduced experimentally in v21.2, which keeps an overloaded node serving critical work such as the health endpoint.

Follow the manifest’s guidance and leave the liveness probe disabled in production. A liveness probe failure triggers a pod restart, which causes lease transfers, connection resets, and brief unavailability. Restarting a node that is slow but not dead often makes the situation worse.

Tradeoffs and common misuses

TCP checks versus the readiness endpoint

A plain TCP check against port 26257 succeeds whenever the process is listening. It cannot detect:

  • A draining node (accepting no new connections).
  • A node with active write stalls (storage engine refusing writes).
  • A node in GC-thrashing oscillation (alternating between alive and unresponsive).
  • A node that has lost liveness but whose process is still running.

If your load balancer uses TCP checks, it will route traffic to impaired nodes during incidents.

Plain /health versus /health?ready=1

The /health endpoint without ?ready=1 returns 200 whenever the process is running. It does not reflect draining state. Using it as a readiness check means traffic continues to route to a draining node during graceful shutdown.

Readiness endpoint versus synthetic SQL probe

A SELECT 1 probe via cockroach sql -e "SELECT 1" exercises TCP connect, TLS handshake, pgwire protocol, authentication, and SQL execution. It catches TLS certificate issues, authentication problems, and pgwire-level errors that the HTTP readiness check misses.

The tradeoff is cost: a SQL probe opens a real connection and is more expensive to run frequently. A reasonable pattern is /health?ready=1 for the load balancer’s fast check (every 2-5 seconds) and a SELECT 1 probe for deeper monitoring (every 10-30 seconds).

The partitioned-node gap

/health?ready=1 does not detect network partitions. A node isolated from quorum can still report 200. For partition detection, correlate the readiness check with:

  • ranges_unavailable metric (nonzero means ranges have lost quorum).
  • Node liveness status changes.
  • Raft leader-not-found errors in logs.
  • Inter-node RPC latency spikes.

If you rely solely on the readiness endpoint for health routing, a partitioned node will keep receiving traffic it cannot serve.

Signals to watch in production

SignalWhy it mattersWarning sign
/health?ready=1 response codeDirect readiness signal for LB routing200 to 503 transition outside planned maintenance
SELECT 1 probe latencyTests the full SQL path the readiness endpoint skipsLatency exceeding 5s, or intermittent failures
ranges_unavailableDetects quorum loss the readiness endpoint cannotAny nonzero value sustained beyond 5 minutes
Node liveness statusCluster-level view of node participationUnexpected transition to not-live, or flapping
CPU utilization per nodeCPU starvation can stall health check responsesSustained above 70-80% correlating with probe timeouts
storage_write_stallsWrite-stalled nodes appear ready but cannot serve writesRate above 1/sec sustained beyond 1 minute
round_trip_latency (RPC)Inter-node network health; partitions affect quorumAny node pair exceeding 5x baseline sustained
leases_transfers_successElevated rate indicates nodes losing and regaining leasesMore than 10x baseline without an operational cause

How Netdata helps

  • Per-second metrics collection captures liveness transitions, write stalls, and RPC latency spikes at the resolution they occur, not 15 or 30 seconds after the fact.
  • Correlate /health?ready=1 probe results with ranges_unavailable, storage_write_stalls, and node liveness on the same timeline. If the readiness endpoint returns 200 but ranges are unavailable or writes are stalling, the node is impaired despite its own self-assessment.
  • Per-node CPU utilization and Go GC pause metrics distinguish “health check failed because the node is unhealthy” from “health check timed out because the HTTP server was CPU-starved.”
  • Admission control queue depth signals when the system is at capacity, which often precedes degradation that eventually surfaces as health check failures.

Netdata’s CockroachDB monitoring brings these signals together with per-second metrics and anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.