The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-certificate-expired

Operations Guides

CockroachDB certificate expired: TLS handshake failures and online rotation

CockroachDB uses mutual TLS for every connection: node-to-node, client-to-node, and admin UI. There is no plaintext fallback and no grace period. When a certificate expires, affected connections fail immediately.

Node certificate expiry prevents nodes from completing TLS handshakes with each other, which can look like a network partition or quorum loss. Client certificate expiry prevents applications from connecting. CA certificate expiry invalidates the entire trust chain at once, breaking every connection simultaneously.

CockroachDB defaults certificate lifetimes to 87840h (10 years), so expiry often surfaces years after initial setup, when the original operator has moved on and nobody configured monitoring for it. CA certificate expiry is the most commonly forgotten item in multi-year deployments.

Three certificate types exist and expire independently:

  • Node certificates are presented by each node to other nodes and to SQL clients. Expired node certs cause inter-node TLS handshake failures, ranges_unavailable spikes, and potential quorum loss.
  • Client certificates are presented by SQL clients (applications, CLI tools). Expired client certs cause client connection failures with no inter-node errors.
  • CA certificate signs all node and client certificates. An expired CA breaks the trust chain for every certificate it signed, causing complete cluster isolation.
flowchart TD
    A["Certificate expired"] --> B{"Which cert type?"}
    B -->|Node cert| C["Inter-node TLS failures"]
    B -->|Client cert| D["Client connections fail"]
    B -->|CA cert| E["All connections fail"]
    C --> F["Ranges unavailable"]
    C --> G["Lease transfers spike"]
    C --> H["Quorum loss possible"]
    D --> I["sql_conn_failures rises"]
    D --> J["New connections rejected"]
    E --> K["Complete cluster isolation"]
    F --> L["Fix: SIGHUP rotation"]
    I --> M["Fix: rotate cert, restart clients"]
    K --> N["Fix: CA rotation, combined cert"]

The characteristic log signature:

http: TLS handshake error from <addr>: remote error: tls: bad certificate
serving SQL client conn: remote error: tls: bad certificate

Common causes

CauseSymptomsFirst check
Node certificate expiredInter-node communication fails, nodes appear offline, ranges_unavailable rises, quorum loss on affected rangescockroach cert list --certs-dir=<dir>
Client certificate expiredApplications cannot connect, sql_conn_failures rising, no inter-node errorsopenssl x509 -in client.root.crt -enddate -noout
CA certificate expiredAll connections fail simultaneously (inter-node and client), complete cluster isolationopenssl x509 -in ca.crt -enddate -noout
Key file permissions wrongNode fails to start or cannot reload certs, permission denied on .key filesls -la *.key in certs directory

Quick checks

# List all certificates with file paths, usage, expiry dates, and errors
cockroach cert list --certs-dir=/path/to/certs

# Check expiry of specific certificate files
openssl x509 -in /path/to/certs/node.crt -enddate -noout
openssl x509 -in /path/to/certs/ca.crt -enddate -noout
openssl x509 -in /path/to/certs/client.root.crt -enddate -noout

# Check certificate expiration metrics from the Prometheus endpoint
curl -s http://localhost:8080/_status/vars | grep -E 'security_certificate_expiration_(ca|node)'

# Search for TLS handshake errors in CockroachDB logs
grep "TLS handshake error\|bad certificate" /path/to/logs/cockroach.log

# Check key file permissions (CockroachDB enforces strict ownership)
ls -la /path/to/certs/*.key

Diagnosis

  1. Identify the failing connection type. Inter-node errors point to node or CA cert problems. Client-only errors (with healthy inter-node communication) point to client cert problems.

  2. List all certificates with expiry dates. Run cockroach cert list --certs-dir=<dir> on any node. Compare each expiry date against the current date.

  3. Check the CA certificate specifically. CA expiry is easy to miss in long cockroach cert list output. Run openssl x509 -in ca.crt -enddate -noout directly.

  4. Check Prometheus metrics. Current versions expose the expiration timestamps as separate metrics: security_certificate_expiration_ca for the CA certificate and security_certificate_expiration_node for the node certificate (older releases used a single security_certificate_expiration metric with a type label). There is no built-in metric for an individual client certificate’s expiration; track client cert expiry externally.

  5. Verify key file permissions. CockroachDB enforces strict key file ownership: the key must be owned by the process user or root, with group ownership matching the process group. Wrong permissions cause rotation to fail even with valid certificates.

  6. Check all nodes, not just one. If certificates were provisioned at different times or partially rotated, different nodes may have different expiry dates.

Metrics and signals

SignalWhy it mattersWarning sign
security_certificate_expiration_caCA cert expiration epoch timestampValue approaching current epoch time
security_certificate_expiration_nodeNode cert expiration epoch timestampValue approaching current epoch time
TLS handshake errors in logsDirect evidence of certificate rejectionbad certificate or TLS handshake error messages
sql_conn_failuresClient connection failuresSustained increase correlating with cert expiry timeline
ranges_unavailableRanges without a leaseholder, indicating inter-node failureNonzero value with simultaneous cert expiry
Node liveness statusNodes marked dead due to failed TLS handshakesMultiple nodes losing liveness simultaneously

Fixes

Online rotation via SIGHUP (node and CA certificates)

CockroachDB reloads certificate files from the certs directory on SIGHUP, without restarting. This works for both node certificates and CA certificates.

  1. Generate new certificates using cockroach cert create-node or your PKI tooling.
  2. Replace the certificate files in the certs directory (same filenames).
  3. Send SIGHUP to the cockroach process:
# Reload certificates without restarting the node
kill -HUP $(pgrep -x cockroach)
  1. Verify the new certificate is loaded: run cockroach cert list --certs-dir=<dir> and confirm the updated expiry date. Test with a SQL query.

Rotate one node at a time. Confirm each node has reloaded successfully before moving to the next to avoid simultaneous reloads across the cluster.

CA certificate rotation

CA rotation requires a multi-step process because nodes and clients with old certificates must still be verified during the transition.

  1. Create a combined CA certificate file with the new CA cert first, followed by the old CA cert. This allows CockroachDB to verify both old and new certificates during the transition.
  2. Replace ca.crt with the combined file on all nodes.
  3. Send SIGHUP to each node sequentially to reload the combined CA.
  4. After all nodes have the combined CA, rotate node certificates (signed by the new CA) via SIGHUP on each node.
  5. Rotate client certificates (signed by the new CA). Client certs require the client process to restart, not just SIGHUP on the server.
  6. After all nodes and clients have certificates signed by the new CA, remove the old CA cert from the combined ca.crt file and SIGHUP again.

CA rotation must complete before node and client cert rotation. Clients only pick up new CA certificates after restart, so plan the transition with adequate lead time.

Client certificate rotation

SIGHUP only reloads certificates on the CockroachDB server process, not on client applications. Each application instance must be restarted after the new client certificate is placed on disk.

There is no built-in metric for client certificate expiration. Track client cert expiry externally using openssl x509 -enddate -noout or your PKI tooling.

Operator-managed deployments (Kubernetes)

In operator-managed deployments, SIGHUP-based online rotation is not available. Pods must be restarted to pick up new certificates, whether those certificates are generated automatically (by the self-signer utility or cert-manager) or manually.

Plan for rolling pod restarts during certificate rotation. Ensure the rolling restart does not cause quorum loss by waiting for each pod to rejoin and stabilize before restarting the next.

For cert-manager-managed certificates, ensure fsGroup is set correctly on the volume mount so CockroachDB can read the key files with the expected ownership.

Key file permission issues

CockroachDB enforces strict key file ownership. The key file must be owned by the process user or root, with group ownership matching the process group.

To temporarily bypass permission checks during debugging:

export COCKROACH_SKIP_KEY_PERMISSION_CHECK=true

Do not use this in production. Fix the actual file ownership and mode on the key files.

Prevention

  • Monitor security_certificate_expiration_ca and security_certificate_expiration_node. Alert at 30 days remaining (plan rotation) and at less than 24 hours remaining with no replacement staged (urgent).
  • Track client certificate expiry externally. Run a scheduled job that checks expiry dates with openssl x509 -enddate -noout and alerts before expiry.
  • Alert before expiry, not at expiry. CockroachDB has no grace period. Alert days or weeks in advance.
  • Document the CA certificate. Record its creation date, expiry date, and rotation procedure in your runbook. The CA cert is the most commonly forgotten item in multi-year deployments.
  • Test rotation before expiry. Rotate certificates well before they expire so you have time to fix issues. Discovering that operator-managed deployments do not support SIGHUP at 3 AM under incident pressure is avoidable.

How Netdata helps

  • Per-second collection of security_certificate_expiration_ca and security_certificate_expiration_node keeps the expiry countdown visible rather than only surfacing during a scheduled check.
  • ML anomaly detection on sql_conns and sql_conn_failures can flag unexpected connection drops that correlate with certificate problems, even when the certificate metric itself was not being watched.
  • Cross-signal correlation between TLS handshake errors in logs, ranges_unavailable spikes, and node liveness changes confirms the failure is certificate-related rather than a network partition or disk stall.
  • Node-level visibility shows per-node certificate states, which matters when certificates were provisioned at different times or partially rotated.

Netdata’s database monitoring brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.