The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-sql-error-rate-xx000

Operations Guides

CockroachDB SQL error rate: XX000 internal errors and error-code triage

sql_failure_count is a single Prometheus counter with no error-code labels. A serialization conflict (40001) is expected noise in a contended workload. An internal error (XX000) is a database fault that needs immediate escalation. Treating the aggregate counter as one signal guarantees the wrong response.

This article covers how to break that aggregate apart by error code, when XX000 internal errors warrant paging, and how to triage the other common error classes.

What this means

sql_failure_count increments every time a SQL statement results in a planning or runtime error. It does not break down by error code natively. To understand what is failing, correlate the counter with per-statement statistics, application logs, or the DB Console Insights page.

CockroachDB follows the PostgreSQL error code standard. The distribution across codes reveals the failure type:

CodeClassMeaningSeverity
XX000internal_errorDatabase fault: corruption, assertion failure, software bugPAGE on sustained nonzero
40001serialization_failureTransaction conflict, retriable by clientTICKET on sustained increase
53200out_of_memorySQL memory budget exhaustedTICKET, may precede OOM
57014query_canceledStatement timeout or user-initiated cancelTICKET, not an internal error
08006connection_failureConnection-level failureTICKET

XX000 is the code that warrants paging. It means the database hit an unexpected internal condition: an assertion failure, a corrupted data structure, or a software bug. These errors are not client-retriable. A sustained nonzero rate means the database is actively malfunctioning.

The other codes represent workload or configuration issues. They need attention but do not indicate database faults.

Triage flow when sql_failure_count spikes:

flowchart TD
    A["sql_failure_count rising"] --> B["Break down by error code"]
    B --> C{"XX000 internal error?"}
    C -->|Yes| D["PAGE: database fault"]
    D --> D1["Capture stack trace from logs"]
    D1 --> D2["Check version for known bugs"]
    D2 --> D3["Escalate to Cockroach Labs support"]
    C -->|No| E{"40001 serialization?"}
    E -->|Yes| F["Check contention signals"]
    F --> F1["txn_restarts by cause"]
    E -->|No| G{"53200 out of memory?"}
    G -->|Yes| H["Check sql_mem_root_current"]
    H --> H1["Increase budget or optimize queries"]
    G -->|No| I{"57014 query canceled?"}
    I -->|Yes| J["Check timeout configuration"]
    I -->|No| K["Check connection errors"]

Common causes

CauseWhat it looks likeFirst thing to check
XX000 internal error (bug or corruption)Sustained nonzero internal errors with stack traces in DETAIL fieldCockroachDB version and known issues for your release
Serialization conflicts (40001)High retry rate, writetooold restart cause dominant, contention on hot keystxn_restarts broken down by cause
Memory budget exhaustion (53200)Errors correlate with sql_mem_root_current near limit, often large sorts or hash joinsSQL memory budget utilization
Query cancellation (57014)Errors correlate with statement timeouts, long-running queries hitting limitsApplication timeout settings
Connection failures (08006/08001)Clients unable to connect, may correlate with cert expiry or node losssql_conn_failures and node liveness

Quick checks

All commands are read-only and safe to run during an active incident.

# Check the overall SQL failure rate
curl -s http://localhost:8080/_status/vars | grep 'sql_failure_count'
# Check connection failure rate separately
curl -s http://localhost:8080/_status/vars | grep 'sql_conn_failures'
-- Identify which statement fingerprints are failing most often (last_error_code is set on failures)
SELECT key, count, last_error_code FROM crdb_internal.node_statement_statistics
WHERE last_error_code IS NOT NULL
ORDER BY count DESC LIMIT 20;
# Check transaction restart rate and cause breakdown
curl -s http://localhost:8080/_status/vars | grep 'txn_restarts'
# Check SQL memory pressure (53200 precursor)
curl -s http://localhost:8080/_status/vars | grep 'sql_mem_root_current'
# Check SQL throughput to calculate error rate as fraction of total
curl -s http://localhost:8080/_status/vars | grep -E 'sql_(select|insert|update|delete)_count'
# Search logs for XX000 internal errors with stack traces
# Adjust the path to match your actual cockroach-data location
grep -r "internal error" /path/to/cockroach-data/logs/

Note that crdb_internal tables are version-sensitive, require admin privileges, and should be used for manual diagnosis only. Do not build automated monitoring pipelines on them.

How to diagnose it

  1. Confirm the error rate is real. Check sql_failure_count rate over the last 5-10 minutes. A single transient error is noise. A sustained rate over more than 1 minute is actionable.

  2. Determine the error code distribution. The sql_failure_count counter does not break down by code. Use the DB Console Insights page (Failed Execution view), which shows the error code for each failed statement, or application-level error logging to identify which statements are failing and what error codes they return. The crdb_internal.node_statement_statistics table records last_error and last_error_code for failing statements; use the Insights page or application logs to map those fingerprints to specific SQLSTATE codes.

  3. If XX000: treat as a database fault. XX000 errors include a stack trace in the error DETAIL. If the error message does not contain a stack trace, it may not be a true XX000 internal error. Capture the full error text including the stack trace. Check your CockroachDB version against known issues on GitHub.

  4. If 40001: check contention depth. Pull txn_restarts broken down by cause. writetooold indicates write-write contention on hot keys. readwithinuncertainty indicates clock skew. txnpush indicates transaction conflicts. Each cause points to a different layer of the stack.

  5. If 53200: check memory pressure. Verify sql_mem_root_current against the configured --max-sql-memory. A single large query (hash join, sort) can exhaust the per-node budget and reject all concurrent queries.

  6. If 57014: check timeout configuration. These are client-initiated cancellations or server-side statement timeouts. They indicate queries running longer than allowed, which may point to plan regression, missing indexes, or storage degradation.

  7. If 08006/08001: check connectivity. These are connection-level errors, not statement execution errors. Check certificate validity, node liveness, and network reachability. Connection errors increment sql_conn_failures, not sql_failure_count.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
sql_failure_countPrimary error rate counterSustained nonzero rate, especially if error code is XX000
sql_conn_failuresConnection-level failures tracked separately from statement errorsSustained increase indicating network, cert, or auth issues
txn_restarts (by cause)Underlying contention or clock skew driving 40001 errorsreadwithinuncertainty nonzero (clock skew), writetooold rising (contention)
sql_mem_root_currentMemory pressure that precedes 53200 errorsSustained above 70% of --max-sql-memory
SQL statement latency P99Correlates with error causes such as contention and storage degradationP99 rising alongside error rate
Node uptimeCold-start gating to suppress false positives during restartsUptime under 10 minutes: suppress most alerts

Fixes

XX000 internal errors

XX000 errors indicate a database bug or corruption. You cannot fix them in the moment. The operational response is:

  • Capture the full error and stack trace from logs or the Insights page. This is essential for support.
  • Check your CockroachDB version against release notes and GitHub issues. XX000 bugs are often fixed in patch releases.
  • Upgrade if a fix exists. Do not attempt workarounds for assertion failures or vectorized engine panics.
  • Escalate to Cockroach Labs support with the stack trace, query fingerprint, and version number.
  • If corruption is suspected, look for explicit “consistency check failed” or “checksum mismatch” messages from the consistency checker subsystem. Do not act on loose log matches that merely contain the word “corrupt” in an unrelated context.

40001 serialization failures

These indicate contention, not a database fault. Responses include:

  • Identify hot keys via crdb_internal.transaction_contention_events.
  • Check for sequential primary key patterns causing hot ranges.
  • Ensure the application implements proper retry logic using SAVEPOINT cockroach_restart or client-side retry with exponential backoff.
  • Consider schema changes to reduce write conflicts (hash-prefixed keys, smaller transaction scopes).

53200 resource exhaustion

  • Increase --max-sql-memory if the node has available RAM.
  • Identify and optimize the query consuming disproportionate memory (large sorts, hash joins on unindexed columns).
  • Ensure the application is not issuing full table scans without proper predicates.

57014 query canceled

  • Review statement timeout settings.
  • Investigate why queries exceed the timeout: plan regression, storage latency, or lock contention.
  • Do not simply increase the timeout without understanding the root cause.

08006/08001 connection failures

  • Verify certificate validity on all nodes and clients.
  • Check that the node is live and accepting connections.
  • Verify network connectivity and firewall rules.
  • Ensure clients connect through a load balancer using /health?ready=1 for health checks, not plain TCP.

Prevention

  • Alert on error code class, not just the aggregate counter. A single sql_failure_count alert tells you something is wrong but not what. Correlate with the Insights page or application logs to identify the dominant error code.
  • Gate XX000 alerts on sustained duration. A single XX000 event may be a transient bug trigger. Page only on sustained nonzero XX000 rate for more than 1 minute.
  • Track your CockroachDB version against known issues. XX000 bugs are version-specific. Knowing your version narrows the diagnosis immediately.
  • Ensure client retry logic handles 40001 correctly. Serialization failures are expected in CockroachDB’s serializable isolation model. Clients that do not retry 40001 errors will surface them as application errors.
  • Monitor sql_conn_failures alongside sql_failure_count. Connection failures are tracked separately. Monitoring only sql_failure_count creates a blind spot for connectivity issues.
  • Use /health?ready=1 for load balancer health checks, not plain TCP checks. This prevents routing traffic to draining or impaired nodes that will generate connection errors.

How Netdata helps

  • Netdata collects sql_failure_count and sql_conn_failures per second, so you see error rate changes immediately rather than waiting for a longer scrape interval.
  • Correlating sql_failure_count spikes with txn_restarts by cause, sql_mem_root_current, and node liveness in a single view shortens error-code triage from minutes to seconds.
  • ML anomaly detection flags unusual error rate patterns, including subtle increases in serialization failures that a fixed threshold would miss.
  • Per-node breakdowns reveal whether errors are concentrated on one node (hot range, storage issue) or cluster-wide (schema problem, version bug).
  • Node uptime gating suppresses false-positive error alerts during rolling restarts and cold starts.

Netdata’s CockroachDB monitoring brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.