The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-transaction-retry-rate-high

Operations Guides

CockroachDB transaction retry rate high: breaking down txn_restarts by cause

When application latency creeps upward without an obvious cause, or you start seeing PostgreSQL error code 40001 (serialization failure) in application logs, the transaction retry rate is likely climbing. CockroachDB exposes txn_restarts as a set of counters, each tagged by cause. The aggregate number tells you retries are happening. It does not tell you why.

The five dominant causes are writetooold, serializable, readwithinuncertainty, txnaborted, and txnpush. Three additional sub-metrics (asyncwritefailure, commitdeadlineexceeded, unknown) cover edge cases. A spike dominated by readwithinuncertainty is a clock infrastructure problem. A spike dominated by writetooold is a schema or application contention problem. The fix for one does nothing for the other.

For well-designed OLTP workloads, retry rates below 2% of total transactions are typical. The 2-10% range warrants investigation. Above 10%, you are burning significant resources on retry overhead and effective throughput is lower than it appears.

What this means

CockroachDB uses serializable isolation by default. Every transaction gets a timestamp from the node’s Hybrid Logical Clock. When two concurrent transactions touch overlapping keys, one may need to restart at a new timestamp. Some restarts happen transparently (automatic retry for implicit single-statement transactions); others surface as error code 40001 to the application.

Each restart cause maps to a different failure mode:

  • writetooold: A transaction tried to write a key that another transaction already wrote at a higher timestamp. Most common under write contention. Typically indicates hot keys or sequential key patterns.
  • readwithinuncertainty: A read on one node encountered a write from another node within the uncertainty window created by clock offset. Nearly diagnostic of clock skew. Any sustained nonzero rate means NTP needs attention.
  • serializable: A write-write conflict where the writer’s timestamp was pushed, invalidating its prior reads. Encompasses high-priority transactions pushing lower-priority ones and long-running transactions hitting the closed timestamp interval.
  • txnaborted: The transaction was aborted, typically due to priority-based deadlock resolution where two transactions hold intents on keys the other needs.
  • txnpush: A transaction conflict where one transaction needs to push another’s timestamp forward to resolve a read-write or write-write conflict.

Break down the total before investigating.

flowchart TD
    A["txn_restarts elevated"] --> B{"Dominant cause?"}
    B -->|"writetooold"| C["Contention: schema or app design"]
    B -->|"readwithinuncertainty"| D["Clock skew: NTP infrastructure"]
    B -->|"serializable"| E["Write conflicts: contention events"]
    B -->|"txnaborted"| F["Deadlocks: lock ordering"]
    B -->|"txnpush"| G["App tx conflicts: patterns"]

Common causes

CauseWhat it looks likeFirst thing to check
writetoooldDominates under write contention. Hot keys, sequential primary keys, single-row counters.Hot Ranges page or crdb_internal.ranges ordered by QPS
readwithinuncertaintyAppears when clock offset between nodes is elevated. Often correlates with clock_offset_meannanos rising.Clock offset metrics and NTP status on all nodes
serializableWrite-write conflicts, intent pushes by higher-priority transactions, long-running transactions exceeding closed timestamp.crdb_internal.transaction_contention_events
txnabortedTransactions acquiring locks in reverse order, causing deadlock resolution.Application transaction patterns and lock ordering
txnpushRead-write or write-write conflicts requiring timestamp pushes. Often from concurrent transactions on overlapping key ranges.Concurrent transaction patterns and query isolation
commitdeadlineexceededLong-running transactions whose commit deadline expires before completion.Transaction duration and kv.closed_timestamp.target_duration setting

Quick checks

These commands assume localhost:8080 is the CockroachDB Admin UI / Prometheus endpoint. For secure clusters, add --cert flags as appropriate.

# Aggregate retry rate and all sub-metrics
curl -s http://localhost:8080/_status/vars | grep 'txn_restarts'
# Clock offset on this node relative to peers
curl -s http://localhost:8080/_status/vars | grep 'clock_offset'
# Intent accumulation (correlates with contention-driven retries)
curl -s http://localhost:8080/_status/vars | grep -E 'intentcount|intentbytes'
# NTP synchronization status
chronyc tracking
# Transaction commit and abort counts for rate calculation
curl -s http://localhost:8080/_status/vars | grep -E 'sql_txn_commit_count|sql_txn_abort_count'
-- Find the most recent contention events with blocking transaction details
SELECT * FROM crdb_internal.transaction_contention_events
ORDER BY contention_duration DESC LIMIT 50;

crdb_internal.node_txn_stats has no restart_count column. It exposes per-application transaction counts and timing (node_id, application_name, txn_count, txn_time_avg_sec, committed_count); restart counts by cause come from the txn_restarts metric family:

-- Compare transaction volume and commit rate per application (node-local, in-memory)
SELECT node_id, application_name, txn_count, committed_count
FROM crdb_internal.node_txn_stats
ORDER BY txn_count DESC LIMIT 20;

Note: crdb_internal.transaction_contention_events requires admin privileges and performs an expensive cluster-wide RPC. Use for diagnosis only, not continuous monitoring.

How to diagnose it

  1. Break down the retry rate by cause. Pull txn_restarts sub-metrics and identify which cause dominates. If multiple causes are elevated, focus on the largest contributor first.

  2. If writetooold dominates, check for hot ranges. Use the DB Console Hot Ranges page, which ranks ranges by QPS (crdb_internal.ranges does not expose a QPS column). A range with 10x the average QPS indicates a hot spot. Check whether the affected table uses sequential primary keys (SERIAL, auto-increment, timestamp-prefixed).

  3. If readwithinuncertainty dominates, check clock synchronization. Query clock_offset_meannanos for all node pairs. Any node above 50% of --max-offset (default 500ms, so above 250ms) is a problem. Run chronyc tracking or ntpstat on all nodes. This cause does not occur for other reasons in meaningful quantities.

  4. If serializable dominates, inspect contention events. Query crdb_internal.transaction_contention_events to find which transactions are blocking each other. Look for the blocking transaction fingerprint and the contended keys. This cause encompasses multiple sub-scenarios that the metric alone cannot distinguish.

  5. If txnaborted dominates, look for deadlocks. The application is likely acquiring locks in reverse order across concurrent transactions. Check whether the same tables are accessed in different orders by different transaction paths.

  6. If txnpush dominates, review application transaction patterns. Look for transactions that read then write the same keys concurrently. SELECT FOR UPDATE can convert read-write conflicts into explicit locks that serialize more predictably.

  7. Correlate with intent accumulation. Check intentcount and intentbytes. Growing intent counts mean abandoned or long-running transactions are leaving unresolved write intents. Other transactions encountering these intents may need to restart.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
txn_restarts by causeEach cause maps to a different layer and requires a different fixAny cause exceeding its baseline or a new cause appearing
clock_offset_meannanosDirectly causes readwithinuncertainty restartsAny node above 50% of --max-offset
intentcount, intentbytesUnresolved intents cause contention and restartsGrowing monotonically without corresponding active transaction increase
SQL P99 latencyRetries add directly to tail latencyP99 rising while P50 is stable (bimodal from retry overhead)
txn_durations P99Transaction commit latency includes retry timeRatio of txn latency to statement latency diverging from 1.0
Per-range QPSIdentifies hot ranges driving writetooold restartsAny range at 10x average QPS
SQL error rate (40001)Serialization failures visible to applicationsSustained nonzero rate in production
sql_txn_abort_countTransactions aborted entirely, not just restartedRate increasing

Fixes

writetooold: reduce write contention

Short-term: Use ALTER TABLE ... SPLIT AT to manually split hot ranges at known boundaries. This distributes traffic across more leaseholders. This is a schema operation that takes effect immediately.

Long-term: Redesign primary keys to use hash-prefixed or UUID-based values instead of sequential patterns. For counter-like patterns, consider application-side batching.

Interleaved tables are deprecated and unsupported in current versions; since v22.1, upgrades are blocked if interleaved tables or indexes exist (cockroachdb/cockroach#68074). Do not use them in new schemas.

Query-level: Use SELECT ... FOR UPDATE to explicitly lock rows before updating, converting implicit conflicts into explicit locks. This can reduce restart overhead when multiple transactions compete for the same keys.

readwithinuncertainty: fix clock synchronization

Fix NTP on all affected nodes. Use cloud-specific time services where available (AWS Time Sync Service, Google Cloud time). After fixing NTP, monitor clock_offset_meannanos convergence. Time correction may be gradual (slew), not instantaneous.

If you trust your NTP infrastructure and still see persistent low-level readwithinuncertainty restarts, reducing --max-offset tightens the uncertainty window. This is a tradeoff: better read performance but higher risk of node self-termination if clock drift spikes. Changing --max-offset requires a rolling restart of all nodes.

See CockroachDB clock_offset_meannanos high: catching clock drift before self-termination and CockroachDB clock skew cascade: how shared NTP drift causes quorum loss for deeper coverage.

serializable: reduce transaction scope and conflicts

Identify the blocking transaction from crdb_internal.transaction_contention_events. Common fixes include:

  • Shrinking transaction scope (fewer statements, shorter duration)
  • Adding indexes to reduce full-table scans that hold locks
  • Using historical reads (AS OF SYSTEM TIME) for read-only parts of a transaction to avoid conflicts

For long-running transactions hitting the closed timestamp, either reduce transaction duration or increase kv.closed_timestamp.target_duration. The latter has tradeoffs for follower reads and CDC latency.

txnaborted: fix lock ordering

Identify which transactions are deadlocking by checking which keys they hold intents on in reverse order. Fix the application to acquire locks in a consistent order across all transaction paths. In practice, this means ensuring all code paths access the same tables and rows in the same sequence.

txnpush: reduce transaction conflict surface

Review application transaction patterns for concurrent read-then-write cycles on overlapping keys. Options include:

  • Adding SELECT ... FOR UPDATE to serialize access to contested keys explicitly
  • Reducing transaction width (fewer keys touched per transaction)
  • Scheduling conflicting workloads to avoid peak concurrency windows

READ COMMITTED isolation (GA since v24.1) transparently retries individual statements and removes RETRY_SERIALIZABLE errors, which can reduce retry overhead for conflict-prone workloads. Caveats: it requires an enterprise license (without one, READ COMMITTED transactions are upgraded to SERIALIZABLE), and it still has statement-level retry limits and tradeoffs, so validate it against your workload before standardizing on it.

Prevention

Monitor by cause, not just total. Alert on individual restart causes with different thresholds. readwithinuncertainty should alert at any sustained nonzero rate. writetooold and txnpush can tolerate higher rates but should be trended. A sudden shift in the cause distribution is itself a signal.

Track retry rate as a fraction of total transactions. The absolute count matters less than the ratio. Use sql_txn_commit_count as the denominator. Alert when the ratio exceeds 10% sustained.

Watch for gradual increases. A retry rate that creeps from 1% to 3% to 5% over weeks indicates worsening contention. The system feels fine until it cascades. Trend the per-cause rates, not just the total.

Instrument application-side retries. The txn_restarts metric captures CockroachDB’s internal restarts, including transparent automatic retries. Application-side retries (reconnect and re-execute after receiving 40001) are not counted. The true contention cost may be higher than the database reports.

Review schema for sequential patterns. Any monotonic primary key (SERIAL, auto-increment, timestamp-prefixed) is a future hot range. Address these proactively before traffic grows.

How Netdata helps

  • Per-second collection captures transient retry spikes that 15-30 second scrape intervals miss. Brief contention bursts that cause application-visible 40001 errors can appear and disappear between coarser scrapes.
  • txn_restarts sub-metrics collected independently means you see each cause as its own time series without manual PromQL or separate alert rules.
  • Correlation between restart causes and clock offset lets you confirm or rule out NTP as the source of readwithinuncertainty restarts within seconds.
  • Anomaly detection per cause flags distribution shifts even when the total rate has not crossed a fixed threshold. A first appearance of readwithinuncertainty where it was previously zero is an early signal.
  • Correlation with intent count, SQL P99 latency, and per-node CPU helps distinguish contention-driven retries from overload-driven retries during an active incident.

See CockroachDB monitoring with Netdata for per-second metrics and anomaly detection.

The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.