The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-transaction-retry-serializable

Operations Guides

CockroachDB TransactionRetryWithProtoRefreshError: RETRY_SERIALIZABLE causes and fixes

The application logs show TransactionRetryWithProtoRefreshError: TransactionRetryError: retry txn (RETRY_SERIALIZABLE - failed preemptive refresh due to a conflict: committed value on key /Table/.../0). SQLSTATE is 40001. Clients are timing out or failing outright. The cluster looks healthy: nodes are live, ranges are available, disk space is fine. The problem is in the transaction layer.

Under SERIALIZABLE isolation (CockroachDB’s default), the database uses optimistic concurrency. It proceeds as if no conflict exists, then detects conflicts at commit or statement boundary. When a transaction’s preemptive refresh fails because another transaction already committed a write to a key the first transaction read, CockroachDB emits RETRY_SERIALIZABLE. This is expected behavior, not a bug.

The critical mistake teams make is treating all RETRY_SERIALIZABLE errors as one problem. CockroachDB tags transaction restarts by cause: writetooold, readwithinuncertainty, and txnpush are the three dominant categories behind RETRY_SERIALIZABLE. Each points to a different layer and requires a different fix. Alarming on aggregate retry rate without breaking it down wastes diagnostic time.

What this means

RETRY_SERIALIZABLE fires when a transaction cannot be refreshed at its current timestamp. In CockroachDB’s MVCC model, each transaction gets a timestamp from the node’s Hybrid Logical Clock. When a concurrent transaction commits a write to a key this transaction has already read, the read becomes stale. The transaction must either advance its timestamp and re-read, or abort.

If the transaction has already performed writes, advancing the timestamp requires a refresh: proving that no other transaction has written to the keys this transaction touched. If that proof fails, the transaction gets RETRY_SERIALIZABLE. The meta={... key=/Table/...} field in the error identifies the transaction’s anchor key (the first key it wrote), not necessarily the key where the conflict occurred. This is a common misinterpretation that sends operators looking at the wrong table or range.

Under implicit transactions (single statements), CockroachDB retries automatically and the error is invisible to the application. Under explicit transactions (BEGIN ... COMMIT blocks), the client must handle the retry. ORM frameworks (Hibernate, Ecto, Prisma) often hide transaction boundaries, making it unclear when retries are needed or whether they are being handled at all.

Common causes

CauseWhat it looks likeFirst thing to check
writetooold (contention)Retry rate correlates with traffic spikes. Hot ranges on specific tables. One leaseholder node saturating while others idle.txn_restarts sub-metrics and per-range query rates
readwithinuncertainty (clock skew)Retries appear across random keys. Clock offset metric elevated. Read latency P99 climbing without write contention.clock_offset_meannanos and NTP status on all nodes
txnpush (transaction conflicts)Retries correlate with long-running transactions. Intent count elevated. Aborted transactions leaving dangling intents.crdb_internal.transaction_contention_events and intent metrics

Quick checks

# Check transaction restart breakdown by cause
curl -s http://localhost:8080/_status/vars | grep 'txn_restarts'

# Check clock offset between nodes
curl -s http://localhost:8080/_status/vars | grep 'clock_offset'

# Check commit vs abort ratio
curl -s http://localhost:8080/_status/vars | grep -E 'sql_txn_(commit|abort)_count'

# Check intent accumulation
curl -s http://localhost:8080/_status/vars | grep intent

# Check transaction latency (retries inflate this)
curl -s http://localhost:8080/_status/vars | grep 'txn_durations'

For contention diagnosis, run this SQL. This performs a cluster-wide RPC, so use it for diagnosis only, not continuous monitoring:

SELECT * FROM crdb_internal.transaction_contention_events
ORDER BY contention_time DESC LIMIT 50;

For clock skew diagnosis:

SELECT * FROM crdb_internal.gossip_nodes;

How to diagnose it

flowchart TD
    A["40001 RETRY_SERIALIZABLE"] --> B{"Check txn_restarts by cause"}
    B -->|"writetooold dominant"| C["Contention / schema"]
    B -->|"readwithinuncertainty dominant"| D["Clock skew / infra"]
    B -->|"txnpush dominant"| E["Transaction conflicts / app"]
    C --> F["Check hot ranges,
index design, key patterns"] D --> G["Check clock_offset,
NTP sync on all nodes"] E --> H["Check transaction scope,
retry loops, intent cleanup"]
  1. Break down retry rate by cause. Pull txn_restarts sub-metrics. The dominant cause determines your diagnostic path. A retry rate below 2% of total transactions is typical for well-designed OLTP. Above 10% warrants investigation. The cause breakdown matters more than the absolute rate.

  2. If writetooold dominates: Look for hot ranges. Query crdb_internal.ranges (admin-only, version-sensitive, expensive) to find ranges with disproportionate traffic, ordering by queries_per_second. Check the primary key design of affected tables for sequential or monotonic patterns (SERIAL, auto-increment, timestamp-prefixed). Compare per-node CPU utilization: one saturated node with others idle confirms a hot range.

  3. If readwithinuncertainty dominates: Check clock_offset_meannanos across all node pairs. Any sustained offset above 100ms is concerning. The default --max-offset is 500ms; when a node detects its clock offset exceeds 80% of this threshold (400ms) relative to the majority, it self-terminates to prevent data corruption. Verify NTP/chrony is running and synced on every node: chronyc tracking or ntpstat. readwithinuncertainty restarts are nearly diagnostic of clock skew. They do not occur in meaningful quantities for any other reason.

  4. If txnpush dominates: Check for long-running transactions holding locks. Query active sessions for transactions with high elapsed time. Check intent count and intent bytes: growing values indicate abandoned transactions or intent resolution falling behind.

  5. Correlate with SQL P99 latency. Transaction retries inflate latency directly. If P99 is climbing and retry rate is climbing, the retries are the latency source. If retry rate is high but latency is stable, the application’s retry loop is absorbing the cost silently, wasting resources and masking a growing problem.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
txn_restarts by causeDirectly measures serialization conflict rate and typewritetooold above 5%, any sustained readwithinuncertainty
txn_durations (P99)Retries inflate transaction latencyP99 diverging from sum of statement latencies
sql_txn_abort_countTransactions that failed entirely, not just retriedRate trending upward
clock_offset_meannanosClock skew drives readwithinuncertainty restartsSustained above 100ms on any node pair
Intent count / bytesAbandoned intents cause txnpush conflictsGrowing monotonically without workload increase
sql_service_latency (P99)Contention manifests as tail latencyBimodal distribution (P50/P99 ratio widening)

Fixes

writetooold (contention)

Short-term:

  • Manually split hot ranges. Run ALTER TABLE ... SPLIT AT to distribute load across leaseholders. Identify the split point by checking where the hot range’s key boundaries fall. This provides immediate relief but is a band-aid.
  • Reduce transaction scope. Shorter transactions hold locks for less time, reducing the window for conflicts. Move non-database work outside transaction boundaries.

Long-term:

  • Redesign primary keys to avoid sequential patterns. Use UUID or hash-prefixed keys instead of auto-incrementing IDs. This is the highest-impact change for write-heavy tables with contention.
  • Add missing indexes. Queries that scan more rows than necessary take locks on data they do not need. Locking fewer rows means fewer conflicts.
  • Consider table partitioning for large tables with known access patterns.

Tradeoff: Splitting ranges manually works temporarily. The range will eventually merge or rebalance if the split point does not align with data distribution. Schema changes are the real fix but require application coordination and testing.

readwithinuncertainty (clock skew)

  • Fix NTP on all affected nodes. Check chronyc tracking or ntpstat. The correction may be gradual (slew mode), so monitor the offset metric over time.
  • Use cloud-specific time services. For VMs, prefer AWS Time Sync Service or Google Cloud time services over generic NTP pools. VMs are prone to clock drift, especially after live migration.
  • Consider tightening --max-offset to narrow the uncertainty window. Warning: reducing --max-offset increases the risk of node self-termination from transient clock drift. Do not change this without testing.

Tradeoff: Clock skew is an infrastructure problem, not a database tuning problem. The database can only report it and protect itself. The fix is always in the NTP/chrony configuration and the underlying host time sources.

For deeper coverage of clock-related failures, see CockroachDB clock_offset_meannanos high: catching clock drift before self-termination and CockroachDB clock skew cascade: how shared NTP drift causes quorum loss.

txnpush (application transaction conflicts)

  • Implement client-side retry loops. For explicit transactions, wrap the transaction in a retry loop using CockroachDB’s SAVEPOINT cockroach_restart mechanism. On receiving error code 40001, issue ROLLBACK TO SAVEPOINT cockroach_restart and re-execute the entire transaction body. The exact implementation depends on your client library. Many PostgreSQL drivers support this pattern.

  • Cancel abandoned transactions. Find sessions holding long-running transactions via SHOW SESSIONS and cancel them with CANCEL SESSION. Check intent metrics afterward to confirm intent resolution is keeping up.

  • Reduce transaction duration. Move non-database work (API calls, file I/O, computation) outside transaction boundaries. Shorter transactions conflict less and resolve faster.

  • Evaluate READ COMMITTED isolation. READ COMMITTED isolation has been available since v23.2. Set the isolation level to READ COMMITTED for transactions that cannot tolerate serialization retries. READ COMMITTED does not return RETRY_SERIALIZABLE errors to the client. It transparently resolves serialization conflicts by retrying individual statements internally. Enable per transaction with SET TRANSACTION ISOLATION LEVEL READ COMMITTED. The sql.txn.read_committed_isolation.enabled cluster setting defaults to true in current versions, making READ COMMITTED available without additional configuration.

Tradeoff: READ COMMITTED provides weaker isolation guarantees than SERIALIZABLE. It is appropriate for workloads where strict serializability is not required. It does not eliminate contention. It hides it from the client by retrying internally, which still consumes resources. For testing retry logic before production, set the inject_retry_errors_enabled session variable to true to force RETRY_SERIALIZABLE errors.

Prevention

  • Monitor retry rate by cause, not in aggregate. Break txn_restarts into its sub-metrics. Alert on writetooold above 5% sustained and on any sustained nonzero readwithinuncertainty. Teams that alarm on total retry rate without cause breakdown cannot distinguish a schema problem from a clock problem from an application problem.

  • Implement retry logic in every client that uses explicit transactions. This is non-negotiable for SERIALIZABLE isolation. Audit ORM transaction management: Hibernate, Ecto, and Prisma can hide transaction boundaries from the application layer.

  • Design schemas for low contention from the start. Avoid sequential primary keys on write-heavy tables. Use UUID or hash-distributed keys. Keep transactions short.

  • Monitor clock offset proactively. Do not wait for nodes to self-terminate. Alert on any node pair exceeding 100ms sustained. The readwithinuncertainty restart rate is a leading indicator that precedes the fatal max-offset threshold.

  • For workloads that cannot implement retry loops, evaluate READ COMMITTED isolation. Available since v23.2, READ COMMITTED eliminates RETRY_SERIALIZABLE at the cost of weaker isolation guarantees. Validate that your application can tolerate read anomalies under READ COMMITTED before switching.

Monitoring in Netdata

  • Per-second transaction restart metrics surface retry rate changes immediately, not at 15-30 second scrape intervals. The sub-metric breakdown (writetooold, readwithinuncertainty, txnpush) lets you classify the problem before opening a SQL shell.
  • Clock offset correlation appears alongside retry metrics on the same dashboard. When readwithinuncertainty restarts spike, the clock offset chart tells you immediately whether clock skew is the cause.
  • Intent count and transaction latency correlate with retry metrics, making it clear whether contention is driving latency or whether a different bottleneck (storage, CPU, network) is involved.
  • Per-node breakdown catches asymmetry: one node showing elevated writetooold retries while others are quiet indicates a hot range, not a workload-wide contention problem.
The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.