The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / cockroachdb / cockroachdb-clock-offset-high

Operations Guides

CockroachDB clock_offset_meannanos high: catching clock drift before self-termination

clock_offset_meannanos measures the mean clock offset between a CockroachDB node and its peers. When it climbs, you are on a path that ends in silent performance degradation from widened read uncertainty windows, or a node self-terminating to preserve data consistency.

The thresholds are unforgiving. CockroachDB uses a default --max-offset of 500ms. A node self-terminates when its mean offset exceeds 80% of that value (400ms) relative to at least half of its peers. But a constant 200ms offset, stable and below any threshold, still doubles the uncertainty interval for every read. Transactions silently restart more often, P99 read latency creeps up, and nobody suspects the clock.

What this means

CockroachDB uses Hybrid Logical Clocks (HLC) to order transactions. Each HLC timestamp combines physical wall-clock time with a logical counter. When a node reads data, it uses the clock offset to define an uncertainty window. If a write might have happened within that window, the read retries at a higher timestamp to guarantee serializable correctness.

Larger clock offsets widen the uncertainty window. More reads fall within it. The readwithinuncertainty transaction restart cause is nearly diagnostic of clock skew: it does not occur in meaningful quantities for any other reason.

One limitation of the metric: clock_offset_meannanos is the mean of signed offsets between this node and its peers. Positive and negative offsets cancel out. If one peer is +300ms ahead and another is -300ms behind, the mean reports near zero. The companion metric clock_offset_stddevnanos detects this: a high standard deviation with a low mean means outliers are being averaged away.

At 80% of max-offset (400ms at the default 500ms), CockroachDB’s safety mechanism fires. The node logs a clock synchronization error and self-terminates. It will not rejoin until the clock is corrected. If multiple nodes share the same NTP infrastructure and drift together, quorum loss across ranges follows.

flowchart TD
    A["NTP drift or VM clock stall"] --> B["clock_offset_meannanos rises"]
    B --> C["Uncertainty interval widens"]
    C --> D["readwithinuncertainty restarts climb"]
    D --> E["SQL read P99 increases"]
    B --> F{"Offset > 400ms?"}
    F -->|Below threshold| G["Silent: 200ms doubles uncertainty window"]
    F -->|Above threshold| H["Node self-terminates"]
    H --> I["Multiple nodes drift? Quorum loss"]

Common causes

CauseWhat it looks likeFirst thing to check
NTP daemon stopped or misconfiguredOffset rising steadily on one node while peers stay stablechronyc tracking on the affected node
VM live migration (vMotion)Sudden offset spike immediately after a migration eventHypervisor migration logs, chronyc tracking
Shared NTP server unreachableMultiple nodes drifting simultaneously, often same AZ or regionchronyc sources -v from affected nodes
Cloud time service regressionOffset creeping on all nodes in a specific regionCloud provider status page, NTP source reachability
Hardware clock failureOffset rising on one physical node despite NTP runningdmesg for RTC errors, hwclock --show

Quick checks

Run these on the node exhibiting drift or from a host with access to the cluster’s HTTP endpoint.

# Check chrony synchronization status on the affected node
chronyc tracking

# Check clock offset metrics from the local node
curl -s http://localhost:8080/_status/vars | grep clock_offset

# Check readwithinuncertainty restart rate (nearly diagnostic for clock skew)
curl -s http://localhost:8080/_status/vars | grep readwithinuncertainty

# Check NTP sources and their reachability
chronyc sources -v

# Verify node liveness status across the cluster
curl -s http://localhost:8080/_status/nodes | python3 -c "
import json, sys
for n in json.load(sys.stdin)['nodes']:
    print(f'Node {n[\"desc\"][\"node_id\"]}: liveness={n.get(\"liveness\",{}).get(\"liveness\",\"UNKNOWN\")}')"

How to diagnose it

  1. Identify which nodes are drifting. Pull clock_offset_meannanos and clock_offset_stddevnanos for every node. The metric is per remote node, so each node reports its offset relative to every peer. Look for the outlier, not just the aggregate. If node A reports high offset to node B, but node B reports normal offset to everyone else, node A is the problem.

  2. Check NTP status on the affected node. Run chronyc tracking. If “Leap status” is not “Normal” or “Last offset” is large, the clock is not synchronized. Run chronyc sources -v to verify NTP source reachability.

  3. Correlate with readwithinuncertainty restarts. Pull the restart breakdown from the txn_restarts metric family and look for the readwithinuncertainty cause. Any sustained nonzero rate is a clock synchronization signal, even if clock_offset_meannanos has not yet reached the 250ms alerting threshold.

  4. Check for recent VM migrations. If nodes run on virtualized infrastructure, check the hypervisor migration log. VM live migration stalls the guest clock during the copy phase. After resumption, the clock jumps and offset metrics spike.

  5. Determine if the problem is shared infrastructure. If multiple nodes drift simultaneously, the root cause is likely shared NTP infrastructure or a cloud time service. Check NTP server reachability from all affected nodes, not just one.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
clock_offset_meannanos (per node pair)Directly measures clock drift between nodes> 250ms is dangerous, > 400ms is fatal
clock_offset_stddevnanos (per node pair)Detects outlier masking in the mean metricHigh stddev with low mean means offsets are canceling
readwithinuncertainty restart countNearly diagnostic of clock skewAny sustained nonzero rate
SQL read latency P99Uncertainty restarts add latency to readsP99 rising without changes to workload or schema
Node liveness statusSelf-termination from clock skew triggers liveness changesUnexpected node transitions to not-live
Node uptimeGates false positives during cold start< 10 minutes may show post-restart clock correction spikes

Fixes

Fix NTP on the affected node

The immediate priority is restoring clock synchronization. If the NTP daemon has stopped, restart it. If it is running but cannot reach its sources, resolve the network or firewall issue blocking NTP traffic.

Use chrony, not ntpd or systemd-timesyncd. Cockroach Labs recommends chrony. ntpd is not included in RHEL 8+ or installed by default on recent Ubuntu. systemd-timesyncd lacks the slew-based correction precision CockroachDB requires.

After fixing NTP, clock correction is gradual (slew), not instantaneous. Monitor clock_offset_meannanos as it converges. Do not force a step change with chronyc makestep unless you understand the consequences: a sudden clock jump can cause a burst of readwithinuncertainty restarts before stabilizing.

Restart a self-terminated node safely

A node that self-terminated from clock skew will crash-loop if restarted before the clock is corrected. Before restarting:

  1. Verify chronyc tracking shows normal synchronization on the host.
  2. Confirm clock_offset_meannanos is converging toward baseline (under 100ms).
  3. Restart the CockroachDB process.

Cloud-specific clock configuration

Virtualized environments introduce clock stalls that bare metal does not. Common hardening steps:

  • AWS: Use the Amazon Time Sync Service (169.254.169.123). It implements leap second smearing. AWS also exposes a PTP hardware clock at /dev/ptp0 on supported instance types.
  • GCP: Disable live migration by setting the host maintenance policy to TERMINATE. GCE does not provide an uninterrupted clock to guest VMs during live migration.
  • Azure: Disable the Hyper-V time synchronization integration service. The native Azure time service does not implement leap second smearing. Use Google Public NTP or Amazon Time Sync Service instead, which both smear.

For VMs subject to live migration (vMotion), CockroachDB supports the --clock-device flag (Linux only) to bind the clock to a PTP hardware clock device. This prevents stale-clock reads after migration suspension. When using --clock-device, do not enable the server.clock.forward_jump_check_enabled cluster setting, as forward jumps are expected with PTP devices.

Adjusting max-offset (use with caution)

--max-offset is a node-start flag, not a cluster setting. Lowering it tightens the uncertainty window (better read performance, fewer restarts) but increases the risk of self-termination from minor clock drift. Cockroach Labs recommends lowering max-offset to 250ms for multi-region clusters using global tables, where tighter uncertainty bounds improve follower read freshness.

Raising max-offset reduces self-termination risk but widens the uncertainty interval, causing more read restarts and longer stale-read windows. This trades availability for correctness headroom and is rarely the right answer.

Changing --max-offset requires a rolling restart of each node. During the update, nodes can run with different --max-offset values, but only for the purpose of updating the setting across the cluster; once every node has restarted with the new flag, all nodes should run the same value.

Prevention

  • Alert on clock_offset_meannanos at 250ms (50% of max-offset). By the time offset reaches 400ms, self-termination is imminent.
  • Alert on readwithinuncertainty restarts at any sustained nonzero rate. This catches clock skew that has not yet reached the offset threshold but is already degrading performance.
  • Monitor clock_offset_stddevnanos alongside the mean. High standard deviation reveals outlier masking that the mean hides.
  • Use chrony on all nodes. Configure all nodes to use the same leap-second-smearing NTP sources.
  • Ensure all NTP sources implement leap second smearing. Google Public NTP and Amazon Time Sync Service smear. The default NTP pool does not. Leap seconds have historically caused mass CockroachDB node crashes.
  • Disable VM features that stall the guest clock. GCE live migration and Azure Hyper-V time synchronization both cause clock stalls. Use PTP clock devices where vMotion is unavoidable.
  • Gate clock offset alerts on node uptime greater than 10 minutes. Brief post-restart clock corrections cause transient spikes that should not page.

How Netdata helps

Netdata collects CockroachDB metrics at per-second resolution, which matters here: a node can drift tens of milliseconds in a single 15-30 second Prometheus scrape gap. Key capabilities for this failure mode:

  • Per-second clock_offset_meannanos and clock_offset_stddevnanos collection catches drift faster than typical scrape intervals.
  • ML anomaly detection on clock offset trends identifies gradual drift before it crosses static thresholds. A node creeping from 5ms to 50ms over an hour is anomalous even if neither value triggers a threshold.
  • Correlation between clock_offset_meannanos and readwithinuncertainty restarts in a single view confirms clock skew as the root cause of latency degradation without manual cross-referencing.
  • Node liveness changes displayed alongside clock metrics show the full failure path from drift to self-termination.
The Netdata solution

CockroachDB monitoring with Netdata

Netdata monitors CockroachDB with per-second metrics and automatic dashboards. Watch LSM compaction, Raft liveness, clock skew, hot ranges, and intent buildup so the distributed-systems failure modes in these runbooks surface early.