The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / ceph / ceph-mon-clock-skew

Operations Guides

Ceph MON_CLOCK_SKEW: clock drift between monitors and election churn

The MON_CLOCK_SKEW health check fires when the leader monitor detects clock drift beyond mon_clock_drift_allowed (default 0.05 seconds, 50 milliseconds) on any monitor in the quorum. It raises HEALTH_WARN, not HEALTH_ERR, and on its own it does not stop client I/O. But it is the precursor to a class of failure that does: repeated Paxos elections, slow map distribution, and, if the skew grows, monitor quorum loss.

A single transient warning is usually NTP convergence after boot, after a hypervisor live migration, or under brief network pressure. Sustained or recurring skew is almost always an NTP/chrony problem on the MON host: a stopped daemon, an unreachable source, an undersized polling interval, or a virtualized monitor on an overloaded hypervisor. Paxos uses wall-clock time for lease management and internal ordering, so even sub-second drift degrades election stability.

What this means

Monitors maintain cluster maps (OSD map, MON map, PG map, MDS map, CRUSH map) via Paxos consensus. A majority quorum must agree on every map update. The leader periodically checks the clocks of its peers. Under normal conditions the check runs on mon_timecheck_interval (default 300 seconds). Once skew is detected, it tightens to mon_timecheck_skew_interval (default 30 seconds).

When the leader sees any peer’s clock beyond mon_clock_drift_allowed, it raises MON_CLOCK_SKEW. The health detail output names the offender and shows the magnitude and network latency to it:

mon.<id> addr <ip>:6789/0 clock skew <value>s > max 0.05s (latency <value>s)

The 50ms threshold must stay well below the monitor lease interval. Upstream Ceph documentation is explicit: do not raise mon_clock_drift_allowed without testing, and even then prefer fixing the underlying time source. The documentation states plainly that you should run NTP on bare metal, because VM-virtualized clocks are not suitable for steady timekeeping in monitor workloads.

The downstream symptom is election churn. Paxos leadership and lease renewals are time-bound. When clocks disagree, leases appear to expire early or late from the perspective of different monitors. The result is monitors calling for elections that should not be needed, freezing map updates for the duration of each round.

flowchart TD
    A[NTP/chrony stopped or source unreachable] --> B[MON clock drifts]
    C[Hypervisor overloaded or live migration] --> B
    B --> D[Leader detects skew over 50ms]
    D --> E[MON_CLOCK_SKEW fires, HEALTH_WARN]
    E --> F[Paxos leases misfire on skewed MON]
    F --> G[Election epoch increments]
    G --> H[Map distribution stalls, CLI commands slow]
    H --> I[Possible quorum loss if skew grows]

Common causes

CauseWhat it looks likeFirst thing to check
NTP/chrony daemon stoppedsystemctl status chronyd inactive; chronyc tracking reports no sourceRestart daemon and inspect chronyc sources
NTP source unreachableDaemon running, chronyc tracking shows Leap status: Not synchronisedNetwork egress, firewall, DNS to upstream
VM clock driftSkew only on virtualized MONs; worse after live migration or under host CPU pressureMove MONs to bare metal, or fix hypervisor timekeeping
Hypervisor live migrationSkew spikes at the moment of vMotion/migration, often clears afterAvoid migrating MON VMs, pin them
systemd-timesyncd onlySkew persists even though timedatectl reports NTP=activeReplace with chrony; timesyncd’s discipline is too loose for MONs
Hardware clock drift (CMOS battery)Single bare-metal host drifts consistently over hours or daysInspect chronyc tracking offset trend, replace battery
Overloaded MON hostCo-located services steal CPU; daemon running but skew oscillatestop, iostat; ensure MON host is dedicated

Quick checks

These are read-only and safe to run on any cluster with read credentials.

# Cluster-level clock-skew status from the monitor leader
ceph time-sync-status

# Cluster health detail, filtered to clock-related lines
ceph health detail | grep -i clock

# Election epoch and quorum leader
ceph quorum_status -f json | jq '{epoch: .election_epoch, leader: .quorum_leader_name, quorum: .quorum_names}'

# Monitor stats
ceph mon stat

# Per-MON Paxos and election counters (run on the MON host)
ceph daemon mon.$(hostname -s) perf dump | jq '{paxos: .paxos, mon: .mon}'

# Monitor store size (large stores slow elections)
du -sh /var/lib/ceph/mon/ceph-$(hostname -s)/store.db

# Time daemon status on the MON host
chronyc tracking
chronyc sources -v

# Generic timedatectl status (covers timesyncd too)
timedatectl status
timedatectl timesync-status

How to diagnose it

  1. Confirm the warning is current and identify the offender. ceph health detail names the skewed monitor and shows the magnitude. If only one monitor is named, the other two are the reference. If multiple are named, the cluster has no common time base and the underlying NTP failure is broader.

  2. Establish whether this is transient or structural. The first occurrence after boot, after a live migration, or after a network blip may clear within minutes as NTP reconverges. Use ceph_health_detail{name="MON_CLOCK_SKEW"} sustained for more than 60 seconds as the operational signal; anything shorter is likely convergence noise.

  3. Check the time daemon on every MON host, not just the leader. The leader reports the skew; the offender is usually elsewhere. On each monitor, run chronyc tracking and look at System time and Last offset. If the daemon reports Not synchronised, the source is unreachable or the daemon is misconfigured.

  4. Correlate skew with election churn. Pull Paxos counters from each monitor with ceph daemon mon.<id> perf dump. The relevant fields are under paxos (commit latency, accept latency) and mon (election_call, election_win, election_lose). Election counts incrementing more than once per hour while skew is active confirm the cascade.

  5. Distinguish host-level from cluster-level. If only one MON host drifts, the problem is local (daemon, hardware clock, hypervisor pressure). If all MONs drift together, the NTP source itself is wrong or unreachable everywhere (broken upstream, restrictive firewall, shared dependency).

  6. Rule out the monitor store. A large or uncompacted RocksDB store slows Paxos rounds and can mimic election instability. Check du -sh /var/lib/ceph/mon/ceph-*/store.db on every MON. Healthy clusters keep this well under 5GB. Run ceph daemon mon.<id> compact during a maintenance window if a single store is bloated.

  7. In Rook or containerized deployments, check the host, not the pod. Monitor containers inherit the host clock. Running chrony inside the MON pod is a common but ineffective workaround; the host kernel clock is what matters. The fix lives on the Kubernetes node, not in the MON container.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ceph_health_detail{name="MON_CLOCK_SKEW"}Canonical Ceph signal for this conditionActive for more than 60 seconds
ceph_mon_quorum_status per MONConfirms whether skew has become quorum lossSum in quorum below majority threshold
Election epoch rate from ceph quorum_statusElection churn is the user-visible damageEpoch incrementing more than once per hour
chronyc tracking system offset on MON hostsThe actual drift value the cluster is reacting toSystem offset > 50ms or Not synchronised
chronyc sources reachabilityDetermines whether the upstream is healthyReach counter at 0 or never-increasing
MON host CPU steal (virtualized)Identifies hypervisor-driven driftHigh steal coinciding with skew events
MON store sizeLarge stores amplify election latencyStore > 5GB or growing faster than 100MB/day

Fixes

Restart and reconfigure the time daemon

If chronyd is stopped or wedged, restarting it is the first action. Verify the source list is sensible before restarting so the daemon converges on a real upstream:

# Inspect current sources
chronyc sources -v

# Restart chrony
sudo systemctl restart chronyd
sudo systemctl enable chronyd

# Watch convergence
watch -n 2 chronyc tracking

If the source list is empty or only points to unreachable hosts, add reachable servers and restart. On AWS use 169.254.169.123; on GCP use metadata.google.internal (the metadata server also serves NTP on GCP instances). For bare metal, prefer multiple low-stratum sources from the NTP Pool or your network’s stratum-2 servers.

Replace systemd-timesyncd with chrony

systemd-timesyncd is fine for general hosts but its clock discipline is too loose for monitors. Community reports from Proxmox, ServerFault, and Rook issues consistently show that switching to chrony resolves persistent MON_CLOCK_SKEW that timesyncd could not. Disable timesyncd before enabling chronyd to avoid two daemons fighting:

sudo systemctl disable --now systemd-timesyncd
sudo systemctl enable --now chronyd

Move virtualized MONs to bare metal

The Ceph documentation is explicit: run NTP on bare metal; VM-virtualized clocks are not suitable for steady timekeeping. If your monitors are VMs and skew recurs, the durable fix is to deploy at least the leader and one peer on bare metal. If that is not possible, reduce the surface for drift:

  • Pin MON VMs to dedicated cores to reduce scheduling jitter.
  • Disable the balloon driver and memory over-commit on MON VMs.
  • Avoid live migration of MON VMs. If migration is required, schedule it during a maintenance window and expect brief skew warnings.
  • Install the hypervisor’s guest integration (VMware Tools, qemu-guest-agent) and any host timekeeping integration it exposes.

Fix SELinux denials

On SELinux-enforced hosts, denials can prevent chronyd from synchronizing. The supported fix is to configure the correct policy, not to disable SELinux. Inspect denials and apply the recommended policy fix:

# Inspect denials
sudo ausearch -m AVC -ts recent | grep chrony
sudo sealert -a /var/log/audit/audit.log

Do not raise mon_clock_drift_allowed

The 0.05s value is tight because Paxos leases are time-bound; raising it lets drift get worse before Ceph warns, which moves the failure closer to quorum loss rather than further from it. Red Hat and upstream Ceph guidance both warn against changing this value without testing, and testing in this context means deliberately inducing skew on a non-production cluster and observing election behavior.

If, after fixing timekeeping, residual warnings persist from a single host with known small drift (for example, false positives at 0.011s during NTP convergence), the right action is to wait for convergence and watch the trend, not to widen the threshold.

Reduce election churn while you fix the cause

If skew is producing election storms and CLI commands are slow, you can stop the worst monitor temporarily to restore stable quorum while you fix the underlying clock:

# Stop the offending MON daemon on its host
sudo systemctl stop ceph-mon@<id>

Removing one monitor from a 3-MON cluster leaves a 2-of-3 majority, so quorum holds. This is a temporary measure, not a fix; a 3-MON cluster has no fault tolerance with one monitor down. Restart the monitor as soon as its clock is stable.

Prevention

  • Run chrony on every MON host. Prefer chrony over ntpd and over systemd-timesyncd. Configure at least three independent upstream sources and monitor chronyc tracking offset as a time-series metric.
  • Run MONs on bare metal when possible. If you must virtualize, pin the VM, disable balloon and over-commit, and never live-migrate MON VMs.
  • Monitor skew as a time series, not just a health check. Track ceph_health_detail{name="MON_CLOCK_SKEW"} with a 60-second sustain for tickets. Treat any individual monitor losing quorum as a separate, more urgent signal.
  • Monitor election epoch rate. A flat or slowly growing epoch is healthy. Spikes coincide with user-visible slowness in ceph -s, ceph status, and map distribution.
  • Track MON store size. Large stores amplify election latency and make the cluster more sensitive to skew. Alert if any MON store exceeds 5GB or grows faster than 100MB/day.
  • Track chrony offset on MON hosts directly. Netdata’s chrony collector exposes the system offset and source reachability as per-second metrics. This catches drift before it crosses the 50ms threshold that triggers Ceph.
  • In Rook/Ceph, configure time on the Kubernetes nodes, not in the MON pods. Verify with chronyc tracking on each node that runs a MON pod.

How Netdata helps

  • Per-second ceph_health_detail{name="MON_CLOCK_SKEW"} distinguishes transient post-boot convergence (clears in seconds) from structural drift (sustained for minutes). The 60-second sustain aligns with the operational threshold.
  • Correlate MON_CLOCK_SKEW with ceph_mon_quorum_status to confirm whether skew has graduated to quorum loss. Two signals, one chart, immediate context.
  • Netdata’s chrony collector exposes system offset, last offset, and source reachability per-second on every MON host, so you see drift building before it crosses 50ms and triggers Ceph.
  • ML anomaly detection on the election epoch counter surfaces election churn that correlates with skew events, even when no static threshold would have fired.
  • Host-level CPU steal and scheduling metrics on MON hosts identify hypervisor-driven drift in virtualized deployments, distinguishing a Ceph problem from a VM placement problem.
  • Filesystem metrics on the MON data directory flag bloated stores that amplify election sensitivity to skew.