The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / mysql / mysql-seconds-behind-source-unreliable

Operations Guides

MySQL Seconds_Behind_Source is unreliable: measuring real replication lag

You run SHOW REPLICA STATUS and see Seconds_Behind_Source = 0. Both replication threads say Yes. You assume the replica is current. During a source write burst, you fail over and discover hours of missing transactions. This happens because Seconds_Behind_Source is not a real-time lag measurement. It is a timestamp diff between the replica’s clock and the timestamp of the event the SQL thread is currently applying. When the SQL thread reaches the end of the relay log, the metric snaps to zero even if the I/O thread has not fetched newer events. When the SQL thread stops, it returns NULL, which most lag alerts ignore. When the source was idle and then bursts, it jumps to the idle duration even though real lag is minimal. With parallel replication, it oscillates based on worker timing. For failover-critical decisions, you need a different measurement strategy.

What it is and why it matters

Seconds_Behind_Source (or Seconds_Behind_Master in MySQL 5.7) is the default replication lag metric in SHOW REPLICA STATUS. It is the first value most operators check when they suspect a replica is stale. Because it is built in and requires no extra setup, it often becomes the only lag signal in monitoring systems and failover scripts.

This is dangerous. Failover decisions, read-after-write consistency checks, and RPO estimates all depend on knowing exactly how far behind a replica is. If the metric reports 0 when the replica is actually hours behind, or NULL when replication is broken, your automation makes wrong decisions silently. Knowing when this metric is lying is a prerequisite for safe replicated topologies.

How it works

MySQL estimates Seconds_Behind_Source as the difference between the replica’s current time and the timestamp of the event the SQL thread is currently applying. It does not compare the replica against the source’s current time. It does not measure network delay or I/O thread fetch delay directly. It is purely a function of the event timestamp inside the relay log.

Because the calculation is bounded by the relay log, the SQL thread can be fully caught up with every event it has received and still report 0, even when the I/O thread has not pulled new events from the source for hours. The metric tells you whether the SQL thread is busy, not whether the replica is current.

flowchart TD
  S[Source event
timestamp T1] --> IO[Replica I/O thread] IO --> RL[Relay log] RL --> SQL[Replica SQL thread] SQL --> SBS[Seconds_Behind_Source
now - T1] IO -->|Stalled| E1[Relay log stale] E1 --> F1[SBS = 0
real lag hidden] S -->|Idle then burst| E2[Old timestamp] E2 --> F2[SBS = 3600
real lag seconds] SQL -->|Stopped| F3[SBS = NULL
alert silent]

Where it shows up in production

The metric fails in four specific ways that have each caused production incidents.

False zero when the I/O thread lags

When the SQL thread drains the relay log and goes idle, Seconds_Behind_Source snaps to 0. If the I/O thread is stalled, reconnecting after a network blip, or slower than the source’s write rate, the relay log itself falls behind. The operator sees a flat zero during the exact moment lag is accumulating. This is common during heavy write bursts, cross-region replication, or when the source’s binlog dump thread is overloaded.

NULL instead of infinity when replication breaks

If the SQL thread stops due to a duplicate key error, schema mismatch, or an unsupported statement, Seconds_Behind_Source returns NULL. An alert threshold like Seconds_Behind_Source > 30 never fires, because NULL is not greater than 30. Replication is broken and lag is growing unboundedly, but the metric appears blank rather than alarming. This is one of the most dangerous silent failure modes in MySQL replication monitoring.

Spikes after source idle periods

If the source was idle for an hour and then receives a burst of writes, the next event in the relay log carries a timestamp from an hour ago. Seconds_Behind_Source jumps to 3600 even though the replica applies the burst within seconds of arrival. This produces false-positive lag alerts and pages that do not represent real data staleness. The metric answers “how old is this event?” not “how long did it take to arrive?”.

Inaccuracy with parallel replication

With replica_parallel_workers > 0, the value oscillates as workers apply transactions with different original commit timestamps. It does not aggregate backlog across workers, so it cannot represent true apply lag in multi-threaded configurations. A replica can have a deep apply queue and report a low number because one fast worker happened to commit recently.

Tradeoffs and when to use it

Seconds_Behind_Source is not worthless. On lightly loaded, single-threaded topologies where the source is continuously active and both replication threads are confirmed healthy, it offers a rough trend of apply performance. It is useful for answering “is the SQL thread keeping up with the relay log?”.

It is dangerous for automated failover, read-after-write routing, SLA enforcement, and any scenario where the I/O thread could lag independently of the SQL thread. If your orchestrator uses SBS as the sole lag input, it will miss I/O stalls, treat broken replication as healthy, and panic over idle-source artifacts.

How to measure real replication lag

For production decisions, use one of the following methods.

Heartbeat tables with pt-heartbeat

A heartbeat table on the source is updated with the current timestamp every second. The replica reads this row and computes the wall-clock delta. This measures end-to-end latency, including network delay, I/O thread fetch time, and SQL thread apply time.

Percona Toolkit’s pt-heartbeat is the standard implementation. It creates a single-row table, updates it on the source, and provides a --monitor mode on the replica that reports real lag in seconds. Many failover tools prefer heartbeat lag over Seconds_Behind_Source because it does not suffer from the false zero and idle-spike problems.

Run the updater on the source and the monitor on the replica. The output is a simple integer: seconds of lag. Feed this into your orchestrator if you run automated failover.

GTID set comparison

If GTID replication is enabled, compare the source’s executed set against the replica’s.

On the source:

SELECT @@global.gtid_executed;

On the replica:

SELECT @@global.gtid_executed;

Then compare:

SELECT GTID_SUBTRACT('source_gtid_set', 'replica_gtid_set');

A non-empty result shows the exact transactions the replica has not yet executed. This gives precise transaction-level lag. It is the safest way to determine whether a replica is caught up enough to fail over to, because it tells you exactly which transactions are missing.

The tradeoff is that it returns a transaction gap, not a time duration. It answers “are we missing data?” perfectly, but it does not answer “how many seconds behind are we?” without correlating transaction volume with time. For RPO decisions, however, the transaction gap is usually what you actually care about.

Performance Schema worker timestamps

For multi-threaded replication diagnostics, MySQL 8.0 exposes per-worker apply timestamps in performance_schema.replication_applier_status_by_worker. Comparing LAST_APPLIED_TRANSACTION_ORIGINAL_COMMIT_TIMESTAMP against the current time gives the exact apply lag for each worker individually.

This is useful for identifying which specific worker is slow, but it covers only the apply side. It does not measure I/O thread fetch delay.

Signals to watch in production

SignalWhy it mattersWarning sign
Seconds_Behind_SourceRough SQL thread catch-up indicator onlySudden drop to 0 during source write bursts, or NULL when SQL thread stops
Replication thread stateBinary health of I/O and SQL threadsReplica_IO_Running = No/Connecting, or Replica_SQL_Running = No
Relay log spaceGrowing relay log means I/O thread is outpacing SQL threadContinuous growth correlating with write peaks
Heartbeat lagReal end-to-end latencySustained value above RPO threshold
GTID set differenceExact transaction gap for failover safetyGTID_SUBTRACT result growing monotonically
Per-worker apply lagMicrosecond precision for MTS bottlenecksOne worker lagging far behind others

How Netdata helps

  • Netdata collects Seconds_Behind_Source alongside Replica_IO_Running and Replica_SQL_Running, enabling composite alerts that require healthy thread states before trusting lag values. This prevents the silent failure mode where a NULL from a stopped SQL thread is missed.
  • Netdata charts relay log growth and source write rate on the same dashboard as lag metrics, letting you distinguish an I/O thread stall (relay log flat, SBS zero) from genuine apply catch-up.
  • For operators running pt-heartbeat, Netdata can chart custom SQL metrics, letting you plot heartbeat lag directly against SBS to visualize the divergence during idle-source bursts.
  • Netdata alerts when either thread leaves the Yes state, independent of lag calculations.
The Netdata solution

MySQL monitoring with Netdata

Netdata monitors MySQL and MariaDB with per-second metrics and ML anomaly detection. Track connection usage, query throughput, slow queries, redo-log pressure, and replication lag alongside the host and storage signals that explain them.