The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / mysql / mysql-fsync-pressure-slow-commits

Operations Guides

MySQL slow commits with idle CPU: redo and binlog fsync pressure

Slow commits or transaction timeouts with idle CPU usually mean the bottleneck is in the durability path, not the query engine.

With innodb_flush_log_at_trx_commit=1 and sync_binlog=1, every COMMIT triggers an InnoDB redo log fsync() and a binary log fsync() before returning to the client. These are sequential, latency-bound operations. When storage cannot drain fsync requests as fast as the server generates them, commits queue up while the CPU waits.

This is not a checkpoint stall. A checkpoint stall happens when the redo log fills and InnoDB must flush dirty pages synchronously to reclaim log capacity. Fsync pressure happens when the log flusher thread cannot complete fsync() calls fast enough to keep up with the commit rate, even though redo log capacity is abundant. Check Innodb_os_log_pending_fsyncs to tell the difference.

What this means

When a client issues COMMIT, InnoDB writes redo entries to the log buffer. A dedicated log flusher thread issues fsync() to make them durable. If binary logging is enabled and sync_binlog >= 1, the server also fsyncs the binary log in the ordered commit pipeline. User threads wait until their log sequence number is flushed.

If storage latency is elevated, the flusher thread falls behind. Pending fsyncs accumulate in OS and device queues. Committing transactions block on the durability barrier. Because the wait is I/O-bound, CPU stays low while query latency spikes.

The critical signals are Innodb_os_log_pending_fsyncs and Innodb_os_log_pending_writes. In a healthy system these are zero. Sustained values greater than zero mean storage cannot keep up with the commit rate.

flowchart TD
    A[Client COMMIT] --> B[Redo log buffer write]
    B --> C[Redo log fsync]
    C --> D[Binlog fsync]
    D --> E[Commit returns]
    C -.->|Slow storage| F[Pending redo fsyncs]
    D -.->|Slow storage| G[Pending binlog fsyncs]
    F --> H[Innodb_os_log_pending_fsyncs > 0]
    G --> I[Commit latency spikes]
    H --> I

Common causes

CauseWhat it looks likeFirst thing to check
Storage latency spikeInnodb_os_log_pending_fsyncs > 0 sustained; OS disk await elevated; no query plan regressionsiostat -x on the redo and binlog devices
Durability settings too strict for storage tierEvery commit triggers two fsyncs; commit rate exceeds storage IOPS ceilinginnodb_flush_log_at_trx_commit and sync_binlog
Redo log and binlog sharing the same deviceBoth sequential fsync streams compete for the same disk queue; latency doubles under loadFilesystem paths for redo log and binlog
Shared storage contentionPeriodic latency spikes that correlate with neighbor activity; no single query culpritHypervisor or cloud volume metrics
Network-attached storage latencyFsync latency measured in tens to hundreds of milliseconds; affects all durable writesNFS or SAN latency metrics

Quick checks

-- Pending redo fsyncs and writes
SHOW GLOBAL STATUS LIKE 'Innodb_os_log_pending%';
-- Aggregate pending fsyncs across all tablespaces
SHOW GLOBAL STATUS LIKE 'Innodb_data_pending_fsyncs';
-- Durability configuration
SHOW GLOBAL VARIABLES WHERE Variable_name IN ('innodb_flush_log_at_trx_commit', 'sync_binlog', 'innodb_flush_method');
-- InnoDB FILE I/O for pending log flushes
SHOW ENGINE INNODB STATUS\G
-- Look for: Pending flushes (fsync) log: N
-- Checkpoint age to rule out capacity stall
SHOW ENGINE INNODB STATUS\G
-- Compare Log sequence number to Last checkpoint at
# OS-level disk latency on the transaction log device
iostat -x 1

How to diagnose it

  1. Confirm the pattern. Check SHOW GLOBAL STATUS for Threads_running and Threads_connected. If Threads_running is well below the CPU core count, the buffer pool hit ratio is healthy, and queries are still slow, suspect the durability path rather than query complexity or CPU saturation.
  2. Check pending fsyncs. Run SHOW GLOBAL STATUS LIKE 'Innodb_os_log_pending%';. Sustained nonzero values for Innodb_os_log_pending_fsyncs or Innodb_os_log_pending_writes indicate the redo log flusher thread is behind and fsyncs are stacking up in the storage queue.
  3. Read InnoDB status. Run SHOW ENGINE INNODB STATUS; and look at the FILE I/O section. Pending flushes (fsync) log: N with N > 0 confirms the storage layer is not draining fsync requests fast enough.
  4. Rule out checkpoint stall. Calculate checkpoint age from the LOG section (Log sequence number minus Last checkpoint at) and compare it to total redo log capacity. In MySQL 8.0.30+, check Innodb_redo_log_capacity_resized. If checkpoint age is well below 75% of capacity but fsyncs are pending, the issue is fsync speed, not log capacity.
  5. Quantify the fsync load. Check innodb_flush_log_at_trx_commit and sync_binlog. If both are 1, every commit generates two fsyncs. Multiply your commit rate by two and compare that to the sustained IOPS and latency profile of your storage. If the product exceeds what the device can deliver at low latency, you have a mathematical mismatch.
  6. Inspect disk latency. Run iostat -x 1 on the devices hosting the redo logs and binary logs. Elevated await with low throughput means the device is struggling with synchronous operations, not bandwidth. Watch aqu-sz (average queue size); values above the device queue-depth limit indicate saturation.
  7. Check for large transactions. Large transactions increase the amount of data flushed per fsync and hold resources in the commit pipeline longer. Review Binlog_cache_disk_use and information_schema.INNODB_TRX for abnormally large writers or long-running transactions.
  8. Correlate with application metrics. If end-to-end commit latency tracks storage fsync latency one-for-one, the diagnosis is confirmed.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Innodb_os_log_pending_fsyncsDirect gauge of redo log flusher thread backlogSustained value > 0
Innodb_os_log_pending_writesPending log writes waiting for I/OSustained value > 0
Innodb_data_pending_fsyncsAggregate fsync backlog across InnoDB data filesSustained value > 0
Checkpoint age / redo log capacityDistinguishes fsync pressure from capacity exhaustionAge < 75% but commits are still slow = fsync speed issue
Innodb_os_log_fsyncs rateBaseline for fsync throughput capacityRate collapses while pending fsyncs rise = storage is stalling
Questions or commit rateApplication-visible throughputSudden drop with steady Threads_connected while CPU is idle

Fixes

Reduce fsync frequency

If the storage tier cannot sustain the current fsync rate, relax durability knobs. Each change trades safety for speed.

  • Set innodb_flush_log_at_trx_commit=2. Redo log entries are written to the OS buffer on every commit, but fsync() happens only once per second. This eliminates the per-commit redo fsync. The tradeoff is up to one second of data loss on an OS crash or power failure. It is safe against a mysqld process crash.
  • Set sync_binlog=0. The binary log fsync is deferred to the operating system. The tradeoff is risk of binlog corruption or data loss on crash. Alternatively, set sync_binlog to a value greater than 1 to fsync the binlog once every N commits instead of every commit.
  • Do not change both settings to zero unless you explicitly accept the risk of losing committed transactions.

Isolate redo and binlog I/O

If the redo log and binary log reside on the same physical device, their sequential fsync streams compete for the same disk queue. Move them to separate volumes backed by distinct physical devices or controllers so that a spike in one does not inflate latency for the other. Update innodb_log_group_home_dir and log_bin accordingly; the change requires a restart.

Upgrade or tune storage

If the storage subsystem is the bottleneck, increase its synchronous IOPS capacity. Options include moving to local NVMe, increasing provisioned IOPS on cloud volumes, resolving RAID rebuilds, or moving off oversubscribed shared storage. If the storage has a write-back cache, ensure it is battery-backed or non-volatile before relying on it for durability.

Prevention

  • Monitor Innodb_os_log_pending_fsyncs continuously. Any sustained nonzero value means storage headroom is gone and commit latency will follow.
  • Size storage for peak commit rate multiplied by the fsync multiplier imposed by your durability settings. With innodb_flush_log_at_trx_commit=1 and sync_binlog=1, every commit generates two fsyncs.
  • Keep transaction logs and binlogs on separate volumes when infrastructure allows.
  • Avoid colocating high-fsync workloads such as backups, schema migrations, or batch imports on the same storage controller as the transaction logs.
  • Track transaction latency in performance_schema so that fsync pressure is visible before application timeouts trigger.

How Netdata helps

  • Netdata tracks Innodb_os_log_pending_fsyncs and Innodb_os_log_pending_writes without manual polling of SHOW GLOBAL STATUS.
  • It plots MySQL commit latency against OS disk await and utilization on the same timeline, confirming whether storage is the bottleneck.
  • Checkpoint age and pending fsyncs are visualized together, distinguishing a capacity stall from a write-speed stall in one view.
  • Alerts on sustained nonzero pending fsyncs provide early warning before commit latency degrades into application timeouts.
The Netdata solution

MySQL monitoring with Netdata

Netdata monitors MySQL and MariaDB with per-second metrics and ML anomaly detection. Track connection usage, query throughput, slow queries, redo-log pressure, and replication lag alongside the host and storage signals that explain them.