The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / mysql / mysql-online-ddl-blocking

Operations Guides

MySQL online DDL still blocking: ALGORITHM, LOCK, and the copy phase

Online DDL in MySQL is not lock-free. Even with ALGORITHM=INPLACE or ALGORITHM=INSTANT, and even with LOCK=NONE, every online DDL operation passes through a brief window where it upgrades its metadata lock (MDL) to exclusive. That window is short, but it is real, and it is the single most common reason an “online” schema change still stalls a busy production table.

This article explains the three-phase MDL model that online DDL uses, why ALGORITHM=INPLACE with LOCK=NONE can still block, what the copy phase actually does, and how third-party tools such as pt-online-schema-change and gh-ost change the picture. It assumes you already understand the broader MySQL mental model. See the hub page for the failure-pattern catalogue that frames metadata lock stalls as one of MySQL’s characteristic outage archetypes.

What online DDL actually promises

“Online” in MySQL terminology means concurrent DML (INSERT, UPDATE, DELETE) is permitted during the execution phase of the DDL. It does not mean the DDL acquires no exclusive locks. The promise is narrower than most operators assume:

  • ALGORITHM=INPLACE with LOCK=NONE allows concurrent reads and writes during the bulk of the operation.
  • ALGORITHM=INSTANT modifies only data dictionary metadata and touches no table data.
  • Both still require metadata locks, and both still upgrade to an exclusive MDL at commit time.

If your expectation is “the DDL will never block another session,” that expectation is wrong. The correct expectation is “the DDL blocks for a very short time at the start and at the end, and during that window anything holding a shared MDL on the table will block the DDL, which in turn blocks everything queuing behind the DDL.”

For the full background on how metadata locks cascade into partial outages, see the hub page section on the metadata lock cascade pattern.

The three-phase MDL model

The MySQL 8.4 Reference Manual documents online DDL as three phases. Understanding the phases is the key to understanding every “online DDL still blocking” symptom.

flowchart TD
  A[Phase 1: Initialization] -->|shared upgradeable MDL| B[Phase 2: Execution]
  B -->|concurrent DML permitted for INPLACE/INSTANT| C[Phase 3: Commit table definition]
  C -->|exclusive MDL, brief| D[New table definition visible]
  D -->|all new DML on table resumes| E[Operation complete]
  C -->|any session holding shared MDL blocks commit| F[DDL waits, new DML queues behind DDL]

Phase 1: initialization

The DDL takes a shared, upgradeable metadata lock on the table. This lock is compatible with concurrent DML, which is why the operation can proceed while queries run. The server assesses whether the operation can be done INPLACE or must fall back to COPY.

Phase 2: execution

For ALGORITHM=INPLACE with LOCK=NONE, concurrent DML is permitted. InnoDB may build a new version of the table (or index) in the background while the original table continues to serve reads and writes. For ALGORITHM=INSTANT, this phase does effectively nothing to table data because only data dictionary metadata changes.

However, the manual notes that an exclusive metadata lock may be taken briefly during the execution phase, during statement preparation. Whether the upgrade happens depends on factors assessed in the initialization phase. If it is required, it is brief, but it is still exclusive.

Phase 3: commit table definition

This is where the “still blocking” symptom originates. The manual states that in the commit table definition phase, the metadata lock is upgraded to exclusive. The server must evict the old table definition from every session’s table cache and publish the new one. While the exclusive MDL is held:

  • New DML against the table queues behind the exclusive request.
  • Any session that still holds a shared MDL on the table prevents the DDL from acquiring the exclusive MDL.
  • The DDL therefore waits for that shared MDL to release, and every new DML statement queues behind the DDL.

This is the metadata lock cascade, triggered by a DDL that is otherwise fully “online” during execution.

Why ALGORITHM=INPLACE with LOCK=NONE still blocks

LOCK=NONE governs what level of table access the DDL requires during execution. It does not remove the exclusive MDL requirement at commit. The two clauses operate at different layers:

  • LOCK=NONE controls InnoDB-level table access during the bulk of the operation.
  • Metadata locks are a server-level layer, separate from InnoDB row locks, and they apply regardless of the LOCK clause.

Bug #106480 (filed February 2022, last publicly modified May 2022) traced this to wait_while_table_is_used() in sql/sql_base.cc, which unconditionally upgrades to an exclusive MDL before commit regardless of the locking mode requested. The MySQL 8.4 manual text now acknowledges that an exclusive MDL may be taken briefly during execution and is taken at commit. The bug remains open in the public tracker as of the last available record.

The practical consequence is that the presence of any long-running transaction holding a shared MDL on the table will block the DDL’s commit, and every new DML statement on the table will queue behind the DDL. The DDL is not stuck because of row locks or because the copy is slow. It is stuck because the server cannot get exclusive control of the table definition to swap in the new one.

Why ALGORITHM=INSTANT still blocks

ALGORITHM=INSTANT modifies only data dictionary metadata. No table data is copied, no rows are rebuilt, and concurrent DML is permitted during execution. Despite this, INSTANT still requires the exclusive MDL at commit. The same wait_while_table_is_used() code path applies.

INSTANT in MySQL 8.4 supports operations such as adding or dropping a column (at any position since 8.0.29), renaming a column, setting or dropping a column default, adding or dropping a virtual generated column, and changing index type. INSTANT is the default algorithm in MySQL 8.4. The server tries INSTANT first, then INPLACE, then COPY.

Always specify ALGORITHM explicitly. If you request ALGORITHM=INSTANT and the operation cannot be done INSTANT, MySQL returns an error rather than silently degrading to INPLACE or COPY. That error is cheap insurance against an unexpected full table rebuild.

One subtlety: ALGORITHM=INSTANT combined with LOCK=NONE, LOCK=SHARED, or LOCK=EXCLUSIVE produces an error. Only LOCK=DEFAULT is compatible with ALGORITHM=INSTANT.

The copy phase

ALGORITHM=COPY is the legacy path. It creates a temporary table, copies every row from the original, applies the schema change to the copy, and then swaps the new table in for the old. The copy phase blocks concurrent DML and holds an exclusive MDL for the entire operation, plus the brief commit-phase upgrade.

You can confirm a COPY operation by checking the rows-affected count. A nonzero rows-affected count after an ALTER indicates data was copied. INPLACE operations that do not rebuild the table report zero rows affected.

INPLACE rebuild operations

Some INPLACE operations rebuild the table internally even though they are classified as INPLACE rather than COPY. Examples include adding or dropping a primary key, changing a column type, and reordering columns. These operations consume disk space roughly equal to the original table size in the InnoDB temporary directory, controlled by innodb_tmpdir. They still permit concurrent DML when combined with LOCK=NONE, but the rebuild itself is more expensive than a metadata-only change.

The online alter log and the commit window

For ALGORITHM=INPLACE with LOCK=NONE, InnoDB buffers concurrent DML changes during execution in the online alter log, sized by innodb_online_alter_log_max_size (default 128 MB). These buffered changes are replayed at the end of the operation, during the exclusive MDL window at commit. The larger the log, the more replay work happens under the exclusive MDL, and the longer the commit window.

If the online alter log overflows during execution, the DDL fails and rolls back any uncommitted concurrent changes. Monitoring innodb_online_alter_log_max_size relative to write rate during a long INPLACE operation is important on write-heavy tables.

The INSTANT row-version limit

Every INSTANT ADD COLUMN or DROP COLUMN increments a per-table counter, visible as TOTAL_ROW_VERSIONS in INFORMATION_SCHEMA.INNODB_TABLES. In MySQL 8.0 through 8.4, the limit is 64. MySQL 9.1.0 raised the limit to 255.

When the limit is reached, MySQL silently degrades to INPLACE or COPY unless ALGORITHM is explicitly specified. If you always specify ALGORITHM=INSTANT, the server returns an error instead of silently downgrading, which is the safer operational behavior.

Monitor row versions with:

SELECT NAME, TOTAL_ROW_VERSIONS
FROM INFORMATION_SCHEMA.INNODB_TABLES
WHERE TOTAL_ROW_VERSIONS > 0
ORDER BY TOTAL_ROW_VERSIONS DESC;

Reset the counter with OPTIMIZE TABLE or ALTER TABLE ... ENGINE=InnoDB. Both rebuild the table and reset TOTAL_ROW_VERSIONS to zero.

pt-online-schema-change and gh-ost

Third-party schema change tools work around the copy phase by doing the copy outside of MySQL’s online DDL machinery, but they have their own MDL behavior.

pt-online-schema-change

pt-online-schema-change (pt-osc) creates a shadow table, copies rows in chunks, and keeps the shadow in sync with the original using INSERT, UPDATE, and DELETE triggers on the original table. Because of the trigger design, pt-osc cannot run on tables that already have triggers, and it requires exclusive access to the original table during trigger creation.

pt-osc holds metadata locks during the copy phase for operations such as trigger creation, the final table swap, and foreign key updates. Its lock_wait_timeout defaults to 60 seconds for metadata lock operations. If any session holds a long transaction on the table, pt-osc blocks at the same MDL layer as native online DDL.

gh-ost

gh-ost also holds MDL during the copy phase but uses row-based binlog streaming instead of triggers. This allows it to run on tables that already have triggers, which pt-osc cannot do. The trigger-free design changes some of the operational constraints but does not eliminate metadata lock requirements at the swap phase.

Both tools are useful when the operation cannot be done INPLACE or INSTANT, or when you need finer control over chunking, throttling, and pause-resume than native online DDL provides. Neither tool is lock-free.

Where this shows up in production

The most common production symptom is an ALTER TABLE that appears to hang, followed within minutes by a flood of application timeouts against the affected table. The pattern is:

  1. A DDL is issued during a window when a long transaction is open on the same table. The long transaction may be a monitoring query, a reporting query, an ORM that began a transaction without committing, or a mysqldump --single-transaction backup.
  2. The DDL reaches commit and requests the exclusive MDL.
  3. The DDL waits for the long transaction to release its shared MDL.
  4. New DML statements on the table queue behind the DDL’s exclusive request.
  5. The connection pool fills with waiting sessions. Other tables continue to work, producing a confusing partial outage.

The DDL is not doing work during the wait. It is blocked at the MDL layer. CPU and disk I/O may be completely normal, which makes the incident harder to diagnose if you are only watching resource metrics.

Tradeoffs and when to use what

GoalRecommended approachTradeoff
Add or drop a column, rename a column, change index typeALGORITHM=INSTANTStill needs exclusive MDL at commit, but commit is fast
Add secondary index, change column type, rebuild primary keyALGORITHM=INPLACE, LOCK=NONEConcurrent DML allowed, but watch online alter log size and commit replay
Operation not supported by INPLACEALGORITHM=COPY or pt-osc / gh-ostCOPY blocks DML for the whole operation; pt-osc / gh-ost avoid that but add tooling complexity
Force failure instead of silent downgradeAlways specify ALGORITHM explicitlyOperation errors out if the requested algorithm is not supported

Run DDL when no long transactions are open on the target table. This is the single most effective way to avoid the commit-phase MDL stall.

Signals to watch in production

SignalWhy it mattersWarning sign
performance_schema.metadata_locks rows with LOCK_STATUS = 'PENDING'Direct evidence of MDL queue buildingMore than 3 sessions pending on the same object, sustained
SHOW PROCESSLIST state “Waiting for table metadata lock”Sessions blocked behind a DDL or behind each otherCount rising on a single table
INNODB_TRX oldest transaction ageLong transactions hold shared MDLsAny OLTP transaction older than 60 seconds
information_schema.INNODB_TABLES.TOTAL_ROW_VERSIONSApproaching the 64-row-version ceiling before silent INPLACE/COPY fallbackCounter approaching 50 for frequently altered tables
Innodb_buffer_pool_pages_dirty ratio during INPLACE rebuildRebuild dirties many pages, can pressure checkpointSustained dirty ratio above 75%
innodb_online_alter_log_max_size relative to write rateOnline alter log overflow aborts the DDLLog filling during long INPLACE on write-heavy table

How Netdata helps

  • The MySQL collector exposes Threads_running, Threads_connected, and Questions rate, which together reveal the moment a DDL starts to stall query processing.
  • Netdata’s correlation across the metadata lock cascade makes it easier to see the relationship between a rising MDL wait count and a corresponding drop in throughput on the affected table.
  • Monitoring long-lived transactions via INNODB_TRX age gives early warning before a DDL is issued against a table with an open transaction.
  • History list length tracking (trx_rseg_history_len) catches the related failure mode where an idle transaction is also blocking purge.
  • Per-table query digest latency from performance_schema.events_statements_summary_by_digest shows whether a stall is isolated to one table or system-wide.
  • Alerting on performance_schema.metadata_locks pending counts detects the cascade before it reaches connection exhaustion.
The Netdata solution

MySQL monitoring with Netdata

Netdata monitors MySQL and MariaDB with per-second metrics and ML anomaly detection. Track connection usage, query throughput, slow queries, redo-log pressure, and replication lag alongside the host and storage signals that explain them.