The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-writelog-waits-high

Operations Guides

SQL Server WRITELOG waits: commit latency from a slow transaction log

WRITELOG is the wait type SQL Server records when a worker thread blocks on a transaction log flush. Write-ahead logging requires the log block to be hardened to disk before the engine acknowledges a commit, so WRITELOG directly bounds write throughput. When it dominates your top-waits list, every write transaction is paying a latency tax at commit.

The reported symptom is rarely “WRITELOG is high.” It is commit latency, write transaction timeouts, application retry storms, Availability Group replication lag, or batch requests/sec collapsing while CPU sits low. WRITELOG is the in-engine signature you find in sys.dm_os_wait_stats or sys.dm_exec_session_wait_stats. It tells you the bottleneck is in the log write path: the storage below the log file, the I/O stack between SQL Server and that storage, or the commit rate the Log Writer is being asked to service. Tuning queries will not fix it.

What this means

WRITELOG is recorded on the worker thread that issued the commit (or any operation that forces a log flush). A log flush is triggered when:

  • A transaction commits and must be hardened before the engine acknowledges it.
  • A log block fills to its maximum size of 60KB.
  • Write-ahead logging forces a flush before a dirty data page can be written.
  • sp_flush_log is executed (used by natively compiled procedures).

The flush is issued by the Log Writer. Pre-2016 there was a single Log Writer thread. SQL Server 2016 added up to 4 Log Writer threads. SQL Server 2019 raised that to 8 and also allows regular worker threads to issue log writes directly. When the storage below the log file is slow, every commit pays that latency.

The coupling between commit and flush is what makes WRITELOG damaging in OLTP. In a high-rate insert/update workload with thousands of small transactions per second, even 5ms of log write latency becomes a serialization point. There is no batching across transactions unless the application batches. The log is a serial stream within each database.

flowchart TD
    A[App commits txn] --> B[Engine builds log block]
    B --> C[Log Writer flushes block to disk]
    C --> D{I/O completes?}
    D -- yes --> E[Commit ack sent to client]
    D -- slow / no --> F[Worker records WRITELOG]
    F --> C
    C -.->|single serial path| G[Every other commit in this DB]

Every concurrent write transaction in the same database funnels through this single path.

Common causes

CauseWhat it looks likeFirst thing to check
Slow log-file storageavg_write_latency_ms on the LOG file elevated; WRITELOG scales with write ratesys.dm_io_virtual_file_stats joined to sys.master_files, filtered to type_desc = 'LOG'
Filter driver in the I/O pathPerfMon disk counters look healthy but WRITELOG still dominatesfltmc instances from an elevated cmd prompt
Outstanding-I/O cap on fast SSDLog flush write time is sub-millisecond but wait time is in thousands of ms; ~112 outstanding requests per databasesys.dm_io_pending_io_requests and per-database write throughput
Tiny transactions committing individuallyHigh transactions/sec, low log bytes per commit, WRITELOG scales with transaction count not log volumeLog Flushes/sec vs Transactions/sec, or query stats showing single-row writes in tight loops
Synchronous-commit AGHADR_SYNC_COMMIT even larger than WRITELOG on primary; secondary log harden is the capHADR_SYNC_COMMIT wait delta and secondary log-file I/O stall

Quick checks

All read-only. Run them in order.

-- 1. Top waits since last reset. For a live problem, snapshot twice 60s apart and compute deltas.
SELECT TOP 20
    wait_type,
    waiting_tasks_count,
    wait_time_ms,
    signal_wait_time_ms,
    wait_time_ms - signal_wait_time_ms AS resource_wait_time_ms,
    CAST(100.0 * wait_time_ms / SUM(wait_time_ms) OVER() AS DECIMAL(5,2)) AS pct
FROM sys.dm_os_wait_stats
WHERE wait_type NOT IN (
    'SLEEP_TASK', 'BROKER_TO_FLUSH', 'BROKER_TASK_STOP',
    'CLR_AUTO_EVENT', 'CLR_MANUAL_EVENT', 'LAZYWRITER_SLEEP',
    'SQLTRACE_BUFFER_FLUSH', 'WAITFOR', 'XE_TIMER_EVENT',
    'XE_DISPATCHER_WAIT', 'FT_IFTS_SCHEDULER_IDLE_WAIT',
    'BROKER_EVENTHANDLER', 'SP_SERVER_DIAGNOSTICS_SLEEP',
    'HADR_FILESTREAM_IOMGR_IOCOMPLETION', 'DIRTY_PAGE_POLL',
    'DISPATCHER_QUEUE_SEMAPHORE', 'QDS_PERSIST_TASK_MAIN_LOOP_SLEEP',
    'QDS_ASYNC_QUEUE', 'CHECKPOINT_QUEUE', 'REQUEST_FOR_DEADLOCK_SEARCH',
    'LOGMGR_QUEUE', 'ONDEMAND_TASK_QUEUE', 'HADR_WORK_QUEUE',
    'BROKER_TRANSMITTER', 'KSOURCE_WAKEUP'
)
AND waiting_tasks_count > 0
ORDER BY wait_time_ms DESC;
-- 2. Per-file write latency. JOIN sys.master_files; do NOT assume file_id = 2.
SELECT
    DB_NAME(vfs.database_id) AS database_name,
    mf.name AS file_name,
    mf.type_desc,
    vfs.num_of_writes,
    CASE WHEN vfs.num_of_writes > 0
         THEN vfs.io_stall_write_ms * 1.0 / vfs.num_of_writes
         ELSE 0 END AS avg_write_latency_ms
FROM sys.dm_io_virtual_file_stats(NULL, NULL) vfs
JOIN sys.master_files mf
    ON vfs.database_id = mf.database_id AND vfs.file_id = mf.file_id
WHERE mf.type_desc = 'LOG'
ORDER BY avg_write_latency_ms DESC;
-- 3. Outstanding I/O on the log file (helps detect the 112-cap scenario on fast SSD)
SELECT
    DB_NAME(database_id) AS database_name,
    file_id,
    io_pending,
    io_pending_ms_ticks
FROM sys.dm_io_pending_io_requests
WHERE io_pending = 1
ORDER BY database_id, file_id;
-- 4. Log flush wait time per database (cumulative; compute deltas)
SELECT
    instance_name AS database_name,
    cntr_value AS log_flush_wait_time_ms
FROM sys.dm_os_performance_counters
WHERE counter_name = 'Log Flush Wait Time'
  AND object_name LIKE '%Databases%'
  AND instance_name NOT IN ('_Total', 'mssqlsystemresource');
-- 5. Log flush rate vs log bytes flushed per database
SELECT
    instance_name AS database_name,
    cntr_value
FROM sys.dm_os_performance_counters
WHERE counter_name IN ('Log Flushes/sec', 'Log Bytes Flushed/sec', 'Transactions/sec')
  AND object_name LIKE '%Databases%'
  AND instance_name NOT IN ('_Total', 'mssqlsystemresource');
-- 6. Confirm synchronous AG contribution
SELECT wait_type, waiting_tasks_count, wait_time_ms
FROM sys.dm_os_wait_stats
WHERE wait_type IN ('HADR_SYNC_COMMIT', 'WRITELOG');

Query 7 below is a Windows command, not SQL. Run it from an elevated command prompt:

fltmc instances

Queries 1 and 2 are decisive. If query 2 shows the LOG file averaging more than 5ms per write, storage is the cause. If query 2 shows sub-millisecond writes but WRITELOG is still dominant, suspect the outstanding-I/O cap on fast SSDs, filter drivers, or AG synchronous commit.

How to diagnose it

  1. Snapshot sys.dm_os_wait_stats twice, 60 seconds apart, and compute the delta. Cumulative-since-startup values are useless for a live problem. WRITELOG must be a meaningful fraction of the delta.

  2. Run query 2 as a delta as well. Thresholds for log write latency: under 2ms healthy, 2-5ms acceptable, above 5ms degraded, above 15ms severe.

  3. If the LOG file shows sub-millisecond writes but WRITELOG is still dominant, check sys.dm_io_pending_io_requests for sustained near-112 outstanding requests per database. That is the per-database outstanding-I/O cap, raised from 32 in SQL Server 2008 to 112 starting in SQL Server 2012. On very fast SSDs at very high commit rates, this cap becomes the bottleneck rather than disk latency.

  4. Run fltmc instances from an elevated command prompt. If PerfMon shows healthy disk latency but SQL Server reports WRITELOG, the bottleneck is in the I/O stack between SQL Server and the Partition Manager: antivirus, backup agents, encryption products. Microsoft’s official guidance is to exclude SQL Server data, log, and backup files from real-time scanning.

  5. If you are running synchronous-commit Availability Groups, look at HADR_SYNC_COMMIT. In synchronous mode the primary waits for the secondary to harden the log before acknowledging the commit. HADR_SYNC_COMMIT typically dwarfs WRITELOG in that case, but WRITELOG still contributes. Check the secondary’s log-file I/O stall and the AG send queue.

  6. Distinguish storage latency from commit-rate pressure. Compare log bytes flushed per second against transactions per second. If transactions per second is very high and log bytes per transaction is small (single-row inserts or updates), the engine is asking the Log Writer to flush a small block for every commit. The fix is in the application, not the storage.

  7. Check VLF count via sys.dm_db_log_info(DB_ID('<db>')) (SQL 2016 SP2+) or DBCC LOGINFO. Excessive VLFs (more than a few hundred) slow log operations including writes, backup, and recovery. This is a secondary cause but worth ruling out.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
WRITELOG wait deltaIsolates the log path from other waitsAbove 5% of total wait delta correlated with user-visible commit latency
Log file write latencySQL Server’s direct view of storage latency per fileSustained above 5ms degraded, above 15ms severe
Log Flush Wait Time counterEngine-reported time spent waiting for flushesTrending up over the sampling window
Log Flushes/sec and Log Bytes Flushed/secCommit pressure on the Log WriterHigh flushes/sec with low bytes per flush indicates application-pattern issue
HADR_SYNC_COMMIT wait deltaDistinguishes local log latency from synchronous AG contributionDominant on primary with sync-commit AGs
Outstanding I/O on log filesDetects the 112-cap bottleneck on fast storageSustained near-112 outstanding requests per database
Transactions/sec to Batch Requests/sec ratioDetects transaction-rate-driven WRITELOGSkew toward many tiny transactions
VLF count per databaseExcess VLFs amplify log operationsAbove 1000 warrants consolidation

Fixes

Slow log-file storage

The most common cause and the most direct to fix. Move the log file to faster, dedicated storage. Log writes are sequential, so a single well-placed log file on low-latency disk (NVMe preferred, then SSD) outperforms multiple log files on slower shared storage.

If you cannot move the file immediately, isolate the log volume from data, TempDB, and backups. Log writes are latency-sensitive; any competing I/O on the same spindle hurts.

Do not add a second log file to parallelize log writes. SQL Server does not parallelize across log files within a database. Multiple log files are only useful for space, never for performance.

Filter driver overhead

Run fltmc instances from an elevated command prompt. Antivirus, backup agents, and encryption products insert themselves between SQL Server and the disk. Exclude SQL Server data, log, and backup files from real-time AV scanning. If you cannot exclude the files, schedule scans for maintenance windows.

Fast-SSD outstanding-I/O cap

If writes are sub-millisecond and the system is hitting the 112 outstanding-I/O cap per database, the bottleneck is not storage latency but commit serialization. Options:

  • Batch commits in the application. Group multiple writes into a single transaction so each log flush carries more bytes.
  • Use delayed durability carefully (see below).
  • Upgrade to SQL Server 2019 or later, which adds up to 8 Log Writer threads and allows regular worker threads to issue log writes directly.

Tiny transactions committing individually

The fix is in the application. Each BEGIN TRANCOMMIT for a single row insert forces a separate log flush. Re-architecting to batch (one transaction per N rows, or one transaction per business operation rather than per row) reduces flush count dramatically without changing storage.

Measure before changing code. Compare Log Flushes/sec to Transactions/sec. If they are nearly equal, every transaction is forcing its own flush. If Log Flushes/sec is much lower, transactions are already being grouped by the engine when their log records combine within the 60KB block window.

Synchronous-commit AlwaysOn

In synchronous-commit AGs, HADR_SYNC_COMMIT will be the larger wait, but WRITELOG still contributes. Investigate the secondary’s log-file I/O stall. If the secondary storage is slower than the primary, that latency is reflected back to the primary through the synchronous-commit handshake.

If commit latency is critical and the secondary cannot keep up, options include upgrading secondary storage, investigating network round-trip between replicas, or temporarily switching to asynchronous commit during the incident. Switching to asynchronous commit creates a data-loss window; acknowledge the risk explicitly before making the change.

Delayed durability as a last resort

SQL Server 2014+ supports delayed durability at the database level (DISABLED / ALLOWED / FORCED), the commit level, or the atomic block level. With delayed durability, the engine acknowledges the commit before the log flush completes, removing WRITELOG from the critical path.

This is a tradeoff, not a fix. You gain commit latency and lose ACID durability. On crash, transactions acknowledged as committed may be rolled back during recovery. Delayed durability is incompatible with transactional replication, Change Data Capture, cross-database and DTC transactions, and Azure Synapse Link. Starting with SQL Server 2022 CU2 and SQL Server 2019 CU20, attempting to enable it alongside these features raises errors 22891 or 22892.

It also hides the underlying problem. WRITELOG disappears from the waits list because the engine is no longer waiting, but the storage is still slow. Use it deliberately for known workloads where the data-loss tradeoff is acceptable, not as a general performance tweak.

Prevention

  • Pre-size log files. Log auto-growth is expensive. Instant File Initialization does not apply to log files, so each growth event requires zero-initialization and pauses all log writes to that file. Pre-allocate based on the volume of log generated between log backups.
  • Keep VLF count low. Periodically check via sys.dm_db_log_info or DBCC LOGINFO. If the count is in the thousands, shrink and regrow the log in appropriate increments.
  • Keep the log on dedicated, low-latency storage. NVMe or SSD is the baseline for OLTP.
  • Validate filter driver exclusions as part of deployment. AV exclusions drift over time as security tooling changes.
  • Baseline WRITELOG and log write latency. The most common failure mode is not knowing whether current WRITELOG is normal or new. Snapshot wait stats periodically (30 to 60 seconds) and store the deltas externally. DMVs reset on instance restart.
  • For applications committing in tight loops, plan batching as a capacity practice, not a one-off fix.

How Netdata helps

  • Per-second collection of SQL Server wait statistics makes WRITELOG visible as a time series rather than a single cumulative number, so you can correlate a WRITELOG spike with a deploy, a traffic spike, or a storage event.
  • Correlating WRITELOG with Log Flush Wait Time and per-file write latency from sys.dm_io_virtual_file_stats separates storage-driven WRITELOG from commit-rate-driven WRITELOG in a single view.
  • Per-second Transactions/sec, Batch Requests/sec, and Log Flushes/sec surface the tiny-transactions pattern: high commit rate, low bytes per commit.
  • For AlwaysOn deployments, pairing WRITELOG with HADR_SYNC_COMMIT waits and send-queue size tells you whether the bottleneck is local storage or replica hardening.
  • ML-based anomaly detection on log write latency and outstanding I/O catches gradual storage degradation before WRITELOG becomes the dominant wait.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.