The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-log-autogrow-stall

Operations Guides

SQL Server log autogrow stall: why every write pauses while the log file grows

Applications report periodic, mysterious write stalls. Throughput drops briefly, then recovers. Repeat. The transaction log on the affected database is creeping upward, and there are no alarms on disk space. What you are seeing is almost certainly log autogrow stall: every write transaction in the database pauses while SQL Server expands and zero-initializes the transaction log file.

Log autogrow is one of the most expensive routine operations SQL Server performs against its own files. Unlike data files, the transaction log cannot use Instant File Initialization (IFI). Every byte of new log space must be written with zeros before the engine accepts new log records. While that zeroing happens, every transaction that needs to log a modification waits. The wider and slower your storage, the longer the stall.

This guide covers the mechanism, detection from the default trace and error log, how to distinguish the pattern from other write-latency problems, and how to size and configure the log file so the engine never autogrows during normal operations.

What this means

When SQL Server decides the transaction log file needs to grow, it issues a single serialized operation that pauses all I/O to that file. The expansion is not concurrent with normal work; it gates normal work. While the new extent is allocated and zeroed, every worker thread that needs to log a modification waits. From the application’s perspective, every write query hangs at the same instant.

The reason log growth is so much more expensive than data file growth is IFI. SQL Server can use IFI to skip zero-initialization for data files when the service account holds the SE_MANAGE_VOLUME_NAME privilege. IFI does not apply to log files. Every growth event writes zeros across the entire new allocation. On direct-attached NVMe this may be fast enough to ignore. On cloud-managed disks with IOPS caps, on shared SAN storage, or on network-attached volumes, the zero-initialization can take seconds. On a percentage-based autogrow for a multi-hundred-GB log, it can take tens of seconds and eventually time out.

SQL Server 2022 (16.x) introduced a narrow exception: transaction log autogrowth events up to 64 MB can benefit from IFI. This is helpful but is not a license to keep autogrow on as a sizing strategy. For a multi-TB database, chasing the 64 MB optimization by triggering hundreds of small growth events creates thousands of Virtual Log Files (VLFs), which degrades crash recovery, backup, and log-reader performance. The correct posture is to pre-size the log and reserve autogrow as an emergency pressure relief valve, not as the primary growth mechanism.

flowchart TD
    A[Log file near full] --> B[Autogrow triggered]
    B --> C[SQL Server pauses all log writes]
    C --> D[Zero-initialize new extent]
    D --> E{Completed in time?}
    E -- Yes --> F[Msg 5145 logged
writes resume] E -- No --> G[Msg 5144: autogrow cancelled] G --> H{Retry succeeds?} H -- Yes --> F H -- No --> I[Error 9002
writes fail]

Common causes

CauseWhat it looks likeFirst thing to check
Log file never pre-sizedAutogrow events at predictable intervals during peak load, clustering around write-heavy batchessys.database_files current size vs. expected peak working window
Percentage-based autogrowthEach autogrow event takes longer than the last as the file grows; eventually starts timing outis_percent_growth and growth columns in sys.database_files
Missing log backups in full or bulk-logged recoveryLog climbs steadily between backups; each backup triggers truncation, then the climb resumeslog_reuse_wait_desc in sys.databases; msdb.dbo.backupset for type='L'
Long-running active transactionLog climbs and will not truncate even after log backups; ACTIVE_TRANSACTION reuse waitDBCC OPENTRAN and sys.dm_tran_active_transactions
Slow storage with tiny growth incrementsMany short autogrow events; default trace shows hundreds of 1 MB eventsDefault trace EventClass 93 events; I/O stall on the log file

Quick checks

-- 1. Log space and reuse wait per database
SELECT
    db.name,
    ls.cntr_value AS log_used_pct,
    db.log_reuse_wait_desc,
    db.recovery_model_desc
FROM sys.dm_os_performance_counters ls
JOIN sys.databases db ON ls.instance_name = db.name
WHERE ls.counter_name = 'Percent Log Used'
  AND ls.object_name LIKE '%Databases%';
-- 2. Log file sizing and growth configuration
SELECT
    DB_NAME(database_id) AS db_name,
    name AS logical_file_name,
    size * 8 / 1024 AS current_size_mb,
    max_size,
    growth,
    is_percent_growth
FROM sys.master_files
WHERE type_desc = 'LOG'
ORDER BY DB_NAME(database_id);
-- 3. Recent autogrow events from the default trace
SELECT
    te.name AS event_name,
    t.DatabaseName,
    t.FileName,
    t.Duration AS duration_ms,
    t.StartTime,
    t.IntegerData * 8 / 1024 AS growth_mb
FROM sys.fn_trace_gettable(
    (SELECT [path] FROM sys.traces WHERE is_default = 1), DEFAULT) t
JOIN sys.trace_events te ON t.EventClass = te.trace_event_id
WHERE te.name IN ('Data File Auto Grow', 'Log File Auto Grow')
ORDER BY t.StartTime DESC;
-- 4. VLF count for one database
SELECT DB_NAME(database_id) AS db_name, COUNT(*) AS vlf_count
FROM sys.dm_db_log_info(DB_ID('YourDatabaseName'))
GROUP BY database_id;
-- 5. Autogrow timeouts or slow growth in the current error log
EXEC sp_readerrorlog 0, 1, 'Autogrow';
-- 6. Engine version (16 = SQL Server 2022, qualifies for 64 MB log IFI)
SELECT
    SERVERPROPERTY('ProductVersion') AS version,
    SERVERPROPERTY('ProductMajorVersion') AS major_version;
-- 7. Log file I/O stall
SELECT
    DB_NAME(vfs.database_id) AS db_name,
    mf.name AS file_name,
    vfs.io_stall_write_ms,
    vfs.num_of_writes,
    CASE WHEN vfs.num_of_writes > 0
         THEN vfs.io_stall_write_ms * 1.0 / vfs.num_of_writes
         ELSE 0 END AS avg_write_latency_ms
FROM sys.dm_io_virtual_file_stats(NULL, NULL) vfs
JOIN sys.master_files mf
    ON vfs.database_id = mf.database_id AND vfs.file_id = mf.file_id
WHERE mf.type_desc = 'LOG'
ORDER BY vfs.io_stall_write_ms DESC;

How to diagnose it

  1. Confirm the symptom. Check the default trace for recent Log File Auto Grow events (EventClass 93). If you see events during the window when users reported stalls, you have your culprit. The Duration column is reported in milliseconds for the auto-grow event classes.

  2. Read the error log for autogrow timeouts. Search for “Autogrow” or the message IDs 5144 and 5145. Msg 5144 indicates an autogrow that was cancelled by the user or timed out; this is the smoking gun for a growth event the engine could not complete in time. Msg 5145 indicates an autogrow that completed but took longer than SQL Server considers healthy.

  3. Inspect the growth configuration. is_percent_growth = 1 is almost always wrong for a transaction log. Percentage growth on a large log is catastrophic: a 10% growth on a 345 GB log triggers a 34.5 GB zero-initialization event that will time out and retry in a loop.

  4. Examine log_reuse_wait_desc. If the log is growing because of LOG_BACKUP, fix the backup chain. If ACTIVE_TRANSACTION, find and address the open transaction. If REPLICATION or AVAILABILITY_REPLICA, the secondary or the replication agent is not consuming log fast enough. Adding log space does not fix any of these; it only delays the next autogrow.

  5. Check VLF count. sys.dm_db_log_info(DB_ID()) returns one row per VLF. Hundreds of small autogrow events produce thousands of VLFs. Above 1000 you should plan a log rebuild. Above 10000 you will see it in crash recovery time.

  6. Correlate with wait statistics. PREEMPTIVE_OS_WRITEFILEGATHER is the wait type SQL Server records while zero-initializing a file. A spike in this wait during the stall window confirms the mechanism. WRITELOG may also be elevated, but it represents normal log flush latency, not autogrowth specifically.

  7. Verify whether the SQL Server 2022+ log IFI optimization applies. SERVERPROPERTY('ProductMajorVersion') returns 16 for SQL Server 2022. If the major version is 16 or higher and your autogrowth increment is 64 MB or less, the zero-initialization is partially mitigated. If the growth increment is larger than 64 MB, the optimization does not apply.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Log File Auto Grow events (default trace EventClass 93)Direct evidence the engine paused writes to grow the logAny event during business hours; events clustering in bursts
Autogrow Duration from the default traceMeasures the actual stall lengthEvents over 100 ms are noticeable; events over 1000 ms are user-visible
Percent Log Used per databaseLeading indicator of imminent growthSustained above 70% with a non-NOTHING reuse wait
Log file I/O write latencySurfaces underlying storage contentionSustained average above 5 ms; degraded trend over time
PREEMPTIVE_OS_WRITEFILEGATHER wait timeThe wait type for file zero-initializationNonzero and growing in proportion to total waits
VLF count per databaseFragmentation that worsens with each autogrowOver 1000 warrants consolidation; over 10000 is severe
Msg 5144 and Msg 5145 in the error logGrowth timed out or completed slowlyAny occurrence indicates undersized log or slow storage
log_reuse_wait_descRoot-cause classifier for growthLOG_BACKUP, ACTIVE_TRANSACTION, or REPLICATION are the usual drivers
Free space on the log volumeLimits how much more the log can grow before exhausting the diskBelow 20% is a capacity planning failure

Fixes

Pre-size the log file

Pre-sizing is the only correct long-term answer. Estimate the log space needed to span your longest expected interval between log backups (typically 15 to 60 minutes for full recovery) plus headroom for batch operations and index rebuilds. Grow the file once, deliberately, during a maintenance window when the storage subsystem can absorb the zero-initialization without impacting user workloads.

-- Grow the log file to a fixed size
ALTER DATABASE [YourDatabase]
MODIFY FILE (NAME = [YourDatabase_Log], SIZE = 50GB);

Set a sane fixed growth increment

If you must keep autogrow enabled as a safety net, set a fixed MB increment rather than a percentage. Choose a size large enough that the engine does not autogrow repeatedly under burst load, but small enough that the zero-initialization completes within SQL Server’s growth timeout. For SQL Server 2022+ instances, 64 MB is a useful ceiling because it qualifies for the partial IFI optimization. For older versions, choose an increment based on observed zero-initialization throughput on your storage.

-- Set a fixed 512 MB growth increment
ALTER DATABASE [YourDatabase]
MODIFY FILE (NAME = [YourDatabase_Log], FILEGROWTH = 512MB);

Avoid percentage growth. There is no operationally defensible case for is_percent_growth = 1 on a production transaction log.

Address the underlying growth driver

If the log fills because log backups are failing, fix the backups. If a long-running transaction is holding the log open, find it with DBCC OPENTRAN and either let it complete or kill it. If replication or an AlwaysOn secondary is not consuming log fast enough, address the latency on that path. Adding log space without fixing the root cause only postpones the next stall.

See the SQL Server log_reuse_wait_desc and SQL Server log backups missing guides for the full root-cause tree.

Consolidate VLFs

If VLF count is already in the thousands, plan a log rebuild during a maintenance window. The pattern: take a log backup, shrink the log in small increments using DBCC SHRINKFILE, then grow it back to the target size in fixed chunks. This produces a small number of large VLFs instead of thousands of small ones. Do this once after fixing the autogrow configuration, not on a recurring schedule.

Warning: DBCC SHRINKFILE on the log is disruptive and generates heavy I/O on the volume. Never schedule it during peak load, and never run it on a recurring basis. Shrink-then-grow cycles are a one-time remediation, not maintenance.

When autogrow has already timed out

If the error log shows Msg 5144 followed by Error 9002, the database is no longer accepting writes. Immediate relief options, in order of preference: take a log backup if log_reuse_wait_desc permits; add a second log file on a faster volume as emergency capacity; switch the database to simple recovery model if losing point-in-time recovery between full and differential backups is acceptable. Switching to simple breaks the log chain since the last log backup and requires a full backup to restart it. None of these are good permanent fixes; each is a way to restore writes while you address the underlying sizing problem. See SQL Server Error 9002 for the full recovery procedure.

Prevention

  • Pre-size log files at provisioning time. Treat autogrow as a break-glass mechanism. The log should be sized to span the longest expected gap between log backups.
  • Set a fixed MB growth increment, never percentage. Choose a value that completes within SQL Server’s growth timeout on your slowest storage tier.
  • Alert on autogrow events, not just disk space. Low disk is too late. Any Log File Auto Grow event in the default trace is a sign the log is undersized.
  • Alert on Msg 5144 and Msg 5145. These mean autogrow is no longer completing within healthy timeframes. Treat them as capacity incidents.
  • Monitor VLF count per database. Trend it over time. A growing VLF count means autogrow is happening in small increments.
  • Verify the log backup chain. For full and bulk-logged recovery models, the most common cause of log growth is failed or missing log backups.
  • Track log_reuse_wait_desc persistently. Any non-NOTHING value that persists across log backups is a growth driver that will eventually force autogrow.

How Netdata helps

  • Per-second transaction log percent used makes the climb toward the next autogrow event visible long before the stall. Correlate the climb with batch requests per second to distinguish a workload burst from a stuck backup.
  • Default-trace-based autogrow event capture turns Log File Auto Grow events into first-class time series. The event becomes a marker on the chart, not something you discover after the fact in the error log.
  • Log file I/O latency from sys.dm_io_virtual_file_stats surfaces the storage-side contribution to the stall. If write latency on the log file is already elevated before the autogrow event, the stall will be longer.
  • Wait statistics deltas expose PREEMPTIVE_OS_WRITEFILEGATHER and WRITELOG spikes in real time. The wait-type signature of an autogrow stall is distinctive.
  • VLF count tracking per database flags the slow accumulation that turns one bad autogrow configuration into a recovery-time problem months later.
  • Composite alerting across log used percent, autogrow events, and log_reuse_wait_desc distinguishes “log grew once due to a batch job” from “log is growing every few minutes because backups are broken.”

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.