The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-log-backup-chain-broken

Operations Guides

SQL Server log backups missing: the full-recovery log that grows forever

The application starts throwing write errors. The database is online, reads work, but every INSERT, UPDATE, and DELETE fails with error 9002: the transaction log is full. The log volume is at zero free space, or the log file has auto-grown to many times the size of the data files. When you ask when the last log backup ran, nobody knows.

This is almost always the same root cause: the database is in FULL or BULK_LOGGED recovery, and log backups have stopped running. A full database backup does not truncate the log. Only BACKUP LOG does. If log backups are missing, disabled, or silently failing, the log accumulates every transaction forever until the file hits its max size or the disk fills.

The engine tells you exactly why the log cannot be truncated, and the fix is usually fast once you read the right signal. This guide covers distinguishing a dead backup job from a broken log chain, verifying what actually happened in msdb, and recovering without making things worse.

What this means

In FULL or BULK_LOGGED recovery, SQL Server keeps every log record so you can restore to any point in time by replaying an unbroken sequence of log backups: a full backup, then every log backup after it, in LSN order. That sequence is the log chain. Log space is only marked reusable (truncated) after a log backup has captured it. No log backup, no truncation, unlimited growth.

Two distinct failure modes produce the same symptom:

  • Log backups stopped running. The SQL Agent job was disabled, deleted, misscheduled, or is failing. Or the job “succeeds” but writes to a bad target: a full backup volume, a deleted network share, or NUL. The msdb history is the authority here, not the job status.
  • The log chain is broken. After a restore, a detach/attach, or a recovery model change, the previous chain no longer applies. A database switched from SIMPLE to FULL recovery needs a fresh full backup before any log backup will work; until then, log backups either fail or the database behaves as if it were still in SIMPLE (“pseudo-simple”). After a restore, a new full backup restarts the chain.

Either way, the database keeps writing log records it cannot reuse, and the file grows until writes stop.

Common causes

CauseWhat it looks likeFirst thing to check
Log backup job disabled, deleted, or failinglog_reuse_wait_desc = LOG_BACKUP; last type ‘L’ row in msdb.dbo.backupset is hours or days oldLast log backup per database in msdb.dbo.backupset
Job “succeeds” but writes to a bad targetJob history green, but backup files missing, zero-length, or on a full/deleted share; backups recorded to NULphysical_device_name in msdb.dbo.backupmediafamily; confirm the file exists and has size
Recovery model switched SIMPLE to FULL, no new full backupLog backups fail or log grows anyway; chain never restartedrecovery_model_desc in sys.databases plus last full backup (type ‘D’) in backupset
Restore or detach/attach broke the chainLog backups error about the chain; no common basebackupset history around the restore time
Someone ran BACKUP LOG to NUL to “clear” the logPoint-in-time recovery silently destroyed; backupset shows physical_device_name = ‘NUL’backupmediafamily for NUL devices
Long-running transaction, replication, or AG lag (not a backup problem at all)log_reuse_wait_desc = ACTIVE_TRANSACTION, REPLICATION, or AVAILABILITY_REPLICAlog_reuse_wait_desc in sys.databases before touching backups
Third-party VSS snapshot backups assumed to truncateSnapshot backups may not truncate the log; log keeps growing despite “backups”is_snapshot / is_copy_only flags in backupset for recent backups

Truncation by VSS snapshot backups depends on the requestor: full backups taken through the SQL Writer are real full backups (eligible as differential bases), and tools that request log truncation through the writer will truncate. Copy-only VSS snapshots never truncate, VM-level snapshots without SQL Writer integration do not truncate (and non-component-based VSS backups support only databases in SIMPLE recovery), and the SQL Writer does not support log backups at all.

One thing that does not break the chain: copy-only backups. BACKUP DATABASE … WITH COPY_ONLY and BACKUP LOG … WITH COPY_ONLY are safe for ad-hoc copies and leave the chain intact. Copy-only log backups also do not truncate the log, which surprises people who run one and expect space back.

Quick checks

All read-only.

-- 1. Why can't the log truncate? This is the authoritative first question.
SELECT name, recovery_model_desc, log_reuse_wait_desc
FROM sys.databases;
-- 2. Log usage per database (percentage and absolute sizes).
SELECT
    db.name,
    ls.cntr_value AS log_space_used_pct,
    db.log_reuse_wait_desc
FROM sys.dm_os_performance_counters ls
JOIN sys.databases db ON ls.instance_name = db.name
WHERE ls.counter_name = 'Percent Log Used'
  AND ls.object_name LIKE '%Databases%';

DBCC SQLPERF(LOGSPACE);
-- 3. Backup freshness: last full, diff, and log backup per database.
-- msdb.dbo.backupset is the authoritative record, not job history.
SELECT
    d.name AS database_name,
    d.recovery_model_desc,
    MAX(CASE WHEN bs.type = 'D' THEN bs.backup_finish_date END) AS last_full_backup,
    MAX(CASE WHEN bs.type = 'I' THEN bs.backup_finish_date END) AS last_diff_backup,
    MAX(CASE WHEN bs.type = 'L' THEN bs.backup_finish_date END) AS last_log_backup
FROM sys.databases d
LEFT JOIN msdb.dbo.backupset bs ON d.name = bs.database_name
WHERE d.database_id > 4
  AND d.state_desc = 'ONLINE'
GROUP BY d.name, d.recovery_model_desc
ORDER BY last_log_backup;
-- 4. Where did recent backups actually go, and were they copy-only or snapshots?
SELECT TOP 50
    bs.database_name,
    bs.type,                       -- D=full, I=diff, L=log
    bs.is_copy_only,
    bs.backup_start_date,
    bs.backup_finish_date,
    bmf.physical_device_name
FROM msdb.dbo.backupset bs
JOIN msdb.dbo.backupmediafamily bmf ON bs.media_set_id = bmf.media_set_id
ORDER BY bs.backup_start_date DESC;

Look for physical_device_name = ‘NUL’ (someone discarded log to force truncation and broke the chain), paths on volumes that no longer exist or are full, and snapshot backups from VSS-based tools that do not truncate.

-- 5. Is an open transaction the real blocker?
DBCC OPENTRAN;
-- 6. Free space on the log volume.
SELECT DISTINCT
    vs.volume_mount_point,
    vs.available_bytes / 1048576 AS available_mb
FROM sys.master_files mf
CROSS APPLY sys.dm_os_volume_stats(mf.database_id, mf.file_id) vs
ORDER BY available_mb;

How to diagnose it

flowchart TD
    A[Log usage high or error 9002] --> B{log_reuse_wait_desc?}
    B -->|LOG_BACKUP| C{Last log backup in backupset?}
    B -->|ACTIVE_TRANSACTION| D[DBCC OPENTRAN - find and end it]
    B -->|REPLICATION / AVAILABILITY_REPLICA| E[Check replication agent / AG secondary lag]
    C -->|Recent, regular| F[Backups exist but log still grows - check for NUL backups, tiny backups clearing zero VLFs]
    C -->|Missing or stale| G{Job enabled and succeeding?}
    G -->|Job fine| H[Check backup target: share deleted, volume full, files zero-length]
    G -->|Job broken or gone| I[Re-enable / fix job, take log backup now]
    C -->|Chain broken| J[Take a full backup to restart the chain, then log backups]
  1. Read log_reuse_wait_desc first. Do not touch backup jobs until you know why the log cannot truncate. LOG_BACKUP means a log backup will fix it. ACTIVE_TRANSACTION, REPLICATION, or AVAILABILITY_REPLICA mean the backup chain is not your problem; find the long transaction (DBCC OPENTRAN), the stalled replication agent, or the lagging AG secondary instead. Adding space does not fix any of these.

  2. Verify the backup history in msdb, not the job status. The SQL Agent job can complete “successfully” while writing to a deleted share or a full volume. msdb.dbo.backupset (joined to backupmediafamily) is the authoritative record of what was actually backed up and where. Confirm the last type ‘L’ row is recent and that the file at physical_device_name exists with non-zero size.

  3. Check for a broken chain. If the database was recently restored, detached/attached, or switched from SIMPLE to FULL recovery, the previous chain is gone. The database needs a new full backup before log backups resume meaningfully. Check backupset for the gap in type ‘L’ rows and whether a full backup exists after the switch or restore.

  4. Check for NUL backups. If someone previously “fixed” log growth with BACKUP LOG TO DISK = ‘NUL’, you will see it in backupmediafamily. That backup discarded log records and broke point-in-time recovery from that point. The chain continues afterward only from the next real log backup. Warn whoever did it: TO NUL is not a substitute for TRUNCATE_ONLY (which was removed in SQL Server 2008); the supported escape hatch is switching to SIMPLE recovery.

  5. Rule out the red herring: one log backup did not shrink anything. A log backup truncates the log logically (marks VLFs reusable); it does not shrink the physical file. If log_reuse_wait_desc still shows LOG_BACKUP immediately after a backup, the backup may have cleared zero VLFs because almost nothing was in the inactive portion; the next backup clears it. And if the file itself is now absurdly large, that is a separate shrink decision, not a backup problem.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Percent Log Used per databaseDirect measure of how close the log is to stopping writesAbove 70% with non-NOTHING reuse wait; above 90% is imminent write failure
log_reuse_wait_descTells you exactly why truncation is blockedAnything other than NOTHING persisting across intervals
Time since last log backup (msdb.dbo.backupset, type ‘L’)Recovery point exposure and truncation healthOver 1 hour for a FULL recovery database
Time since last full backup (type ‘D’)Restore capability; also restarts chains after recovery model changesOver 24 hours for production
Log file auto-growth eventsEach event pauses all log-writing transactions; IFI does not apply to log filesAny growth during business hours
VLF count (sys.dm_db_log_info or DBCC LOGINFO)Thousands of tiny VLFs slow recovery, backup, and restoreOver 1000; over 10,000 is severe
Error 9002 in the error logWrites are already failingAny occurrence is a page
Backup destination validityA recorded backup to a dead target is not a backupNUL devices, unreachable shares, zero-length files

Fixes

Log backups stopped running

Take a log backup immediately, then fix the job:

BACKUP LOG [YourDb] TO DISK = N'<your normal backup path>';

Then re-enable or recreate the scheduled log backup job on a cadence that matches your RPO, typically every 15-60 minutes for production FULL recovery databases. Verify the next few scheduled runs land in backupset with valid files. If the job was writing to a bad target, fix the target and confirm free space on the backup volume.

Chain broken after restore, detach/attach, or recovery model switch

Take a full database backup to start a new chain:

BACKUP DATABASE [YourDb] TO DISK = N'<your normal backup path>';

Log backups work from this point forward. Be explicit with the team: recovery is only possible to points after this full backup.

Log is full right now and the disk is also full

If there is no room anywhere for a log backup to truncate into, add a second log file on a different volume as emergency relief, take a log backup, and then remove the temporary file once the log truncates. Do not shrink first; a full log has nothing to shrink until it truncates.

-- Emergency: add a second log file on a volume with free space.
ALTER DATABASE [YourDb] ADD LOG FILE (
    NAME = N'YourDb_log2',
    FILENAME = N'<path on a different volume>\YourDb_log2.ldf',
    SIZE = 1GB,
    FILEGROWTH = 256MB
);

After the log truncates and pressure is off, take a log backup and then remove the emergency file:

-- Remove the emergency log file only after the log has truncated.
-- DBCC SHRINKFILE with EMPTYFILE moves active log out before removal.
DBCC SHRINKFILE (N'YourDb_log2', EMPTYFILE);
ALTER DATABASE [YourDb] REMOVE FILE [YourDb_log2];

Log file physically oversized after the incident

Once the log is truncating normally (log_reuse_wait_desc = NOTHING) and usage is low, you can shrink the file back to a sane pre-allocated size:

-- One-time corrective action. Do not schedule this.
-- Check current size and VLF count before shrinking.
DBCC SHRINKFILE ([YourDb_log], <target_size_MB>);

Repeated shrink/regrow cycles fragment the log into thousands of VLFs and every regrowth zero-initializes (IFI does not apply to log files), stalling writers. Size the file for peak volume between log backups and leave it there.

The database never needed point-in-time recovery

If the honest answer is “we would restore last night’s full backup and accept the data loss,” the database belongs in SIMPLE recovery:

-- WARNING: This breaks the existing log chain. Any differential or log
-- backups taken before this point become unusable for point-in-time restore.
-- Take a full backup immediately after switching.
ALTER DATABASE [YourDb] SET RECOVERY SIMPLE;

SIMPLE truncates the log at checkpoint, so no log backups are needed and this failure mode disappears. The tradeoff is real: you lose point-in-time restore and you break any existing log chain (take a full backup after switching). Do not use this as a panic button on databases that actually have an RPO; use it as a deliberate per-database decision.

Prevention

  • Alert on log backup age, not job status. Query msdb.dbo.backupset for the last type ‘L’ backup per database and alert when it exceeds your RPO. This catches disabled jobs, deleted jobs, and successful-jobs-writing-nowhere in one check.
  • Alert on Percent Log Used above 70% with a non-NOTHING log_reuse_wait_desc. You want to know about truncation blockers hours before error 9002, not during it.
  • Alert on log auto-growth events. Growth during business hours means the log was undersized or truncation is lagging, and every growth event stalls writers while the new space is zero-initialized.
  • Never allow BACKUP LOG TO NUL. Treat physical_device_name = ‘NUL’ in backupmediafamily as a finding. It silently destroys point-in-time recoverability.
  • Match recovery model to actual RPO. Audit databases in FULL recovery. Any database without a working log backup schedule either gets one or moves to SIMPLE.
  • Test restores. RESTORE VERIFYONLY validates media but does not guarantee a restore works. Periodic actual restore tests are the only proof the chain is usable.
  • Watch the chain across operational events. Restores, detach/attach, recovery model changes, and DR drills all restart or break chains. After any of these, verify a fresh full backup exists.

How Netdata helps

  • Log usage over time: Netdata’s SQL Server collector surfaces per-database Percent Log Used and log file size, so you can see log percent climbing over time instead of discovering it at error 9002.
  • Backup freshness as a first-class signal: tracking hours since last full and last log backup per database turns “the job said success” into “the authoritative backupset record says nothing landed in 26 hours.”
  • Growth events in context: log file growth correlated with write latency stalls explains the periodic commit hiccups teams usually blame on storage.
  • Volume free space on log and backup targets: catching the backup volume filling up before the job starts “succeeding” into a dead target.
  • Trend over time: a log that creeps upward week over week is a missing-backup problem in slow motion; per-second and historical views make the trend obvious long before the cliff.
The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.