The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-restore-readiness

Operations Guides

SQL Server restore readiness: why a backup you never test is not a backup

The most common SQL Server backup monitoring answers one question: did the job succeed? A green checkmark means SQL Server wrote bytes to a file. It does not mean those bytes can be read back, decrypted, applied through a recovery sequence, and brought online within your RTO.

The gap between “backup succeeded” and “we can recover” is where most DR failures actually live. Backups fail to restore because of media corruption the backup job never validated, because a TDE or backup-encryption certificate was not backed up alongside the data, because an ad-hoc operation broke the log chain and silently invalidated hours of log backups, or because the restore takes four times longer than anyone measured. None of those surface in backup-completion monitoring.

This article covers how to validate backups beyond the job’s success flag, how to schedule and measure real restore tests, what cryptographic material you must protect alongside the data, and how to keep log chains intact when you pull ad-hoc copies for testing.

What restore readiness actually means

Restore readiness is the verified ability to recover a database to a defined point in time, on a target instance, within a measured duration. It has four components.

  • Backup freshness. A recoverable backup exists and is recent enough to meet RPO.
  • Backup integrity. The backup media is readable and the data inside it is internally consistent.
  • Restore feasibility. Every prerequisite the restore needs is in place: certificates, private keys, reachable storage, a compatible target instance, and an unbroken log chain.
  • Measured RTO. Someone has actually run the restore end-to-end and timed it.

Backup-completion monitoring covers only the first. It is necessary and not sufficient.

Why backup success is not recoverability

flowchart TD
    A[Backup job success] --> B[VERIFYONLY passes]
    B --> C[Restore test passes]
    C --> D[CHECKDB clean on restored copy]
    D --> E[Restore duration measured]
    A -.- A1[Only proves bytes were written]
    B -.- B1[Misses page corruption without CHECKSUM]
    C -.- C1[Catches missing cert and broken chain]
    D -.- D1[Only true corruption validation]
    E -.- E1[Only real RTO number]

The first trap is RESTORE VERIFYONLY. It is widely used as the post-backup integrity check, and Ola Hallengren’s maintenance solution wires it in via @Verify = 'Y'. Per Microsoft’s documentation, VERIFYONLY checks that the backup set is complete and the backup is readable, but it does not attempt to verify the structure of the data contained in the backup volumes. In practice it catches media-level and header problems but not page-level corruption unless BACKUP ... WITH CHECKSUM was used at backup time. A backup that passes VERIFYONLY can still fail to restore, or can restore a corrupt database.

The second trap is msdb.dbo.backupset. A job can mark itself successful without a row ever being written to backupset, for example when the job script swallows errors or when the target path is unreachable after the file handle closes. Always check the authoritative record, not the job history table.

The third trap is cryptographic material. TDE-encrypted databases and backup-encrypted backups require the server certificate (and its private key) to be present on the restore target. If the only copy of that certificate lived on the source instance, and the source instance is the thing you are recovering from, the backup is unrecoverable. Certificate loss equals backup loss.

The fourth trap is log chain continuity. Point-in-time recovery requires an unbroken sequence of log backups anchored to a full backup. Switching recovery models from FULL to SIMPLE and back, detaching and reattaching a database, or letting a log shipping target come online with RECOVERY can all break the chain. A new full or differential backup re-anchors it, but the window between the break and the re-anchor is not point-in-time recoverable.

The restore-readiness procedure

1. Track backup freshness from the authoritative source

-- Authoritative backup history per database
SELECT
    d.name AS database_name,
    d.recovery_model_desc,
    MAX(CASE WHEN bs.type = 'D' THEN bs.backup_finish_date END) AS last_full_backup,
    MAX(CASE WHEN bs.type = 'I' THEN bs.backup_finish_date END) AS last_diff_backup,
    MAX(CASE WHEN bs.type = 'L' THEN bs.backup_finish_date END) AS last_log_backup,
    DATEDIFF(HOUR, MAX(CASE WHEN bs.type = 'D' THEN bs.backup_finish_date END), GETDATE()) AS hours_since_full
FROM sys.databases d
LEFT JOIN msdb.dbo.backupset bs ON d.name = bs.database_name
WHERE d.database_id > 4  -- exclude master, tempdb, model, msdb
  AND d.state_desc = 'ONLINE'
GROUP BY d.name, d.recovery_model_desc
ORDER BY hours_since_full DESC;

This is the source of truth for backup freshness. Job status alone is insufficient: a job can return success while the backupset row is never written because the target path was unreachable after the file handle closed.

2. Validate media with RESTORE VERIFYONLY

-- Validate a backup file's media and header integrity
RESTORE VERIFYONLY
FROM DISK = N'/backups/prod/MyDb_full_20260720.bak'
WITH CHECKSUM;

WITH CHECKSUM only adds value if the backup was taken with BACKUP ... WITH CHECKSUM. Without checksums at backup time, VERIFYONLY falls back to header and media checks. Treat VERIFYONLY success as “media is readable”, not “restore will succeed”.

3. Take ad-hoc backups safely with COPY_ONLY

-- Ad-hoc backup that does NOT break the log chain or reset the differential base
BACKUP DATABASE [MyDb]
TO DISK = N'/restoretests/MyDb_copyonly_20260720.bak'
WITH COPY_ONLY, CHECKSUM, STATS = 10;

Copy-only backups are the correct way to capture a database for restore testing on a separate target without disrupting the production recovery chain. A regular full backup would establish a new differential base LSN and force the next differential cycle to rebuild from scratch.

4. Run periodic end-to-end restore tests

This is the gold standard. On a separate, non-production instance, restore the most recent full backup, the latest differential, and the log sequence to a chosen point in time. Time the entire operation. The measured duration is your empirical RTO for that database, and it is almost always longer than the number in the DR plan.

-- Destructive on the target instance only. Restores under a different name.
RESTORE DATABASE [MyDb_RestoreTest]
FROM DISK = N'/restoretests/MyDb_full.bak'
WITH MOVE N'MyDb' TO N'/data/MyDb_RestoreTest.mdf',
     MOVE N'MyDb_log' TO N'/log/MyDb_RestoreTest.ldf',
     NORECOVERY, REPLACE, STATS = 10;

RESTORE DATABASE [MyDb_RestoreTest]
FROM DISK = N'/restoretests/MyDb_diff.bak'
WITH NORECOVERY, STATS = 10;

RESTORE LOG [MyDb_RestoreTest]
FROM DISK = N'/restoretests/MyDb_log.trn'
WITH RECOVERY, STOPAT = '2026-07-20T12:00:00', STATS = 10;

The REPLACE and MOVE clauses let you land the restore under a different name on different paths. NORECOVERY between full and differential keeps the restore sequence open until the final RECOVERY.

Cadence depends on tier. Tier-1 systems typically need monthly restore tests; lower tiers may test quarterly or annually. Whatever you choose, record the date and the duration per database. A backup that has never been restored is not a recovery asset, it is a hypothesis.

5. Verify DBCC CHECKDB against the restored copy

-- Run integrity check against the restored database
DBCC CHECKDB ('MyDb_RestoreTest') WITH NO_INFOMSGS, ALL_ERRORMSGS;

CHECKDB against the restored copy is the only way to detect corruption that lived undetected in the source. VERIFYONLY cannot find it. A clean CHECKDB on the restored database is the strongest evidence the backup is usable.

6. Back up and track certificates and keys

-- Back up the TDE / backup-encryption certificate with its private key
BACKUP CERTIFICATE MyTdeCert
TO FILE = N'/securekeys/MyTdeCert.cer'
WITH PRIVATE KEY (
    FILE = N'/securekeys/MyTdeCert.pvk',
    ENCRYPTION BY PASSWORD = '<strong-password>'
);

Store certificate and key backups in a separate location from the database backups, with their own access controls. If the database backup and the certificate that decrypts it sit on the same array, a single failure takes both.

-- Certificate expiry (certificates expiring within 90 days)
SELECT name, subject, expiry_date,
       DATEDIFF(DAY, GETDATE(), expiry_date) AS days_until_expiry
FROM sys.certificates
WHERE expiry_date < DATEADD(MONTH, 3, GETDATE())
ORDER BY expiry_date;

-- TDE encryption state per database
SELECT d.name, dek.encryption_state,
    CASE dek.encryption_state
        WHEN 1 THEN 'Unencrypted'
        WHEN 2 THEN 'Encryption in progress'
        WHEN 3 THEN 'Encrypted'
        WHEN 4 THEN 'Key change in progress'
        WHEN 5 THEN 'Decryption in progress'
        WHEN 6 THEN 'Protection change in progress'
    END AS state_desc
FROM sys.dm_database_encryption_keys dek
JOIN sys.databases d ON dek.database_id = d.database_id;

After TDE certificate rotation, the old certificate is still required to restore log backups taken while it was active. Dropping the old certificate before every dependent backup has aged out of retention breaks the restore chain silently.

Common pitfalls

  • Treating backup job success as proof of recoverability. Always corroborate with msdb.dbo.backupset and periodic restore tests.
  • Skipping WITH CHECKSUM. Without checksums at backup time, VERIFYONLY cannot detect page-level corruption. Make CHECKSUM the default on production backup jobs.
  • Losing the certificate. TDE and backup encryption depend on a certificate and private key that are not inside the backup file. Lose the cert, lose the backup. Back it up separately and store it off-array.
  • Dropping rotated certificates too early. Older log backups still in the retention window need the certificate under which they were taken.
  • Taking regular full backups for restore testing. Use COPY_ONLY to avoid resetting the differential base and forcing a rebuild on the next production differential.
  • Switching recovery models to SIMPLE and back. This breaks the log chain. A full or differential backup is required to re-anchor it, and during the gap point-in-time recovery is impossible.
  • Ignoring restore duration. A backup that takes 20 minutes to write but 6 hours to restore does not meet a 2-hour RTO. Measure end-to-end.
  • Untested restore targets. The test target instance must be the right version, edition, and patch level, with enough storage. Discovering at restore time that the target is one version too old is the worst possible moment to learn that.
  • Backup files on the same volume as live databases. A volume failure takes both the database and the backup that was supposed to recover it.

Signals to monitor

SignalWhy it mattersWarning sign
Hours since last full backup per databaseDirect measure of recovery point exposureProduction database more than 24h without a full; more than 7d is critical
log_reuse_wait_desc per databaseTells you why the log cannot truncate; chain breaks surface hereACTIVE_TRANSACTION, REPLICATION (when not configured), or AVAILABILITY_REPLICA persisting. Note: LOG_BACKUP is normal between log backups on FULL recovery databases
Certificate days until expiryTDE and backup encryption break silently on expiryLess than 90 days warrants rotation planning
TDE encryption_stateStuck states indicate stalled encryption workProlonged state 2, 4, 5, or 6
Restore-test date per databaseUntested backups are unprovenMore than 30 days for tier-1, more than 90 days for any production database
Measured restore duration per databaseEmpirical RTO; the only real numberDuration exceeds SLA RTO
msdb.dbo.backupset row per expected backupAuthoritative record that a backup actually existsExpected backup type missing for a database
DBCC CHECKDB result on restored copyCatches corruption VERIFYONLY cannotAny error output

How Netdata helps

  • Backup-freshness tracking from msdb.dbo.backupset surfaces hours-since-last-full and hours-since-last-log per database, so a job that “succeeds” without writing a backupset row is visible.
  • Per-database transaction-log usage with log_reuse_wait_desc exposes a broken or stalled log chain as a non-LOG_BACKUP, non-NOTHING value long before error 9002.
  • Certificate expiry tracking puts TDE and backup-encryption certificates on a clock, so rotation happens before the cert blocks a restore.
  • Database-state monitoring catches databases stuck in RESTORING or RECOVERY_PENDING after a test or a real recovery.
  • Correlating restore-test windows with I/O stall, CPU, and TempDB metrics on the test target reveals whether measured RTO is dominated by storage, by the recovery pass, or by something else.
  • Disk-space monitoring on backup and restore-test volumes catches the common failure where a restore cannot complete because the target has no room.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.