The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-tde-certificate-expiry

Operations Guides

SQL Server TDE and endpoint certificate expiry: silent AG and backup failures

SQL Server does not alert when a certificate is about to expire. There is no performance counter, no error log entry at expiry time, and no DMV flag that flips the moment the date passes. The engine keeps using the certificate silently until something forces a re-evaluation, and at that point the failure mode depends entirely on what the certificate was protecting.

Three consumers of the same self-signed certificate pattern behave in three different ways:

  • TDE protector certificates are not enforcement-checked at expiry, so an expired TDE cert keeps the database working.
  • Backup encryption certificates fail the next backup immediately with Msg 3096 / 3013.
  • AG / database mirroring endpoint certificates keep working until SQL Server restarts or the endpoint is stopped and started, at which point the endpoint refuses to authenticate and every secondary disconnects.

None of these are surfaced as runtime metrics. All of them are detected by an explicit query against sys.certificates.

The dangerous scenario is not the expiry itself but the cascade that follows a restart, a failover, or a backup rotation that happens to land in the expiry window. AG endpoint certificate expiry has been observed to take down distributed AGs after a routine Windows patching cycle, with error log messages that point at networking rather than certificates. Backup encryption expiry silently breaks the restore chain. TDE cert expiry is operationally silent but becomes a problem during key rotation or restore.

What this means

Certificates in SQL Server back three independent subsystems:

  • TDE protector certificates encrypt the database encryption key (DEK) for databases with TDE enabled. sys.dm_database_encryption_keys.encryption_state reflects whether encryption is in progress (2), encrypted (3), key change in progress (4), decryption in progress (5), or protection change in progress (6).
  • AG / database mirroring endpoint certificates authenticate the HADR endpoint, common on domain-less or cross-domain setups where Windows Negotiate auth is not available.
  • Backup encryption certificates encrypt the backup media when WITH ENCRYPTION is specified on BACKUP DATABASE.

The default expiry window for self-signed certificates created without EXPIRY_DATE is one year from creation: Microsoft’s CREATE CERTIFICATE documentation sets the default expiry to one year after START_DATE, which itself defaults to the creation date. Most operators hit their first rotation cycle inside the first 12 months after deployment, and most have not built alerting for it because the engine does not surface it.

A single expired certificate can manifest in multiple subsystems at once if it is overloaded, but in practice the failure surfaces wherever a consumer first re-evaluates:

  • New encrypted backups fail immediately (Msg 3096, Msg 3013 terminating the BACKUP).
  • AG endpoints keep working until the next endpoint stop/start or SQL Server restart.
  • TDE keeps working indefinitely for encrypt/decrypt, but key rotation is blocked and restores that need the original cert may fail.
flowchart TD
    Cert[SQL Server certificate expires]
    Cert --> TDE[TDE protector: silent, no enforcement]
    Cert --> AG[AG endpoint: silent until restart or endpoint stop/start]
    Cert --> BU[Backup encryption: fails immediately on next backup]
    TDE --> TDERestore[Restores needing original cert fail if cert dropped]
    AG --> AGDown[All secondaries disconnect after restart]
    BU --> BUDown[Backup jobs fail with Msg 3096 + 3013]

Common causes

CauseWhat it looks likeFirst thing to check
TDE protector certificate expiredDatabase still online; encryption_state still 3; no errors; key rotation blockedsys.certificates.expiry_date joined to DEK via thumbprint
AG endpoint certificate expiredReplicas worked until restart or Windows patching; secondaries now DISCONNECTED; error log mentions “connection timeout” or “availability replica”sys.database_mirroring_endpoints joined to sys.certificates on certificate_id
Backup encryption certificate expiredBACKUP DATABASE ... WITH ENCRYPTION fails with Msg 3096 / 3013msdb.dbo.backupset.encryptor_type + sys.certificates.expiry_date
Old TDE cert dropped after rotationRestore of older backup fails with “Cannot find server certificate”sys.certificates for the missing thumbprint; check old cert backup exists
Endpoint certificate-to-login mapping lost during rotationEndpoint auth fails after cert replacement even with new cert in placeEndpoint certificate_id and login-to-certificate mapping

Quick checks

Run these in the master database context. All are read-only.

-- Check 1: List certificates expiring within 90 days, ordered by expiry
SELECT name, subject, start_date, expiry_date,
       DATEDIFF(DAY, GETDATE(), expiry_date) AS days_until_expiry
FROM sys.certificates
WHERE expiry_date < DATEADD(DAY, 90, GETDATE())
ORDER BY expiry_date;
-- Check 2: Map TDE protector certificates to databases and their encryption state
SELECT d.name AS database_name,
       c.name AS certificate_name,
       c.expiry_date,
       dek.encryption_state,
       CASE dek.encryption_state
            WHEN 0 THEN 'No database encryption key'
            WHEN 1 THEN 'Unencrypted'
            WHEN 2 THEN 'Encryption in progress'
            WHEN 3 THEN 'Encrypted'
            WHEN 4 THEN 'Key change in progress'
            WHEN 5 THEN 'Decryption in progress'
            WHEN 6 THEN 'Protection change in progress'
       END AS encryption_state_desc
FROM sys.dm_database_encryption_keys dek
JOIN sys.databases d ON dek.database_id = d.database_id
LEFT JOIN sys.certificates c ON dek.encryptor_thumbprint = c.thumbprint;
-- Check 3: AG / database mirroring endpoint authentication method and bound cert
SELECT name, state_desc, protocol_desc,
       connection_auth_desc, certificate_id
FROM sys.database_mirroring_endpoints;
-- If connection_auth_desc contains CERTIFICATE, join certificate_id to sys.certificates
-- Check 4: Join endpoint certificate to its expiry
SELECT dme.name AS endpoint_name, dme.connection_auth_desc,
       c.name AS certificate_name, c.expiry_date
FROM sys.database_mirroring_endpoints dme
LEFT JOIN sys.certificates c ON dme.certificate_id = c.certificate_id;
-- Check 5: Recent backup encryption usage and which certificate thumbprint was used
SELECT database_name, encryptor_type, encryptor_thumbprint,
       backup_start_date, type AS backup_type
FROM msdb.dbo.backupset
WHERE encryptor_type IS NOT NULL
ORDER BY backup_start_date DESC;
-- Check 6: AG replica state and connectedness (correlate with endpoint cert expiry)
SELECT ag.name AS ag_name,
       ar.replica_server_name,
       ars.role_desc,
       ars.connected_state_desc,
       ars.synchronization_health_desc,
       ars.last_connect_error_description
FROM sys.dm_hadr_availability_replica_states ars
JOIN sys.availability_replicas ar ON ars.replica_id = ar.replica_id
JOIN sys.availability_groups ag ON ar.group_id = ag.group_id;
-- Check 7: Stalled TDE encryption state (2/4/5/6 stuck beyond expected scan duration)
SELECT d.name AS database_name, dek.encryption_state,
       dek.encryption_scan_modify_date
FROM sys.dm_database_encryption_keys dek
JOIN sys.databases d ON dek.database_id = d.database_id
WHERE dek.encryption_state IN (2, 4, 5, 6);
-- Check 8: Error log for the misleading "availability replica" pattern after restart
EXEC sp_readerrorlog 0, 1, 'availability replica';

How to diagnose it

  1. Inventory every certificate and its consumer. Run checks 1 through 4. Confirm which certificates are bound to TDE, which to AG endpoints, and which to backup encryption. Treat each binding independently: a TDE-only certificate does not carry endpoint-style rotation urgency, and an endpoint certificate that is also a TDE protector is two problems at once.
  2. Confirm the failure mode with the error log. After a restart with an expired endpoint certificate, the error log typically reports connection timeouts to the availability replica, not “certificate expired”. This is misleading. If check 6 shows DISCONNECTED secondaries and check 4 shows an expired endpoint certificate, the certificate is the root cause even though the error message reads like a network problem.
  3. Distinguish expiry from other failure causes. A failed encrypted backup could be certificate expiry (Msg 3096), a missing certificate, or a permissions issue on the certificate private key. The Msg 3096 / 3013 pair specifically indicates expiry. Use check 5 to confirm which cert the backup job is referencing.
  4. Check whether old certificates still exist. If a TDE rotation was done recently and older backups fail to restore, the cause is the dropped original certificate. SQL Server needs every TDE protector that ever encrypted a log block you are trying to restore. KB4534430 documents a related case where dropping the original cert after rotation breaks log backups taken with COMPRESSION and MAXTRANSFERSIZE (it applies to SQL Server 2012, 2014, 2016, and 2019).
  5. Cross-check required_synchronized_secondaries_to_commit. If the AG has required_synchronized_secondaries_to_commit > 0 and the endpoint cert has expired, commits on the primary will block once secondaries disconnect after a restart.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
sys.certificates.expiry_date per certificatePrimary signal; SQL Server does not expose this as a metric< 90 days to expiry: ticket. < 7 days: page.
sys.dm_database_encryption_keys.encryption_state per databaseStuck 2/4/5/6 indicates a stalled TDE operation that may relate to a cert problemStays in 2/4/5/6 longer than the expected scan duration
sys.dm_hadr_availability_replica_states.connected_state_descDISCONNECTED secondaries after a restart with an expired endpoint certDISCONNECTED on synchronous replica sustained > 120 seconds
msdb.dbo.backupset.encryptor_thumbprintTies backups to specific certificates; flags when one cert is the only restore path for many backupsA single cert protecting more backups than your retention can survive losing
Error log entries referencing “availability replica” or “connection timeout” following a restartMost reliable in-engine indicator that endpoint auth is brokenCluster of these entries within minutes of instance startup
sys.database_mirroring_endpoints.connection_auth_desc containing CERTIFICATEFlags endpoints that will break on cert expiryCombined with cert expiry < 90 days

SQL Server does not surface certificate expiry as a runtime metric. Whatever monitoring platform you use, this needs an explicit query job that runs at least daily, persists the result, and alerts on a threshold rather than waiting for the engine to complain.

Fixes

The fix differs by consumer. Treat them separately even when they share a certificate.

Replace an expired TDE protector certificate

  1. Back up the existing certificate (private key included) before doing anything: BACKUP CERTIFICATE <old> TO FILE = ... WITH PRIVATE KEY .... Keep this backup for the lifetime of any backup that depends on it.
  2. Create or restore the new certificate on the primary and on every AG secondary that hosts the database.
  3. Rotate the DEK protector: ALTER DATABASE ENCRYPTION KEY ENCRYPTION BY SERVER CERTIFICATE <new>;.
  4. Wait for sys.dm_database_encryption_keys.encryption_state to return to 3 (Encrypted). Expect brief transitions through 4 (key change in progress) and 6 (protection change in progress).
  5. Do not drop the old certificate yet. Drop it only after confirming every restore path that depends on it has been migrated, and after taking at least one full backup under the new cert.

Replace an expired AG endpoint certificate

  1. Create the new certificate on the primary, back it up, and restore it on every secondary that authenticates via this endpoint.
  2. Create or confirm the corresponding login on each partner that maps to the certificate.
  3. Alter the endpoint to use the new certificate: ALTER ENDPOINT ... FOR DATA_MIRRORING AUTHENTICATION = CERTIFICATE [<new>] ....
  4. Stop and restart the endpoint with ALTER ENDPOINT ... STATE = STOPPED followed by STATE = STARTED. The endpoint will not pick up the new cert without this. This disconnects replicas briefly; do it inside a planned window.
  5. Verify sys.dm_hadr_availability_replica_states.connected_state_desc returns to CONNECTED on all replicas.

If the endpoint was configured with CERTIFICATE, NEGOTIATE and the certificate side has expired but Windows auth is available, a temporary mitigation is to drop certificate auth and rely on NEGOTIATE only, since NEGOTIATE uses Windows authentication and therefore only works where Windows auth is available. This is a workaround, not a fix. Bring certificate auth back once the new certificate is in place.

Replace an expired backup encryption certificate

  1. Create the new certificate, or restore it from a known-good backup.
  2. Update the backup job, maintenance plan, or Ola Hallengren job to reference the new certificate by name in the WITH ENCRYPTION (SERVER CERTIFICATE = ...) clause.
  3. Re-run the failed backup. Msg 3096 should disappear.
  4. Keep the old certificate installed in master until every backup encrypted with it has been overwritten or has reached end of retention. Restore of an existing encrypted backup still works with an expired certificate as long as the certificate is present: Microsoft documents that you can still use a certificate past its expiration date to encrypt and decrypt data, and restore needs the certificate and its private key, not a valid expiry date.

Prevention

  • Build an external expiry check. Schedule a daily job that runs check 1 and persists the output. Alert at 90 days out (ticket), 30 days (warning), and 7 days (page). Do not rely on the SQL Server engine to tell you.
  • Track certificates by consumer, not by name. A single expiry query without context will not tell you which certificates are AG-critical. Join to sys.dm_database_encryption_keys, sys.database_mirroring_endpoints, and msdb.dbo.backupset.encryptor_thumbprint.
  • Standardize on a longer EXPIRY_DATE for self-signed certs. The default one-year window is appropriate for some compliance regimes but operationally short. Pick a window that matches your patching and rotation cadence.
  • Keep every retired TDE protector until its backups age out of recovery requirements. Treat dropped TDE certs as equivalent to dropping the ability to restore.
  • Document restart-dependent failures in your runbooks. An endpoint cert that expired six months ago can sit silent through many failover drills until a restart exposes it. Add a pre-restart certificate check to the patching runbook.
  • Test AG endpoint auth rotation in non-production. If your AG endpoint uses certificate auth, validate the full rotation procedure (create, back up, restore, alter endpoint, stop/start endpoint) on a non-production AG at least once before doing it under pressure.

How Netdata helps

  • Netdata collects per-second SQL Server metrics including AG replica state (connected_state_desc, synchronization_health_desc), so a DISCONNECTED secondary after a restart shows up immediately and correlates with the restart event itself.
  • Netdata can run custom SQL queries against sys.certificates and sys.dm_database_encryption_keys on a schedule, persist the expiry window, and alert on the 90/30/7 day thresholds without depending on the engine to surface expiry as a counter.
  • Correlating a sudden AG disconnection with a recent SQL Server service restart is the fastest way to separate a network issue from an endpoint auth issue. The two failure modes look identical in the error log.
  • TDE encryption_state transitions (2/4/5/6) and their duration are visible as a state timeline alongside AG and backup freshness signals, which lets you see a stalled rotation as it happens rather than after the fact.
  • ML-based anomaly detection on the AG send queue, redo queue, and HADR_SYNC_COMMIT wait time flags the secondary-side impact of an endpoint auth failure before the operator has to read the error log.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.