The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-ag-failover-readiness

Operations Guides

SQL Server AlwaysOn failover readiness: quorum, health checks, and the failover you assume works

An AG that reports “healthy” in the dashboard can still fail to failover when you need it. The synchronization state tells you data is flowing between replicas. It says nothing about whether the cluster can orchestrate a failover, whether the health detection policy will catch the specific failure you are about to have, or whether the cluster has already exhausted its automatic failover budget for the period.

Automatic failover is the intersection of cluster quorum, replica synchronization state, health detection policy, the failure-condition-level setting, and the WSFC failover threshold. Break any one and the AG sits in RESOLVING during the incident you built it to survive.

This article covers the conditions automatic failover actually requires, the cluster and health-check settings that silently disable it, and the periodic validation checks that catch a broken failover path before the incident does.

What automatic failover actually requires

Automatic failover is not triggered by the AG itself. The cluster layer (WSFC on Windows, Pacemaker on Linux) triggers it after SQL Server’s health detection reports a failure condition. For the failover to execute, every one of the following must be true at the moment of failure.

flowchart TD
    A[Primary failure detected] --> B{Cluster quorum present?}
    B -- No --> X[AG goes offline. No failover.]
    B -- Yes --> C{Failover threshold exceeded?}
    C -- Yes --> Y[AG stuck FAILED. Manual reset required.]
    C -- No --> D{Auto-failover target exists?}
    D -- No --> Z[Manual failover only.]
    D -- Yes --> E{Secondary SYNCHRONIZED?}
    E -- No --> W[No auto-failover. Potential data loss.]
    E -- Yes --> F{failure-condition-level met?}
    F -- No --> V[No failover triggered.]
    F -- Yes --> G[Automatic failover executes]

Synchronous-commit on both replicas. The primary and the failover target must both be configured for synchronous-commit (availability_mode = SYNCHRONOUS_COMMIT). Asynchronous-commit replicas are never automatic failover targets.

Automatic failover mode on both replicas. The failover target must have failover_mode = AUTOMATIC. A synchronous replica set to manual failover provides zero-data-loss protection but will not failover without human action.

Secondary in SYNCHRONIZED state. The secondary must have hardened all log records up to the current primary LSN. A secondary in SYNCHRONIZING is catching up and is not a valid automatic failover target.

Cluster quorum present. The cluster must have quorum. If the cluster loses quorum, the AG goes offline on every node regardless of replica health. This is the most commonly overlooked requirement.

Failure-condition-level met. SQL Server’s health detection (via sp_server_diagnostics) must report a condition matching the configured failure-condition-level. The default level does not detect database-level problems.

Within the WSFC failover threshold. The cluster must not have exhausted its automatic failover budget for the period. See the section on the failover threshold below.

Quorum and vote configuration

Quorum is the cluster’s agreement on which node should own the AG resource. Without quorum, the cluster cannot bring the AG online on any node, and no failover of any kind occurs.

Witness requirement for even-numbered clusters. A two-node cluster without a witness cannot sustain a node crash. If either node goes down unexpectedly, quorum is lost and the AG goes offline on both nodes. Any even-numbered cluster requires a witness, either a file share witness or a cloud witness, to break the tie. Cloud witness (available on Windows Server 2016 and later) is the recommended option for Azure VMs and multi-site deployments.

Dynamic quorum is not a substitute for a witness. Dynamic quorum and dynamic witness adjust vote weights as nodes leave and join the cluster. They protect against sequential graceful node shutdowns. They do not protect against simultaneous failures. If two of three nodes crash at once, the remaining node plus witness hold half the original vote count, and quorum is lost.

Vote configuration matters in multi-site deployments. The primary replica node and at least one automatic failover target should hold a quorum vote. The Always On Availability Group Wizard warns when a node that could become primary lacks a vote. In single-site deployments this warning can be ignored. In multi-site deployments it is critical.

Check the current quorum configuration:

# WSFC: cluster quorum and node votes
Get-ClusterQuorum
Get-ClusterNode | Select-Object Name, State, NodeWeight
# Pacemaker: cluster status and quorum
pcs status
corosync-quorumtool

On Linux with Pacemaker, quorum is managed by corosync. The concepts (vote count, witness requirement for even-numbered clusters) apply, but the tooling differs entirely.

Health check and timeout settings

SQL Server reports its health to the cluster via sp_server_diagnostics. The cluster uses this input, along with the lease mechanism, to decide when the primary is unhealthy and a failover should be triggered. Three settings control this behavior.

Failure-condition-level (default: 3, OnCriticalServerError). Controls which conditions trigger an automatic failover. Level 3 triggers on critical server errors such as serious internal errors. It does not trigger on database-level problems. A database going SUSPECT, a transaction log filling up, or a data file going missing will not trigger an automatic failover at the default level. Levels range from 1 (least sensitive, server down only) through 5 (most sensitive, includes any qualifying error condition).

Health-check-timeout (default: 30000 ms). How long the cluster waits for sp_server_diagnostics to respond before considering the primary unhealthy. The sp_server_diagnostics polling interval is one-third of this value, which is 10 seconds at the default. The documented minimum is 15,000 ms (15 seconds), which sets the polling interval to 5 seconds.

Lease timeout (default: 20000 ms). A separate mechanism where the SQL Server resource DLL and the SQL Server process exchange lease renewals. If either side stops responding, the lease expires and the cluster considers the resource failed. The lease timeout must be less than 2 * SameSubnetThreshold * SameSubnetDelay, the maximum-lease equation Microsoft documents, to prevent split-brain: if the subnet detection window expires before the lease, the cluster could bring the resource online on a second node while the original primary still holds the lease. At the default 20000 ms timeout, leases renew every 10 seconds (half the timeout). Lease timeouts can be triggered by resource pressure such as high CPU, low memory, or disk latency, not just SQL Server process crashes. Microsoft’s troubleshooting documentation confirms the default lease timeout is 20 seconds across SQL Server 2016-2022.

Check the current AG-level health detection settings:

SELECT
    name AS ag_name,
    failure_condition_level,
    health_check_timeout,
    db_failover
FROM sys.availability_groups;

Database-level health detection (DB_FAILOVER)

The default failure-condition-level operates at the server instance level. At the default level, sp_server_diagnostics reports server health but does not trigger failover for individual database failures. This is by design.

The consequence is significant. If a single database in the AG goes offline due to corruption, a full transaction log (error 9002), a deleted data file, or any other database-specific failure, the AG will not automatically failover. The primary instance is still healthy from the cluster’s perspective, so the health detection never fires.

DB_FAILOVER is a separate AG-level option that enables database-level health detection. When enabled, the AG treats any database transitioning to an unhealthy state as a failover condition.

DB_FAILOVER is OFF by default. Enable it per AG:

-- Warning: changes failover behavior. A single database going offline
-- will trigger an AG-level failover after this change.
ALTER AVAILABILITY GROUP [YourAG] SET (DB_FAILOVER = ON);

For production AGs where a single database failure should trigger failover, this should be ON. Evaluate the tradeoff before enabling it. If your AG carries databases with different criticality levels, DB_FAILOVER will failover the entire AG when any one database has a problem, which may not be what you want for a mixed-workload AG.

The WSFC failover threshold trap

WSFC tracks automatic failovers per resource over a rolling time window. The default policy is “Maximum Failures in Specified Period,” which allows N-1 failures, where N is the number of cluster nodes, within a 6-hour window.

Failover Cluster Manager shows these same defaults on the Failover tab of an availability group resource’s properties: “Maximum failures in the specified period” is n-1 and “Period” is 6 hours. Microsoft documents these as the default WSFC failure-threshold values and exposes them for change in the Failover Manager Console.

Once this threshold is exceeded, the AG resource enters a FAILED state and the cluster stops attempting automatic failovers entirely. No further automatic failover will occur until a human intervenes. The AG stays in RESOLVING or FAILED.

This is a silent failover blocker that operators frequently misdiagnose. The AG looks broken, the replicas look healthy, but the cluster has given up on automatic failover for this resource. Check the cluster log:

# Generate cluster log and search for failover threshold events
Get-ClusterLog -Node <NodeName> -TimeSpan 60
# Then search the cluster.log for: "failoverCount", "IsAlive", "MaximumFailures"

The immediate fix is to bring the AG resource back online manually, either in Failover Cluster Manager or via PowerShell, which resets the failover count. But the real fix is to understand why the AG failed over repeatedly in the first place. A flapping AG that keeps failing over has an underlying problem, and the threshold is masking it.

Periodic failover readiness validation

Failover readiness decays silently. Cluster configuration drifts, nodes lose votes, certificates expire, and the failover path you validated at deployment may not exist six months later. Run the following checks on a regular schedule.

Replica configuration and health state:

SELECT
    ag.name AS ag_name,
    ar.replica_server_name,
    ar.availability_mode_desc,
    ar.failover_mode_desc,
    ag.failure_condition_level,
    ag.health_check_timeout,
    ag.db_failover,
    ars.role_desc,
    ars.connected_state_desc,
    ars.synchronization_health_desc
FROM sys.availability_groups ag
JOIN sys.availability_replicas ar ON ag.group_id = ar.group_id
LEFT JOIN sys.dm_hadr_availability_replica_states ars ON ar.replica_id = ars.replica_id
ORDER BY ag.name, ar.replica_server_name;

Secondary synchronization and queue depth:

SELECT
    DB_NAME(drs.database_id) AS database_name,
    ar.replica_server_name,
    drs.is_local,
    drs.synchronization_state_desc,
    drs.redo_queue_size,
    drs.log_send_queue_size
FROM sys.dm_hadr_database_replica_states drs
JOIN sys.availability_replicas ar ON drs.replica_id = ar.replica_id
ORDER BY DB_NAME(drs.database_id), drs.is_local DESC;

NT AUTHORITY\SYSTEM permissions on each replica (required for health detection):

SELECT permission_name, state_desc
FROM sys.server_permissions sp
JOIN sys.server_principals p ON sp.grantee_principal_id = p.principal_id
WHERE p.name = 'NT AUTHORITY\SYSTEM'
AND sp.permission_name IN ('CONNECT SQL', 'VIEW SERVER STATE', 'ALTER ANY AVAILABILITY GROUP');

Certificate expiry on AG endpoints:

SELECT name, start_date, expiry_date
FROM sys.certificates
WHERE expiry_date < DATEADD(day, 90, GETUTCDATE());

Additional checks to perform on schedule:

Quorum state. Verify the cluster has quorum and the witness is online. A witness that has been offline for weeks is a silent failover blocker that AG-level monitoring will never surface.

AG endpoint connectivity. Verify that the database mirroring endpoints on all replicas can connect to each other. Certificate expiry on endpoint authentication causes silent disconnection between replicas.

Failover threshold state. Confirm the AG resource is not in a state where the cluster has stopped attempting automatic failovers due to threshold exhaustion.

Periodic failover testing. The only definitive way to validate that automatic failover works is to perform a planned failover during a maintenance window. Fail over to each automatic target, confirm the application reconnects through the listener, verify data integrity, and fail back. If you have never performed a failover on this AG in production, you do not actually know if it works.

Signals to monitor for failover readiness

SignalWhy it mattersWarning sign
synchronization_health_descWhether the secondary is a valid failover targetNot HEALTHY on a synchronous replica
connected_state_descSecondary must be connected to primary to be a failover targetDISCONNECTED on a synchronous replica
Redo queue sizeDetermines RTO on failover (new primary applies pending log)Growing queue or estimated catch-up exceeds RTO target
Send queue sizePrimary accumulating un-replicated log means data loss on failoverGrowing queue on a synchronous replica
WSFC cluster quorum stateWithout quorum, no failover of any kind occursQuorum loss or witness offline
NT AUTHORITY\SYSTEM permissionsRequired for health detection on every potential primaryMissing any of the three required permissions on any replica
Certificate expiry on AG endpointsExpired certificates disconnect replicas silentlyExpiry within 90 days
HADR_SYNC_COMMIT wait timeLatency cost of synchronous replication on primary commitsSudden increase indicates secondary or network degradation

How Netdata helps

  • AG replica state as a continuous time series. Per-second sampling of synchronization health and connection state catches the transition from HEALTHY to PARTIALLY_HEALTHY before the incident, not after.
  • Redo queue correlated with secondary resource pressure. A growing redo queue paired with elevated CPU or I/O stall on the secondary tells you why the secondary cannot keep up, not just that it is falling behind.
  • HADR_SYNC_COMMIT waits alongside secondary I/O metrics. Correlating synchronous commit latency with secondary log write latency and network throughput separates network problems from secondary storage problems.
  • Cluster quorum alongside AG health. Monitoring both layers catches quorum loss, the most common silent failover blocker, that AG-only monitoring misses entirely.
  • Per-second granularity during failover events. When a failover occurs, per-second metrics show when the primary stopped responding, when the secondary took over, and how long the application was without a writable endpoint.

The Netdata SQL Server collector provides per-second visibility into these signals.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.