The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-monitoring-checklist

Operations Guides

Microsoft SQL Server monitoring checklist: the signals every production instance needs

Most SQL Server outages are not exotic. The transaction log fills because a backup job silently stopped. A sleeping session with an open transaction blocks forty other sessions until the worker pool runs dry. TempDB runs out of space and every database on the instance stalls at once. All of these are visible hours or days in advance if you collect the right signals. Most teams do not.

This checklist is the minimum set of signals a production SQL Server instance needs, organized so you can audit what you have today and fill the gaps. It targets standalone instances, failover cluster instances, and AlwaysOn AG deployments on-premises or on VMs. Azure SQL Database and Managed Instance share many of the same DMVs but abstract storage and resource governance differently.

Use it in two passes. First, confirm the survival baseline: can you tell, right now, whether the instance is alive, databases are online, logs are not filling, and backups exist? Second, work through the operational signals that catch degradation before it becomes an outage.

Survival baseline: is the instance alive and recoverable

These are binary checks. If any of them fail, you have an incident or a recovery exposure, not a tuning problem.

  • Instance responsiveness: run a real query, not just a TCP connect, against port 1433. A successful connect with a hanging query means the engine is alive but saturated (worker thread exhaustion or severe blocking), which is a different failure than a dead service. sqlcmd -S localhost -Q "SELECT 1" -l 10 is enough. Page on sustained failures (3 or more consecutive probe failures over 60 seconds or more); single failures are often transient.
  • Database state: every production database in sys.databases should be ONLINE. SUSPECT or RECOVERY_PENDING means corruption or failed recovery and is a page. RESTORING or OFFLINE that you did not cause is a ticket. Large databases legitimately spend time in recovery after a restart; track your normal recovery durations so you can tell slow from stuck.
  • Transaction log percent used: alert above 70% used with a non-NOTHING log_reuse_wait_desc, and page above 90% when the log is still rising and there is no autogrow or volume headroom left. Log full (error 9002) halts all writes to that database, and if it is TempDB’s log, it halts the instance.
  • Log reuse wait reason: collect log_reuse_wait_desc alongside the percentage. LOG_BACKUP means backups are not running. ACTIVE_TRANSACTION means a long transaction pins the log. REPLICATION or AVAILABILITY_REPLICA means a downstream consumer is behind. Adding disk space fixes none of these.
  • Disk free on data, log, and TempDB volumes: alert below 20% free. Page only when exhaustion is operationally imminent, meaning free bytes are less than the next configured growth increment plus margin, or autogrow has already failed. A flat “10% free” threshold pages on huge volumes with terabytes left and misses tiny volumes about to die.
  • Backup freshness: for every online user database, track hours since the last full backup and last log backup from msdb.dbo.backupset. No full backup in over 24 hours for production is a ticket; no log backup in over an hour on a full recovery model database is both a recovery exposure and a log-full outage in waiting. Job success status is not proof; check the backupset records, because jobs can “succeed” while writing to a bad target.
  • Critical error log entries: errors 823, 824, and 825 mean storage is failing. Error 825 is the insidious one: the read eventually succeeded, so nothing broke, but the medium is deteriorating. Most teams alert on 823 and 824 and miss 825, which is often the only warning you get.

Memory and buffer pool

Memory pressure in SQL Server is a slow spiral: the buffer pool shrinks, PLE drops, physical reads climb, I/O saturates, latency rises, concurrency piles up. Catching it at step one is cheap; catching it at step four is an incident.

  • Page Life Expectancy (PLE): the seconds a page survives in the buffer pool. The old “300 seconds” rule was calibrated for small 32-bit systems; a widely used community heuristic is (buffer_pool_GB / 4) * 300 seconds, but the trend matters more than the absolute value. A 50% drop from baseline is worth investigating regardless of the number. PLE is instantaneous, not cumulative, so brief dips during checkpoints or large scans are normal.
  • Per-NUMA-node PLE: on multi-NUMA systems the instance-wide Buffer Manager value is an average across nodes, not a minimum. One starved node hides behind a healthy aggregate. Collect PLE per node from the Buffer Node counters.
  • Buffer cache hit ratio: above 99% for OLTP, below 95% investigate, below 90% with elevated PAGEIOLATCH_* waits is serious. Data warehouses legitimately run lower; 80-95% can be normal there. Do not page on this alone, since cold starts, backups, and DBCC all cause legitimate drops. Note the trap: a 99% hit ratio can coexist with a collapsing PLE if a small working set is being churned rapidly. PLE is the more sensitive signal.
  • Memory grants pending: should be zero. Any sustained nonzero value means queries are parsed, optimized, and queued waiting for a memory grant, appearing hung to the application while CPU and I/O look fine. This is one of the most commonly missed signals in SQL Server monitoring. Correlate with RESOURCE_SEMAPHORE waits and sys.dm_exec_query_memory_grants to find the hoarder.
  • Memory state flags: system_memory_state_desc in sys.dm_os_sys_memory and process_physical_memory_low in sys.dm_os_process_memory tell you when the OS is pressuring SQL Server to shrink. On a dedicated host, available physical memory below 500 MB suggests max server memory is set too high.

CPU, schedulers, and worker threads

SQL Server runs its own cooperative schedulers, one per logical CPU. OS CPU percentage is a weak proxy for engine CPU pressure, and worker thread exhaustion is invisible in OS metrics entirely.

  • CPU utilization, split SQL vs other: SQL CPU high with other CPU low means the engine is the bottleneck. SQL low with other high means something else on the host competes (antivirus, backup agents). Both low with slow queries means the bottleneck is not CPU at all; go look at waits. CPU percentage alone should never page; page only as a composite with runnable backlog and user-visible timeouts.
  • Runnable tasks per scheduler: runnable_tasks_count in sys.dm_os_schedulers (filtered to status = 'VISIBLE ONLINE') is the cleanest in-engine CPU queue signal. Sustained values above 1 per scheduler mean CPU pressure; above 5, significant contention. On VMs this can be elevated while the host reports headroom, because steal time is invisible to the guest.
  • Signal wait ratio: signal_wait_time_ms as a share of total wait_time_ms in sys.dm_os_wait_stats. Above roughly 20% means threads get their resource and then wait for a CPU to run on. A ratio rising week over week is CPU headroom quietly disappearing.
  • Worker thread utilization: track active workers against max_workers_count from sys.dm_os_sys_info. Investigate at 80% of max; 90% is imminent exhaustion. Any sustained THREADPOOL wait with work_queue_count > 0 means the instance is actively refusing work, which is a page. Do not just raise max worker threads; find what is consuming them, usually a blocking cascade or a parallel query explosion.
  • Compilations per second: the ratio of SQL Compilations/sec to Batch Requests/sec should stay under 10%. Above that, CPU is being burned on plan compilation, usually from non-parameterized ad-hoc SQL or plan cache eviction under memory pressure. These are cumulative counters; compute deltas between samples.

Waits, blocking, and deadlocks

Wait statistics are the single most diagnostic signal SQL Server exposes, and the most commonly wasted one because sys.dm_os_wait_stats is cumulative since startup. A single point query shows the entire lifetime profile of the instance, which is useless for identifying what is wrong now.

  • Delta-sampled wait stats: snapshot every 30-60 seconds, compute deltas, and exclude the known benign idle waits (SLEEP_TASK, LAZYWRITER_SLEEP, XE_TIMER_EVENT, REQUEST_FOR_DEADLOCK_SEARCH, and similar), or background noise will dominate. Any single wait type above 30-40% of total waits that correlates with user-visible latency deserves investigation.
  • Specific high-signal wait types: PAGEIOLATCH_* means storage or buffer pool pressure. LCK_M_* means blocking. WRITELOG means log flush latency. RESOURCE_SEMAPHORE means memory grant queuing. THREADPOOL means worker exhaustion. HADR_SYNC_COMMIT means a synchronous AG secondary is slowing your commits. CXPACKET is routinely the top wait on healthy systems and is expected in isolation; after the SQL 2016 SP2 split, CXCONSUMER carries the typically benign consumer side.
  • Blocking chain depth and head blocker state: poll sys.dm_exec_requests for blocking_session_id <> 0 every 30 seconds. The dangerous shape is a sleeping head blocker (no active request) with an uncommitted transaction: it will not resolve on its own, every blocked session holds a worker thread, and the cascade ends in THREADPOOL waits. Ticket on chains deeper than 5 sessions or blocks older than 60 seconds; page when a sleeping head blocker has blocked 10 or more sessions for 5 or more minutes with throughput impact. When you kill the blocker, rollback can take as long as the original transaction, so do not expect instant relief.
  • Deadlocks per second: Number of Deadlocks/sec from the Locks counters, with deadlock graphs pulled from the system_health Extended Events ring buffer. Occasional deadlocks that applications retry are a ticket; a sustained storm above roughly 10 per minute is a page. The ring buffer has limited retention, so capture graphs externally.

TempDB

TempDB is shared by every database on the instance. Space exhaustion halts queries instance-wide, and allocation contention throttles throughput without a single error message.

  • TempDB free space: ticket below 20% free, urgent below 10%. Page only on genuine exhaustion: free space under 5%, still falling across samples, autogrow unavailable or maxed, and query failures beginning. TempDB is recreated at its initial size on every restart, so growth after a restart is expected, not a leak.
  • Space by consumer: from sys.dm_db_file_space_usage (database-scoped, so run it in the TempDB context), split user objects (temp tables), internal objects (sort and hash spills), and version store (RCSI, snapshot isolation, readable secondaries). Version store above 50% of TempDB means find the long-running snapshot transaction. Large internal objects mean queries are spilling, which points back at memory grants.
  • Allocation contention: PAGELATCH_UP/PAGELATCH_EX waits on pages in database ID 2 (PFS, GAM, SGAM allocation pages) visible in sys.dm_os_waiting_tasks where resource_description LIKE '2:%'. The standard mitigation is one TempDB data file per logical CPU up to 8, equally sized. Do not page on contention alone; ticket it.

I/O latency per file

sys.dm_io_virtual_file_stats is SQL Server’s own measurement of storage latency as the engine experiences it, which OS disk metrics cannot tell you. It is cumulative since startup, so snapshot and compute deltas.

File typeExcellentAcceptableDegradedSevere
Data file reads< 10 ms10-20 ms> 20 ms> 50 ms
Log file writes< 2 ms2-5 ms> 5 ms> 15 ms

Log write latency matters most because every commit waits on it. Sustained log write latency above 20 ms is a page; it directly taxes every write transaction and every synchronous AG commit. Identify log files by joining sys.master_files on type_desc = 'LOG', not by assuming a file_id. On SSD or NVMe, anything consistently above 5 ms points at the storage layer, a throttled cloud disk tier, or an IOPS cap.

Availability Groups, if configured

Skip this section for standalone instances. If you run AGs, “synchronization_health says HEALTHY” is not sufficient monitoring.

  • Replica state and sync health: sys.dm_hadr_availability_replica_states should show CONNECTED plus HEALTHY on synchronous replicas. Page on DISCONNECTED or NOT_HEALTHY on a synchronous replica sustained past the failover transition window (120 seconds or more) with no other healthy synchronous target, because failover capability and data durability are both compromised.
  • Send queue: log generated on the primary but not yet sent. For synchronous replicas this should be near zero; sustained growth means the network or the secondary cannot keep up, and with HADR_SYNC_COMMIT rising it means primary commits are blocking. That is a page.
  • Redo queue: log received but not yet applied on the secondary. Divide redo_queue_size by redo_rate to get catch-up time, and compare that to your failover RTO. A secondary that would need an hour of redo after failover is a recovery time liability even while everything reports healthy. Note: sys.dm_hadr_database_replica_states has no database_name column; use DB_NAME(drs.database_id) or join sys.databases.

Throughput and connections

  • Batch requests/sec: the best single workload volume indicator. Baseline it by time of day and day of week, then alert on sustained deviation beyond 2x or below 0.5x of baseline. A drop with high connections and rising waits means SQL Server cannot complete work; a drop with low connections and low waits means the problem is upstream. The counter in sys.dm_os_performance_counters is cumulative (cntr_type = 272696576); reading cntr_value directly gives you a meaningless large number.
  • User connections: useful as a correlate, not a standalone trigger. Connections climbing without a matching rise in batch requests is a pool leak or retry storm. The dangerous relationship is connection and active request count against worker threads, not any absolute number.
  • Autogrow events: every autogrow pauses I/O to the growing file while the new space initializes, and Instant File Initialization does not apply to log files, so log growth zero-fills the whole extent. Any autogrow on a production database during business hours is a ticket: files should be pre-sized, and the event itself is a latency stall someone should know about. Pull events from the default trace or Extended Events.

Collection gotchas that will bite you

  • Cumulative counters everywhere: wait stats, I/O stall, batch requests, deadlocks, compilations, and log counters are all cumulative since startup. Everything above assumes periodic snapshots with deltas. If your tooling reads raw cntr_value, your graphs are fiction.
  • SQL Server 2022 permission change: on SQL Server 2022 and later, many performance DMVs including sys.dm_os_performance_counters require the new VIEW SERVER PERFORMANCE STATE permission instead of VIEW SERVER STATE. A monitoring account that worked on 2019 will silently return nothing after an upgrade until you grant it.
  • DMV data dies with the instance: wait stats, query stats, and I/O stats all reset on restart. If you rely on DMVs alone, every post-mortem after a crash has no forensic data. Persist samples externally.
  • sys.dm_os_performance_counters can return zero rows if performance counters are disabled on the instance. Check your collector actually gets data, not just that it runs without error.

How Netdata helps

Netdata’s SQL Server collector samples the engine directly, which maps well onto this checklist:

  • Per-second collection of connections, buffer cache hit ratio, and throughput counters, with the delta math on cumulative counters handled for you.
  • Wait statistics collected as a time series, so you see the current wait profile rather than a cumulative-since-startup smear.
  • Database state, transaction log usage, and active transaction visibility per database, so log-full trajectories show up before error 9002.
  • Blocking chain detection that surfaces the head blocker and how many sessions are stuck behind it, which is exactly the signal that precedes worker thread exhaustion.
  • Correlation on one dashboard between engine-internal signals (waits, PLE, runnable tasks) and host-level signals (CPU, disk latency, disk free), which is how you tell a memory pressure spiral apart from a storage failure in minutes instead of hours.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.