The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-pageiolatch-waits

Operations Guides

SQL Server PAGEIOLATCH waits: the buffer pool waiting on slow storage

Your top wait type is PAGEIOLATCH_SH or PAGEIOLATCH_EX. Queries that used to be sub-50ms now take seconds. CPU may be low. There are no blocking chains. The buffer pool is waiting on disk.

PAGEIOLATCH waits are the direct fingerprint of physical I/O. A worker needs an 8KB data page that is not in the buffer pool, takes an in-memory latch on the buffer descriptor, and waits for storage to return the page. When the read completes, the worker continues. Sustained PAGEIOLATCH time means the engine is spending wall-clock time waiting on disk.

The trap is assuming the storage is slow. Often it is not. PAGEIOLATCH is the wait type for “any physical read”, and SQL Server may simply be issuing far more reads than it should because the buffer pool is shrinking under memory pressure or because a bad plan is scanning a table it should not. Before opening a ticket with your storage team, separate the failure modes.

The diagnostic fork has three branches: memory pressure spiral, genuinely slow storage, and excess reads from plan regression. For the broader SQL Server failure taxonomy, see the mental model for operators.

What this means

Two PAGEIOLATCH wait types dominate in practice:

  • PAGEIOLATCH_SH (shared): the worker intends to read the page once it is in memory. The dominant form on read-mostly OLTP workloads.
  • PAGEIOLATCH_EX (exclusive): the worker intends to modify the page once it is in memory. The page is still read from disk first, so the wait is on I/O, not on the write.

Both mean “this page is not in the buffer pool, go get it.” Both are distinct from PAGELATCH_*, which is an in-memory latch on a page already in the buffer pool. PAGELATCH_UP and PAGELATCH_EX on pages in database ID 2 are TempDB allocation contention (PFS, GAM, SGAM pages). They are not I/O waits and faster storage does not help them.

The fork has three branches, not two. First, storage is genuinely slow. Second, the buffer pool is shrinking, so the same workload produces more physical reads. Third, SQL Server is issuing more reads than baseline because of plan regression, missing indexes, or implicit conversions. Only the first is a storage problem. The other two are SQL Server problems wearing an I/O mask.

flowchart TD
    A[PAGEIOLATCH_SH/EX dominant] --> B[Snapshot waits, PLE, I/O stall]
    B --> C{PLE dropping?}
    C -- Yes --> D[Memory pressure spiral]
    C -- No, stable --> E{I/O stall per file high?}
    E -- Yes --> F[Storage layer slow]
    E -- No --> G{Wait count rising vs baseline?}
    G -- Yes --> H[Bad plan driving excess reads]
    G -- No --> I[Filter driver or path latency]

A useful operator heuristic: compare wait count and wait time deltas. If the count of PAGEIOLATCH waits is roughly the same as baseline but the average wait time has grown, the I/O subsystem is slower. If the count has grown substantially, SQL Server is driving more I/Os than it should. The two require different fixes.

Common causes

CauseWhat it looks likeFirst thing to check
Memory pressure spiralPLE dropping, buffer cache hit ratio falling, lazy writer active, PAGEIOLATCH rising with itsys.dm_os_sys_memory, per-NUMA-node PLE
Slow storage subsystemPLE stable, I/O stall per file elevated, single-digit to low-double-digit ms latency on readssys.dm_io_virtual_file_stats deltas
Bad plan driving scansPlan regression in Query Store, table scan on a large table, missing index, implicit conversionTop queries by avg_physical_reads
Filter driver latencyI/O stall in SQL but PerfMon disk counters look finefltmc instances from an admin shell
Cold start / restartPLE starts at 1, climbs over minutes to hours, normalizesRecent restart in error log
Maintenance windowDBCC CHECKDB, index rebuild, full backup runningJob history, expected during window

Quick checks

Read-only. Run them in two snapshots 30 to 60 seconds apart for any counter that is cumulative.

-- Top waits (delta the values against a second run):
SELECT TOP 15
    wait_type,
    waiting_tasks_count,
    wait_time_ms,
    wait_time_ms - signal_wait_time_ms AS resource_wait_ms,
    signal_wait_time_ms
FROM sys.dm_os_wait_stats
WHERE wait_type LIKE 'PAGEIOLATCH%'
   OR wait_type LIKE 'PAGELATCH%'
ORDER BY wait_time_ms DESC;
-- PLE, instance-wide and per NUMA node:
SELECT instance_name, cntr_value AS ple_seconds
FROM sys.dm_os_performance_counters
WHERE counter_name = 'Page life expectancy'
  AND object_name LIKE '%Buffer%';
-- I/O stall per file (snapshot twice, compute deltas):
SELECT
    DB_NAME(vfs.database_id) AS db_name,
    mf.name AS file_name,
    mf.type_desc,
    vfs.num_of_reads,
    vfs.io_stall_read_ms,
    CASE WHEN vfs.num_of_reads > 0
         THEN vfs.io_stall_read_ms * 1.0 / vfs.num_of_reads
         ELSE 0 END AS avg_read_latency_ms,
    vfs.num_of_writes,
    vfs.io_stall_write_ms,
    CASE WHEN vfs.num_of_writes > 0
         THEN vfs.io_stall_write_ms * 1.0 / vfs.num_of_writes
         ELSE 0 END AS avg_write_latency_ms
FROM sys.dm_io_virtual_file_stats(NULL, NULL) vfs
JOIN sys.master_files mf
    ON vfs.database_id = mf.database_id AND vfs.file_id = mf.file_id
ORDER BY vfs.io_stall_read_ms + vfs.io_stall_write_ms DESC;
-- Buffer cache hit ratio and OS memory pressure notification:
SELECT
    (a.cntr_value * 1.0 / NULLIF(b.cntr_value, 0)) * 100.0 AS buffer_cache_hit_ratio
FROM sys.dm_os_performance_counters a
JOIN sys.dm_os_performance_counters b
    ON a.object_name = b.object_name
   AND a.instance_name = b.instance_name
WHERE a.counter_name = 'Buffer cache hit ratio'
  AND b.counter_name = 'Buffer cache hit ratio base';

SELECT total_physical_memory_kb / 1024 AS total_mb,
       available_physical_memory_kb / 1024 AS available_mb,
       system_memory_state_desc
FROM sys.dm_os_sys_memory;
-- Memory grants pending (compounding factor in spiral):
SELECT cntr_value AS memory_grants_pending
FROM sys.dm_os_performance_counters
WHERE counter_name = 'Memory Grants Pending'
  AND object_name LIKE '%Memory Manager%';
# Filter drivers (Windows admin shell). Look for AV, backup, encryption, compression:
fltmc instances

How to diagnose it

  1. Snapshot sys.dm_os_wait_stats, wait 30 to 60 seconds, snapshot again. Compute deltas per wait type. The DMV is cumulative since startup; a single read tells you the lifetime average, not current behavior.

  2. Confirm PAGEIOLATCH is actually dominant in the delta, not in the cumulative. A wait that dominated for two days but is now zero will still top the cumulative chart.

  3. Check PLE. Read both the Buffer Manager counter and the per-node Buffer Node counters. The instance-wide value is an average across NUMA nodes; one node can be starved while the aggregate looks healthy.

  4. Check I/O stall per file with sys.dm_io_virtual_file_stats, snapshotted twice. Calculate per-file average read and write latency from the deltas. Apply these bands from the SQL Server signal catalog:

    • Data file reads: under 10ms excellent, 10 to 20ms acceptable, over 20ms degraded, over 50ms severe.
    • Log file writes: under 2ms excellent, 2 to 5ms acceptable, over 5ms degraded, over 15ms severe.
  5. Apply the fork:

    • PLE dropping AND PAGEIOLATCH rising means memory pressure spiral. The buffer pool is evicting pages that will be needed again. The I/O subsystem may be perfectly healthy but is being asked to do too much. Look at sys.dm_os_sys_memory, sys.dm_exec_query_memory_grants, and the memory clerk breakdown in sys.dm_os_memory_clerks.
    • PLE stable AND PAGEIOLATCH high means the storage layer. The buffer pool is doing its job but reads are returning slowly. The per-file I/O stall numbers will confirm which file and which database. Hand that evidence to the storage team.
    • PLE stable AND I/O stall normal AND PAGEIOLATCH count growing means SQL Server is issuing more reads than baseline. Plan regression, missing index, or implicit conversion driving a scan. Open Query Store.
  6. If storage is the suspected branch but Windows PerfMon disk counters show no latency, suspect filter drivers. Antivirus, backup agents, encryption, and compression products sit between SQL Server and the partition manager and add latency that is invisible to PerfMon but visible to SQL Server. fltmc instances from an elevated command prompt lists them.

  7. Check the error log for I/O warnings. Error 823 is a hard I/O error. Error 824 is a logical consistency error. Error 825 means the read succeeded on retry but the medium is deteriorating. Error 833 indicates I/O requests taking longer than 15 seconds to complete on a file. Any of these is immediately actionable.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
PAGEIOLATCH_* wait time (delta)Direct measure of physical read waitSustained above 30 to 40% of total wait time
PLE per NUMA nodeBuffer pool churnSudden 50%+ drop from baseline; one node much lower than others
I/O stall per fileStorage latency as the engine sees itData file reads above 20ms, log writes above 5ms sustained
Buffer cache hit ratioWorking set fit in memoryOLTP under 95% sustained, with PAGEIOLATCH rising
Memory grants pendingMemory pressure compounding the spiralAny sustained nonzero value
Lazy writer pages/secBuffer pool evicting pagesSustained nonzero outside checkpoint
Top queries by avg_physical_readsWhich queries are driving the I/ORegression vs baseline in Query Store
Error log entries 823, 824, 825, 833Storage layer healthAny occurrence

Fixes

If PLE is dropping (memory pressure spiral)

The I/O subsystem is not the problem. The buffer pool is too small for the working set.

  1. Check sys.dm_os_sys_memory. If available_physical_memory_kb is low, the OS is squeezing SQL Server. Either max server memory is set too high (starving the OS) or another process on the host is consuming memory.
  2. Check sys.dm_exec_query_memory_grants. A single query with an oversized grant can evict buffer pool pages. If you find one, KILL the session. Rollback may take time.
  3. Check the memory clerk breakdown. If CACHESTORE_SQLCP (plan cache) is more than 10 to 15% of max server memory, the plan cache is polluted by single-use ad-hoc plans. Enable optimize for ad hoc workloads to cache stubs on first execution.
  4. On multi-NUMA systems, confirm PLE per node. Foreign memory access is invisible in the aggregate counter.
  5. If max server memory is genuinely too low for the working set, raise it after you confirm the OS has headroom. Do not raise it under OS pressure.

If PLE is stable (genuinely slow storage)

The per-file I/O stall numbers from sys.dm_io_virtual_file_stats are your evidence.

  1. Identify the worst file. Join to sys.master_files to confirm type (data vs log) and physical name. Use type_desc, not file_id, to distinguish data from log files.
  2. On cloud VMs, compare current IOPS and throughput to the disk tier’s provisioned cap. Latency that spikes as throughput climbs usually means you have hit the cap. Resize the disk or move the file.
  3. On physical or SAN storage, check controller queue depth, HBA saturation, and any concurrent maintenance such as array rebuild, snapshot, or replication.
  4. Confirm files are placed correctly. Log and data on the same spindle is a classic misconfiguration. TempDB on slow shared storage undermines every spill.
  5. Pre-size database files to avoid auto-grow stalls. Auto-grow pauses I/O to the file while initializing. Log files cannot use Instant File Initialization, so log auto-grow is especially expensive.

If wait count is rising (bad plan driving excess I/O)

The I/O subsystem is fine, PLE is fine, but SQL Server is reading far more pages than baseline.

  1. Open Query Store. Find queries with a recent duration regression and a corresponding increase in avg_physical_reads.
  2. Compare the current plan to the previous known-good plan. Look for: table scan where there used to be an index seek, nested loops driving random I/O on a large set, or a missing join predicate.
  3. Force the known-good plan with sp_query_store_force_plan as a holding measure.
  4. Check for implicit conversions, which kill sargability. Query Store’s plan XML exposes these as warnings.
  5. Update statistics if they are stale. Cardinality estimation errors cause both bad plans and underestimated memory grants.

If filter drivers are the suspect

PerfMon says the disk is fast. SQL Server says it is not. The latency is being added between SQL Server and the partition manager.

  1. Run fltmc instances from an elevated command prompt.
  2. Identify antivirus, backup, encryption, and compression drivers attached to the volume hosting database files.
  3. Exclude the SQL Server data, log, and backup directories from real-time scanning per Microsoft’s published guidance. Do not exclude the SQL Server process binary itself.
  4. Re-snapshot I/O stall after each change. Filter driver latency often disappears immediately after exclusion.

Prevention

  • Snapshot wait stats at 30 to 60 second intervals and store deltas. Cumulative-since-startup values are useless for current-state diagnosis.
  • Track PLE per NUMA node, not just instance-wide. The aggregate hides single-node starvation.
  • Track I/O stall per file, not per volume. A single hot file on a shared volume changes the conversation with the storage team.
  • Establish a batch requests per second baseline by time of day and day of week. Without it you cannot tell a workload surge from a real problem.
  • Pre-size data and log files. Treat auto-grow as an incident, not a routine.
  • Capture top queries by physical reads in Query Store continuously. The plan regression that ends in PAGEIOLATCH usually shows up there days before users notice.
  • Enable optimize for ad hoc workloads. Keeps single-use plans from consuming buffer pool memory.

How Netdata helps

Netdata surfaces the signals for the PAGEIOLATCH fork at per-second resolution, which is the granularity this diagnosis actually needs.

  • The wait statistics stream shows PAGEIOLATCH_SH, PAGEIOLATCH_EX, PAGELATCH_*, WRITELOG, and ASYNC_NETWORK_IO together, so you can confirm which wait is dominant without running a manual delta query.
  • PLE is collected per NUMA node as well as instance-wide, so single-node starvation is visible without a manual DMV query.
  • Buffer cache hit ratio, lazy writer activity, and memory grants pending are on the same dashboard, so the memory pressure spiral is visible as a composite, not as a single counter.
  • ML-based anomaly detection flags the rising PAGEIOLATCH pattern before it crosses a fixed threshold, which is useful on workloads where the absolute baseline varies by time of day.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.