The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-sos-scheduler-yield-waits

Operations Guides

SQL Server SOS_SCHEDULER_YIELD waits: CPU scheduler pressure explained

When SOS_SCHEDULER_YIELD dominates sys.dm_os_wait_stats, the reflex is to assume CPU pressure and start hunting for bad queries. That reflex is right about half the time. The other half, you are chasing a signal that is doing exactly what it was designed to do: recording every time a worker voluntarily yielded its 4ms quantum because other runnable workers were queued.

The wait type has existed in SQL Server since the SQLOS era and the 4ms quantum is fixed in every version. It cannot be tuned. What you can tune is your interpretation. A workload doing efficient set-based scans of pages already in memory will yield constantly and rack up enormous SOS_SCHEDULER_YIELD numbers without anything being wrong. A VM on an oversubscribed host reports the same wait type while the hypervisor silently steals CPU cycles SQL Server cannot see.

This article covers the mechanism, the signals that separate benign from pathological yielding, the production patterns that drive the pathological case (missing indexes, parameter sniffing, compilation storms, VM steal, soft-NUMA surprises), and the diagnostic queries that take you from “SOS_SCHEDULER_YIELD is high” to “here is the plan consuming the scheduler.”

What this means

SQL Server runs its own cooperative, non-preemptive scheduler through SQLOS. There is one scheduler per logical CPU. Each scheduler has three lists: running (the worker currently on CPU), runnable (workers ready to run, queued for their turn), and waiter (workers blocked on a resource such as I/O, a lock, or memory). Workers are not preempted by the OS. They run until they hit a yield point, at which point SQLOS checks whether the 4ms quantum has been exhausted.

When a worker exhausts its quantum and the runnable queue is nonempty, the worker goes to the bottom of the runnable queue. The time it spends there, waiting to get back on CPU, is recorded as SOS_SCHEDULER_YIELD. Two properties follow that operators must internalize:

  • Resource wait is always zero. wait_time_ms - signal_wait_time_ms for SOS_SCHEDULER_YIELD is always 0. The worker is not waiting on a resource. The entire wait is signal wait: time spent runnable, queued for a CPU.
  • Count alone is meaningless. A high waiting_tasks_count with single-digit-millisecond average signal wait is a healthy busy system. The same count with 50ms or higher average signal wait is scheduler saturation.

This is why the signal wait ratio matters more than the raw wait time. If signal_wait_time_ms across all wait types exceeds roughly 20% of total wait_time_ms, the engine is spending a significant fraction of its wait time queued for CPU. On a VM, that number can climb while OS-reported CPU utilization stays low, because the cycles being waited for were stolen by a co-tenant or consumed by hypervisor overhead.

Common causes

CauseWhat it looks likeFirst thing to check
Real CPU-bound workloadTop waits dominated by SOS_SCHEDULER_YIELD, signal wait ratio above 20%, runnable_tasks_count sustained above 1 per scheduler, SQL Server CPU at 80%+sys.dm_exec_query_stats ordered by total_worker_time
VM CPU steal or co-tenant contentionSame wait profile but SQL Server CPU reports low (30-50%), other process CPU also low, no obvious query culpritHypervisor CPU ready (%RDY on VMware, steal in /proc/stat on Linux hosts, CPU credits on Azure burstable VMs)
Missing indexes forcing scansSOS_SCHEDULER_YIELD rising with specific queries, large total_logical_reads on a few plans, often CXPACKET as a co-wait on parallel scanssys.dm_exec_query_stats ordered by total_logical_reads, missing index DMVs
Parameter sniffing bad planSudden spike after plan recompilation, one plan shape (often nested loops on a large set) burning CPU, query duration regressedQuery Store plan regression, or compare plans in sys.dm_exec_query_plan
Compilation stormSOS_SCHEDULER_YIELD correlated with SQL Compilations/sec tracking Batch Requests/sec, plan cache churningCompilations-to-batch-requests ratio, single-use plan count
Auto soft-NUMA on large systemsSevere SOS_SCHEDULER_YIELD accumulation on systems with 9 or more cores per socket, low CPU utilization, limited parallel workloads (documented on SQL 2016/2017; behavior on later versions not confirmed from public sources)sys.dm_os_nodes node type, soft-NUMA configuration
CPU power managementErratic SOS_SCHEDULER_YIELD correlated with CPU frequency transitions, often on idle-to-burst trafficOS power plan, BIOS C-states and P-states
Edition core limitSOS_SCHEDULER_YIELD on Standard or Express when workload exceeds licensed CPU capacity, SQL Server CPU capped below host CPUEdition and licensing, scheduler count vs host CPU

Quick checks

-- Snapshot wait stats delta over 60 seconds (read-only)
DECLARE @t1 TABLE (wait_type nvarchar(60), wait_time_ms bigint, signal_wait_time_ms bigint, waiting_tasks_count bigint);
INSERT @t1 SELECT wait_type, wait_time_ms, signal_wait_time_ms, waiting_tasks_count
FROM sys.dm_os_wait_stats WHERE wait_type = 'SOS_SCHEDULER_YIELD';
WAITFOR DELAY '00:00:60';
SELECT
    s.wait_type,
    s.waiting_tasks_count - t.waiting_tasks_count AS new_yields,
    s.wait_time_ms - t.wait_time_ms AS new_wait_ms,
    s.signal_wait_time_ms - t.signal_wait_time_ms AS new_signal_ms
FROM sys.dm_os_wait_stats s
JOIN @t1 t ON t.wait_type = s.wait_type
WHERE s.wait_type = 'SOS_SCHEDULER_YIELD';
-- Signal wait ratio across all waits (read-only)
SELECT
    SUM(signal_wait_time_ms) AS total_signal_ms,
    SUM(wait_time_ms) AS total_wait_ms,
    CAST(100.0 * SUM(signal_wait_time_ms) / NULLIF(SUM(wait_time_ms), 0) AS DECIMAL(5,2)) AS signal_pct
FROM sys.dm_os_wait_stats
WHERE wait_type NOT IN (
    'SLEEP_TASK','BROKER_TO_FLUSH','BROKER_TASK_STOP','CLR_AUTO_EVENT','CLR_MANUAL_EVENT',
    'LAZYWRITER_SLEEP','SQLTRACE_BUFFER_FLUSH','WAITFOR','XE_TIMER_EVENT','XE_DISPATCHER_WAIT',
    'FT_IFTS_SCHEDULER_IDLE_WAIT','BROKER_EVENTHANDLER','SP_SERVER_DIAGNOSTICS_SLEEP',
    'HADR_FILESTREAM_IOMGR_IOCOMPLETION','DIRTY_PAGE_POLL','DISPATCHER_QUEUE_SEMAPHORE',
    'QDS_PERSIST_TASK_MAIN_LOOP_SLEEP','QDS_ASYNC_QUEUE','CHECKPOINT_QUEUE',
    'REQUEST_FOR_DEADLOCK_SEARCH','LOGMGR_QUEUE','ONDEMAND_TASK_QUEUE','HADR_WORK_QUEUE',
    'BROKER_TRANSMITTER','KSOURCE_WAKEUP');
-- Per-scheduler runnable queue depth (read-only)
SELECT scheduler_id, cpu_id, current_tasks_count,
       runnable_tasks_count, active_workers_count, work_queue_count
FROM sys.dm_os_schedulers
WHERE status = 'VISIBLE ONLINE'
ORDER BY runnable_tasks_count DESC;

How to diagnose it

The flow is: confirm the wait is real, find where the CPU is going, then decide whether the cause is inside or outside SQL Server.

flowchart TD
    A["SOS_SCHEDULER_YIELD dominates waits"] --> B{"signal_wait_time_ms
above 20% of total?"} B -- No --> C["Benign yielding
No action needed"] B -- Yes --> D{"runnable_tasks_count above 0
across schedulers?"} D -- No --> E["Check non-yielding scheduler
or measurement window"] D -- Yes --> F{"SQL Server CPU high?"} F -- Yes --> G["Real CPU pressure
Find top worker_time queries"] F -- No --> H["VM steal or external contention
Check hypervisor CPU ready"] G --> I["Missing index, bad plan,
or compilation storm"] H --> J["Resize, relocate, DRS anti-affinity"]
  1. Snapshot wait stats and compute deltas. sys.dm_os_wait_stats is cumulative since startup, or since the last DBCC SQLPERF('sys.dm_os_wait_stats', CLEAR). A single read tells you nothing about current state. Snapshot, wait 30 to 60 seconds, snapshot again, compute the delta. Do not run DBCC SQLPERF(..., CLEAR) on a production instance unless you accept losing the cumulative history.

  2. Check the signal wait ratio. In the delta window, compute SUM(signal_wait_time_ms) / SUM(wait_time_ms). Above 0.20 means the engine is spending more than 20% of its wait time runnable, queued for CPU. Below 0.10 with high SOS_SCHEDULER_YIELD counts is benign yielding. The ratio is more reliable than any absolute threshold on the wait type itself.

  3. Check runnable queue depth. Sustained runnable_tasks_count above 1 across multiple schedulers is CPU pressure. Any work_queue_count above 0 is worker thread exhaustion: a different and more urgent problem. Check max_worker_threads against current_tasks_count and look for THREADPOOL waits alongside.

  4. Find the queries consuming CPU.

-- Top queries by CPU (read-only, cumulative since plan cached)
SELECT TOP 20
    qs.sql_handle, qs.plan_handle, qs.execution_count,
    qs.total_worker_time / 1000 AS total_cpu_ms,
    qs.total_worker_time / qs.execution_count / 1000 AS avg_cpu_ms,
    qs.total_elapsed_time / qs.execution_count / 1000 AS avg_elapsed_ms,
    SUBSTRING(qt.text, 1, 200) AS query_text
FROM sys.dm_exec_query_stats qs
CROSS APPLY sys.dm_exec_sql_text(qs.sql_handle) qt
ORDER BY qs.total_worker_time DESC;

The gap between avg_cpu_ms and avg_elapsed_ms shows wait time. Queries where the two are close are CPU-bound. Queries where elapsed is much larger than CPU are waiting on something else and are not your SOS_SCHEDULER_YIELD culprit.

  1. Find live requests currently yielding. Queries incurring SOS_SCHEDULER_YIELD do not appear in sys.dm_os_waiting_tasks, because the worker is runnable, not waiting on a resource. Query sys.dm_exec_requests and filter on last_wait_type. A running request’s last_wait_type is the wait it last completed, not necessarily the wait it is currently on.
-- Live requests currently yielding (read-only)
SELECT r.session_id, r.status, r.wait_type,
       r.last_wait_type, r.wait_time,
       r.cpu_time, r.logical_reads,
       SUBSTRING(qt.text, 1, 200) AS query_text
FROM sys.dm_exec_requests r
CROSS APPLY sys.dm_exec_sql_text(r.sql_handle) qt
WHERE r.last_wait_type = 'SOS_SCHEDULER_YIELD'
ORDER BY r.cpu_time DESC;
  1. If SQL Server CPU is low and the wait is high, suspect the VM. This is the pattern that catches teams off guard. The hypervisor is not giving SQL Server the cycles it expects, so workers exhaust their quantum without actually getting 4ms of CPU time. Check VMware %RDY (sustained above 5-10% per vCPU is worth investigating; above 10% is significant contention), steal time on Linux KVM hosts, or Azure VM CPU credits and quota. Also check host power management, because aggressive C-state transitions cause the same artifact.

  2. On SQL 2016 or later with a large socket count, check soft-NUMA. Auto soft-NUMA is enabled by default on systems with more than 8 physical cores per socket and has been documented to cause severe SOS_SCHEDULER_YIELD accumulation on large systems running limited parallel workloads. The documented reproduction is from SQL 2016/2017; whether the behavior persists in SQL 2019 or 2022 is not confirmed from public sources. Check sys.dm_os_nodes and evaluate whether disabling auto soft-NUMA changes the profile. The configuration change requires a service restart.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
SOS_SCHEDULER_YIELD wait time, deltaDirect measure of time spent waiting to get back on CPU after yieldingSustained growth above baseline, especially with throughput drop
signal_wait_time_ms ratioSeparates CPU pressure from resource waitsAbove 20% of total wait time
runnable_tasks_count per schedulerCleanest in-engine CPU queue signalSustained above 1 across multiple schedulers
work_queue_count per schedulerWorker thread exhaustion indicatorAny sustained nonzero value
SQL Server CPU vs other CPU vs idleDistinguishes SQL-bound from external contentionSQL CPU low plus wait high equals VM steal
SQL Compilations/sec to Batch Requests/sec ratioCompilation pressureAbove 10-20% sustained
Batch Requests/sec trendWorkload context for interpreting waitsDrop while waits climb means throughput collapse
Top queries by total_worker_timeIdentifies queries burning CPUSudden change in ranking after deployment or stats update

Fixes

Real CPU-bound queries (missing indexes, bad plans). This is the most common actionable cause. Use the top-worker-time query above to identify culprits, then inspect their plans. Missing indexes show up as large scans with high logical reads. Parameter sniffing regressions show up as plan changes visible in Query Store. Force the known-good plan in Query Store with sp_query_store_force_plan for immediate relief, then address the root cause (hint, statistics update, query rewrite, covering index). On SQL 2022, the Parameter Sensitive Plan optimization can address the most common parameter-sniffing pattern automatically.

Compilation storms. If SQL Compilations/sec tracks Batch Requests/sec at near 1:1, the plan cache is providing no benefit. Enable optimize for ad hoc workloads at the server level to stub single-use plans on first execution and cache the full plan only on second execution. Consider forced parameterization for workloads dominated by non-parameterized ad-hoc SQL. Address the source of the ad-hoc queries if possible.

VM steal. Not fixable inside SQL Server. Work with the virtualization team to reduce host oversubscription, configure DRS anti-affinity rules to keep noisy co-tenants away, size the VM so the vCPU-to-core ratio is 1:1 for SQL Server workloads, and verify power management is set to high performance at both host and guest levels.

Auto soft-NUMA surprises. On large systems with the symptom, test disabling auto soft-NUMA in a controlled change window. The mechanism to disable auto soft-NUMA varies by SQL Server version; consult Microsoft’s documentation for the exact registry key or server configuration option applicable to your version. Measure CPU utilization and SOS_SCHEDULER_YIELD before and after. In the documented reproduction, disabling auto soft-NUMA took CPU utilization from 33% to 90% and eliminated the wait accumulation, because parallel query schedulers were no longer artificially partitioned.

SQL 2022 compilation regression. A documented case shows SQL 2022 producing significantly more SOS_SCHEDULER_YIELD waits during query compilation than SQL 2019 on identical hardware, with first execution taking 14 seconds on 2022 versus 2 seconds on 2019. As of early 2026 this remained under Microsoft support investigation. If you see SOS_SCHEDULER_YIELD concentrated on compile operations after upgrading, open a support case and reference the compilation-time regression.

Edition core limits. If Standard or Express edition is capping CPU below what the workload needs, no amount of query tuning will help. The fix is licensing or workload reduction.

Prevention

  • Baseline the signal wait ratio. Track SUM(signal_wait_time_ms) / SUM(wait_time_ms) as a regular time series. Most teams first notice CPU pressure after it is already impacting users; the signal ratio starts climbing hours or days earlier.
  • Snapshot wait stats on a 30 to 60 second cadence. Cumulative sys.dm_os_wait_stats hides current behavior. Without deltas you cannot tell whether SOS_SCHEDULER_YIELD is climbing now or whether it has been accumulating since the last restart three months ago.
  • Track the top-N queries by total_worker_time over time. Sudden changes in ranking are the earliest indicator of plan regression or workload shift.
  • On VMs, monitor hypervisor CPU ready time alongside SQL Server metrics. SQL Server cannot see CPU steal; the only in-engine symptom is SOS_SCHEDULER_YIELD with low reported CPU utilization.
  • Keep power management on high performance. Aggressive C-state and P-state transitions cause the same artifact as VM steal and are easy to overlook.
  • Validate after every SQL Server upgrade. Compilation behavior and scheduler logic have changed across versions. An upgrade that felt slow may have a SOS_SCHEDULER_YIELD fingerprint worth investigating before you start tuning queries.

How Netdata helps

  • The SQL Server collector surfaces sys.dm_os_wait_stats deltas at per-second granularity, so SOS_SCHEDULER_YIELD shows up as a current trend rather than as a cumulative number that obscures recent behavior.
  • signal_wait_time_ms ratio is computed and charted directly, which removes the most common misread of this wait type.
  • runnable_tasks_count and work_queue_count per scheduler are collected alongside CPU utilization, so real CPU pressure (high runnable, high SQL CPU) is distinguishable from VM steal (high runnable, low SQL CPU, low host CPU) in one view.
  • Wait stats correlate with Batch Requests/sec, SQL Compilations/sec, and the top-query metrics, so a compilation storm or throughput collapse appears next to the wait signal.
  • Anomaly detection on the signal wait ratio and runnable queue depth surfaces slow-build CPU pressure before users report latency.
  • Host-level CPU steal metrics sit alongside the SQL Server metrics where the hypervisor exposes them, making the VM-steal case diagnosable from the same dashboard.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.