The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-connection-count-climbing

Operations Guides

SQL Server user connections climbing: connection pool leaks and retry storms

A rising User Connections counter is easy to misread. It does not mean that many queries are running. Most application connections are pooled and idle, so the count can climb for hours while CPU, I/O, and batch rate look almost normal.

The useful split is shape and correlation. A slow upward trend without a matching rise in Batch Requests/sec usually points to an application-tier connection pool leak or pool fragmentation. A sudden spike, especially after transient errors, points to a retry storm or a burst of new application instances. The dangerous endpoint is the same: enough concurrent active work consumes worker threads until new requests queue on THREADPOOL, which is effectively connection refusal.

Treat the counter as a correlate, not a page trigger by itself. Page when the climb converges with worker exhaustion signals: active requests near max_workers_count, work_queue_count above zero, sustained THREADPOOL waits, and failed connect or query probes.

What this means

SQL Server sees application pool size multiplied by application instance count, not end users. If each app instance keeps 50 pooled connections and you scale from 4 instances to 40, the instance can gain 1,800 mostly idle sessions without any single user doing more work.

Idle pooled sessions are not free, but they are not the immediate killer either. The immediate risk is when those sessions become active together, hold locks, wait on external calls, or retry in a tight loop. Then each request needs a worker thread. When workers run out, SQL Server can still be alive on TCP 1433 while refusing useful work.

flowchart TD
  A[User Connections rising] --> B{Batch Requests/sec rising too?}
  B -- Yes --> C[Real load or retry storm]
  B -- No --> D[Idle pooled sessions accumulating]
  C --> E[Check active requests and THREADPOOL]
  D --> F[Group sessions by host program login]
  E --> G[Worker exhaustion risk]
  F --> H[Find app tier leak]

The first job is to separate three states: idle pooled connections, active requests, and blocked requests that are pinning workers. The second is to find whether the growth is gradual leak behavior or event-driven storm behavior.

Common causes

CauseWhat it looks likeFirst thing to check
Application connection pool leakUser Connections trends up over hours or days while Batch Requests/sec stays flat. Sessions are mostly sleeping.Group sys.dm_exec_sessions by host_name, program_name, and login_name.
Pool fragmentation or misconfigurationMore pools than expected: many hosts, many app identities, or many connection strings. Count tracks app instances, not users.Compare connection count before and after a deploy or scale-out. Look for new program_name or host_name values.
Retry storm after transient errorsSharp spike in connections and logins after timeouts, deadlocks, network blips, or a brief SQL unresponsiveness event.Align the spike with error log entries, app timeout logs, and Batch Requests/sec.
Blocking cascade converting connections into workersConnections are not just idle. Many sessions are suspended on LCK_M_*, CPU drops, and workers climb toward exhaustion.Query blocking chains and identify the head blocker.
Parallel query or external wait pileupActive requests use multiple workers or stall outside the engine. Worker use rises faster than connection count.Check active requests, sys.dm_os_schedulers, and dominant waits.
Orphaned sessions from crashed app processesOld host names or stale sessions persist after an app crash or bad deployment.Look for old last_request_end_time values with no matching active workload.

Quick checks

Run these read-only checks from sqlcmd or SSMS. On named instances, counter object names use the MSSQL$INSTANCENAME: prefix.

# Current User Connections counter
sqlcmd -S localhost -E -Q "SELECT cntr_value AS user_connections FROM sys.dm_os_performance_counters WHERE counter_name = 'User Connections' AND object_name LIKE '%General Statistics%';"
-- Who owns the sessions, and how many are actually active
SELECT
    s.login_name,
    s.host_name,
    s.program_name,
    COUNT(*) AS connection_count,
    SUM(CASE WHEN r.session_id IS NOT NULL THEN 1 ELSE 0 END) AS sessions_with_requests
FROM sys.dm_exec_sessions s
LEFT JOIN sys.dm_exec_requests r ON s.session_id = r.session_id
WHERE s.is_user_process = 1
GROUP BY s.login_name, s.host_name, s.program_name
ORDER BY connection_count DESC;
-- Active request pressure
SELECT COUNT(*) AS active_requests
FROM sys.dm_exec_requests
WHERE status IN ('running', 'runnable', 'suspended');

SELECT max_workers_count FROM sys.dm_os_sys_info;

SELECT scheduler_id, runnable_tasks_count, active_workers_count, work_queue_count
FROM sys.dm_os_schedulers
WHERE status = 'VISIBLE ONLINE'
ORDER BY work_queue_count DESC, runnable_tasks_count DESC;
-- THREADPOOL and lock waits
SELECT wait_type, waiting_tasks_count, wait_time_ms, signal_wait_time_ms
FROM sys.dm_os_wait_stats
WHERE wait_type IN ('THREADPOOL', 'LCK_M_S', 'LCK_M_X', 'LCK_M_U', 'SOS_SCHEDULER_YIELD')
ORDER BY wait_time_ms DESC;
-- Batch request rate over 10 seconds
DECLARE @t1 BIGINT, @t2 BIGINT;
SELECT @t1 = cntr_value
FROM sys.dm_os_performance_counters
WHERE counter_name = 'Batch Requests/sec'
  AND object_name LIKE '%SQL Statistics%';
WAITFOR DELAY '00:00:10';
SELECT @t2 = cntr_value
FROM sys.dm_os_performance_counters
WHERE counter_name = 'Batch Requests/sec'
  AND object_name LIKE '%SQL Statistics%';
SELECT (@t2 - @t1) / 10.0 AS batch_requests_per_sec;
-- Blocking chains turning sessions into pinned workers
SELECT
    r.session_id AS blocked_session,
    r.blocking_session_id AS blocker,
    r.wait_type,
    r.wait_time / 1000 AS wait_seconds,
    DB_NAME(r.database_id) AS database_name,
    t.text AS blocked_query_text
FROM sys.dm_exec_requests r
CROSS APPLY sys.dm_exec_sql_text(r.sql_handle) t
WHERE r.blocking_session_id <> 0
ORDER BY r.wait_time DESC;

If normal connections stop responding during suspected worker exhaustion, use the Dedicated Admin Connection for diagnosis (sqlcmd -A, or ADMIN:servername in SSMS). Do not use DAC as the fix; use it to see what normal sessions cannot.

How to diagnose it

  1. Confirm the shape. Sample User Connections every 30 to 60 seconds and compare it with Batch Requests/sec over the same window. A gradual climb with flat batch rate suggests leak or pooling change. A vertical jump after errors suggests storm behavior.

  2. Break down ownership. Group sessions by login_name, host_name, and program_name. One app identity across many hosts usually means scale-out multiplied by pool size. One host with an abnormal count usually means a leak, a stuck process, or a bad deploy.

  3. Separate idle from active. Join sessions to requests. If thousands of sessions have no row in sys.dm_exec_requests, they are pooled or orphaned. If many have requests in suspended status, they are consuming workers while waiting.

  4. Check worker headroom. Compare active request pressure with max_workers_count and look for work_queue_count > 0 or sustained THREADPOOL. Do not raise max worker threads as the first response; find what is pinning workers.

  5. Look for a head blocker. If LCK_M_* waits rise while CPU falls, find the head blocker. A sleeping head blocker with an open transaction is the classic application bug that turns a normal connection pool into a worker thread fire.

  6. Correlate with application events. Line up the SQL timeline with deployments, autoscaling events, app restarts, timeout bursts, deadlock retries, and transient network errors. SQL Server can show the blast radius; the app tier usually owns the trigger.

  7. Decide the incident class. If connections are idle and workers are healthy, you have time to fix pooling. If active requests, blocked chains, and THREADPOOL are rising, treat it as an availability incident and reduce incoming work before the instance becomes unreachable.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
User ConnectionsTracks pooled and active sessions. Useful for trend and ownership, useless alone.Sustained rise beyond baseline, especially after deploys or scale-out.
Batch Requests/secTells you whether work arrival rose with connections.Connections up but batches flat means idle accumulation. Both spiking means storm or real surge.
Sessions grouped by host, program, loginConverts a vague count into an app-tier suspect list.One new program_name, host set, or identity dominates growth.
Active requestsShows how many connections are actually executing or waiting.Rising active requests while batch completion falls.
max_workers_count vs workersWorker threads are the cliff edge behind connection growth.Active workers sustained above 60 percent of max, urgent above 80 percent (heuristic, not a Microsoft-published limit)
work_queue_count per schedulerDirect evidence that work is queued for lack of workers.Any sustained value above zero.
THREADPOOL waitsNew requests cannot get workers.Sustained waits with probe failures is page territory.
Blocking chain depth and head blocker stateBlocking pins workers and converts modest connection growth into exhaustion.Sleeping head blocker, chain depth above 5, or waits over 60 seconds.
Error log and app timeout logsReveal the transient event that started a retry storm.Burst of timeouts, deadlocks, login failures, or network errors immediately before the spike.

Fixes

Fix a pool leak

Stop the growth at the app tier. Identify the leaking host or process from session ownership, then recycle the smallest blast radius that clears the leak: one pool, one worker process, or one app instance. A full SQL restart clears sessions but hides evidence and usually does not fix the cause.

On the application side, look for code paths that open connections without disposing them, asynchronous paths that abandon operations before completion, and retry logic that creates a new connection instead of reusing the pool. Validate that every execute path has a deterministic close or dispose.

Tradeoff: restarting app instances is fast but can drop in-flight work. Draining instances before restart is safer but slower.

Fix pool fragmentation or misconfiguration

Reduce unnecessary pool multiplication. Common drivers are too many app identities, per-user integrated security in a web tier, inconsistent connection strings, and uncontrolled instance scale-out. The fix is to standardize connection strings, consolidate app identities where appropriate, and set pool limits as part of capacity math.

The capacity equation is simple: expected app instances times maximum pool size must stay far below the worker thread danger zone, with headroom for blocking and parallel queries. Pool limits that look reasonable on one instance become dangerous when autoscaled.

Tradeoff: smaller pools reduce SQL-side connection pressure but can queue work in the app. Larger pools absorb app bursts but move the queue into SQL Server workers.

Stop a retry storm

First reduce amplification. If the application retries aggressively after transient errors, apply backoff and jitter, cap concurrency, and shed nonessential load. Circuit breakers are more useful than hero queries during the storm.

On SQL Server, verify whether the trigger was a real engine event: brief unresponsiveness, a blocking chain, log or I/O stall, failover, or resource saturation. If the engine had a short brownout, fix that cause or the retry storm will repeat.

Tradeoff: disabling retries entirely can turn transient faults into user errors. Unbounded retries turn one transient fault into a self-inflicted denial of service.

Relieve worker exhaustion safely

If THREADPOOL is active, reduce incoming work and unblock workers. Kill only after assessment. A sleeping head blocker with an open transaction is a candidate, but rollback can take as long as the original work. Killing active large queries can also trigger long rollback and more I/O. Warning: KILL is disruptive and its rollback is single-threaded; estimate rollback cost before issuing it against a large transaction.

Do not treat raising max worker threads as the fix. Find what is consuming workers first. Extra workers can buy minutes in a true emergency, but they also let a blocking cascade grow wider before it falls over.

Prevention

  • Baseline connection shape by time and deploy. Track User Connections, batch rate, active requests, and session ownership together. A count without a baseline is noise.
  • Alert on divergence, not count alone. Ticket when connections rise beyond baseline without batch growth. Page only when connection growth converges with worker exhaustion or failed probes.
  • Make pool arithmetic explicit. Record pool size, app instance count, identity model, and connection string variants as capacity inputs.
  • Watch worker headroom continuously. Trend active workers against max_workers_count, plus work_queue_count and THREADPOOL. The cliff is binary.
  • Detect sleeping head blockers early. A 30 second check for blocking chains and idle blockers with open transactions catches the most dangerous leak-adjacent failure before users report it.
  • Test retry behavior under brownouts. Inject short transient faults in staging and observe whether retries create a second outage.
  • Keep DMV history outside SQL Server. DMVs reset on restart and ring buffers roll off. Persist samples so post-incident review can distinguish leak from storm.

How Netdata helps

  • Correlation in one timeline: Netdata helps line up User Connections with Batch Requests/sec, active requests, waits, CPU, and I/O so you can tell idle accumulation from real work arrival.
  • Worker exhaustion context: tracking scheduler backlog, worker pressure, and THREADPOOL alongside connection count shows when a medium signal is becoming an availability event.
  • Ownership drilling: session breakdowns by login, host, and program make it faster to name the app tier, deployment, or instance group driving growth.
  • Blocking cascade detection: combining lock waits, blocked session counts, and head blocker state explains the low CPU plus unresponsive pattern that often follows connection growth.
  • Storm reconstruction: per-second collection preserves the spike shape and the preceding error burst that longer scrape intervals smooth away.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.