The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-tempdb-file-configuration

Operations Guides

SQL Server TempDB file configuration: one data file per CPU and why it matters

TempDB is the shared scratchpad every database on a SQL Server instance writes to. Temp tables, table variables, sort and hash spills, row version stores for RCSI and AlwaysOn readable secondaries, and internal worktables all land here. When TempDB serializes, every database on the instance serializes with it. The most common serialization point is not space or I/O. It is logical latch contention on a handful of allocation bitmap pages.

The standard guidance is to create multiple TempDB data files. The rule most operators learn is “one data file per logical CPU.” That rule is correct as a starting point but incomplete. The file count, the file sizing, the autogrowth configuration, and the SQL Server version all change what the right number actually is.

This is the configuration side of TempDB contention. For active allocation latch waits in production, see the wait-statistics and diagnosis material in this guide cluster.

What it is and why it matters

Microsoft’s baseline guidance is:

  • If the instance has 8 or fewer logical CPUs, create one TempDB data file per logical CPU.
  • If the instance has more than 8 logical CPUs, start with 8 data files.
  • If allocation latch contention persists after that, add files in groups of 4, up to the number of logical CPUs.

The rule exists because of allocation page contention. SQL Server tracks free space in every data file using three bitmap page types: PFS (Page Free Space), GAM (Global Allocation Map), and SGAM (Shared Global Allocation Map). Every allocation or deallocation of an extent or page requires an in-memory latch update on the relevant bitmap page. With a single TempDB data file, every concurrent temp object creation, every sort spill, and every version store row update lines up on the same PFS, GAM, and SGAM pages. The latches are held for microseconds, but at thousands of allocations per second the queue becomes the bottleneck.

The symptom is not a clean error. It is throughput that will not climb no matter how much CPU, memory, or I/O headroom you add. PAGELATCH_UP and PAGELATCH_EX waits on pages in database ID 2 (TempDB) dominate the wait statistics. CPU utilization looks moderate. Disk latency looks fine. The server is idle and slow at the same time.

How it works

Multiple TempDB data files work because each file has its own allocation bitmap pages. Two files means two PFS page sets, two GAM page sets, two SGAM page sets. The latch contention that was concentrated on one set of pages is now spread across N sets. This is the entire mechanism. The file count is a serialization dilution strategy.

The allocation engine uses a proportional-fill algorithm. When SQL Server needs to write to TempDB, it does not round-robin across files. It writes to the file with the most free space, weighted proportionally. If all files are the same size with the same amount of free space, allocations distribute evenly across them. If one file is larger than the others, it receives a disproportionate share of new allocations because it has more free space.

This is why equal sizing is not a recommendation. It is the precondition that makes multiple files reduce contention. Unequal files defeat the algorithm.

flowchart TD
    A[Concurrent temp object creation] --> B[Each allocation needs a PFS/GAM page update]
    B --> C{TempDB file count}
    C -->|1 file| D[All updates serialize on the same bitmap pages]
    C -->|N equal files| E[Updates spread across N bitmap page sets]
    D --> F[PAGELATCH_UP/EX waits on database ID 2]
    E --> G[Latch contention distributed across files]
    F --> H[Throughput throttled without errors]
    G --> I[Allocation scales with file count]

Adding files does not make TempDB faster in the I/O sense. It removes a logical serialization point so the existing I/O and CPU can actually be used.

Where it shows up in production

TempDB contention is a high-concurrency OLTP problem. It appears when many sessions simultaneously create and drop temp objects, spill sorts or hashes, or generate row versions. The classic triggers:

  • Heavy temp table use in stored procedures. Temp tables created and dropped per request, especially under high request rates, hammer the allocation pages.
  • RCSI or snapshot isolation enabled. Every read transaction under Read Committed Snapshot Isolation generates row versions in TempDB. Every update generates the before-image version. The version store is append-heavy and allocation-intensive.
  • AlwaysOn readable secondaries. Readable secondaries use snapshot isolation, so they generate TempDB version store rows on the secondary side.
  • Massive sort or hash spills. A query with a badly underestimated memory grant spills a multi-gigabyte sort or hash into TempDB. One query can consume gigabytes and trigger autogrowth.

The pattern that catches teams off guard is the post-restart slowdown. TempDB is recreated on every SQL Server restart, at the configured file sizes recorded in sys.master_files (tempdb growth is not reflected in that view, so a grown file returns to its configured size). If TempDB was never pre-sized for the workload, the first big workload after a restart forces autogrowth events. Every autogrowth event pauses I/O to that file while the new space is initialized. Log file autogrowth is especially expensive because instant file initialization does not apply to log files. A cluster of autogrowth events on the first morning after a patching reboot looks like a mysterious performance regression.

Tradeoffs and when to use it

The one-file-per-CPU rule is a safe baseline, not a law. Several refinements matter in production.

Equal size and equal growth are non-negotiable. Every TempDB data file must have the same initial size and the same autogrowth increment, specified in fixed megabytes, not percent. If files start equal but one autogrows first (because growth increments are in percent, or because someone added a file at a different size), proportional fill starts concentrating allocations on the larger file. The new files stop absorbing work. You end up with N files but the contention profile of 1.

Pre-size TempDB for the workload. Configure the initial file sizes large enough that autogrowth never fires during normal operations. Ensure instant file initialization is enabled so data file autogrowth does not stall on zero-initialization. The goal is to make autogrowth a safety net, not a routine occurrence. A common operational target is 25 to 30 percent free space under peak load.

The one-file-per-core ceiling is not always 8. Microsoft’s “start with 8, add in groups of 4” guidance assumes typical OLTP concurrency. On very high core count systems, blindly creating one file per core creates hundreds of files, introducing its own overhead in allocation tracking and file management. Community guidance, notably from Paul Randal, suggests one quarter to one half of the logical CPU count is often sufficient on modern SQL Server versions where several allocation contention sources have been removed at the engine level. The right workflow is to start with the Microsoft baseline, measure PAGELATCH waits on TempDB pages, and add files only when the measurement justifies it.

Adding files does not retroactively fix unequal sizing. If existing TempDB files have grown through autogrowth to different sizes, adding new files at the original configured size does nothing useful. The new files are smaller, proportional fill concentrates on the larger existing files, and the new files contribute almost no contention relief. When adding files to a system that has already autogrown unevenly, size the new files to match the current grown size of the existing files, not the original configured size.

Storage placement matters. TempDB should live on the fastest storage available. It is a write-heavy workload with random I/O patterns from spills and version store appends. On local SSD or NVMe, this is straightforward. On shared SAN storage, TempDB competes with data and log I/O. On cloud VMs, TempDB on the OS disk or a slow data disk is a common misconfiguration that caps the whole instance. Splitting TempDB files across multiple physical volumes can improve I/O parallelism, but only after the file count and sizing are correct.

Version-specific improvements change the calculus. Several SQL Server releases have reduced the allocation contention that multiple files exist to mitigate:

  • SQL Server 2016 and later. Trace flags 1117 (uniform autogrowth across files in a filegroup) and 1118 (uniform extent allocation instead of mixed extents) became default behavior for TempDB. Do not enable these trace flags for TempDB on SQL Server 2016 or newer. They are only relevant for TempDB on SQL Server 2014 and earlier.
  • SQL Server 2019 and later. Concurrent PFS page updates are enabled by default, reducing PFS latch contention directly.
  • SQL Server 2022 and later. GAM and SGAM page updates became concurrent, using shared-latch updates similar to the concurrent PFS updates introduced in SQL Server 2019, significantly improving high-concurrency TempDB allocation throughput.
  • SQL Server 2019 and later. Memory-optimized TempDB metadata removes metadata latch contention on system tables but does not address GAM or SGAM page contention. These are separate problems with separate fixes.

On SQL Server 2022 and later, the allocation latch problem is substantially smaller than on older versions. The file-count rule still applies as a baseline, but the threshold at which adding more files produces measurable improvement is higher (the SQL Server 2019+ concurrent PFS and SQL Server 2022 concurrent GAM/SGAM updates are engine-level fixes that postdate Paul Randal’s one-quarter-to-one-half recommendation).

Verify your configuration. After applying any changes, confirm the result:

SELECT name, size/128.0 AS current_size_mb, growth, is_percent_growth
FROM sys.master_files WHERE database_id = 2 ORDER BY file_id;

Every data file should show the same current_size_mb, the same growth value, and is_percent_growth = 0.

Signals to watch in production

SignalWhy it mattersWarning sign
PAGELATCH_UP and PAGELATCH_EX waits with resource_description matching 2:%Direct evidence of allocation page latch contention in TempDB (database ID 2)Sustained waits above 5 percent of total wait time, correlated with a throughput ceiling
TempDB free space from sys.dm_db_file_space_usageTempDB exhaustion halts queries across all databasesBelow 20 percent under normal load, below 10 percent urgent
Version store size from sys.dm_db_file_space_usageRCSI and snapshot isolation push version store rows into TempDB. Long-running transactions prevent cleanup.Version store above 50 percent of TempDB usage indicates a long-running transaction holding the version store open
Autogrowth events from the default traceAutogrowth pauses I/O to the file while initializing new space. Log file autogrowth is especially expensive.Any autogrowth event during business hours on a production instance
TempDB I/O latency per file from sys.dm_io_virtual_file_statsTempDB on slow storage caps every spill, version store append, and temp table operationAverage read or write latency above 20ms sustained
RESOURCE_SEMAPHORE waits and memory grants pendingMemory grant pressure causes queries to spill to TempDB. Spill volume is proportional to the grant shortfall.Any sustained nonzero memory grants pending

How Netdata helps

  • Correlate PAGELATCH waits with throughput. Netdata collects SQL Server wait statistics at per-second granularity. A sustained rise in PAGELATCH_UP or PAGELATCH_EX waits on database ID 2 pages, with no corresponding rise in CPU or I/O utilization, is the signature of TempDB allocation contention. Per-second resolution matters because latch contention spikes are short and easy to miss with 60-second polling.
  • Track TempDB space by consumer. User objects, internal objects, and version store are distinct failure modes. A growing version store points to long-running snapshot transactions. Growing internal objects point to spill-heavy queries. Netdata surfaces the breakdown so you do not chase the wrong consumer.
  • Detect autogrowth events as they happen. Because autogrowth pauses I/O and is a leading indicator of under-sized files, alerting on autogrowth events catches misconfiguration before users report latency spikes.
  • Watch version-specific behavior change. If you upgrade from SQL Server 2017 to SQL Server 2022, the latch-free GAM pages should reduce PAGELATCH contention. Tracking the wait profile before and after the upgrade confirms the improvement rather than assuming it.
  • Correlate TempDB I/O latency with spill activity. High TempDB I/O latency combined with high internal object allocation is the signature of memory grant undersizing. Netdata’s per-second collection lets you align I/O latency spikes with the workload that caused them.

Netdata’s Microsoft SQL Server monitoring with Netdata brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.