The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-suspect-pages

Operations Guides

SQL Server suspect_pages: the durable record of storage corruption

The SQL Server error log recycles on every service restart, and by default only six archived logs are retained. When a storage fault produces errors 823, 824, or torn-page detection, the entries that document it may be gone before anyone investigates. msdb.dbo.suspect_pages is where those events also land, and unlike the error log it survives restarts. Treat it as the durable forensic record of page-level I/O corruption on the instance.

This is an operational reference for teams adding suspect_pages to their corruption-monitoring workflow: what the table records, how rows are written, the event_type lifecycle, the 1000-row ceiling, what it misses, and how to operate it alongside DBCC CHECKDB.

What the table is and why it matters

msdb.dbo.suspect_pages is a system table in msdb. Every time SQL Server encounters a page-level I/O error during a normal read, a DBCC operation, or a backup, the engine writes a row recording the database_id, file_id, page_id, the type of event, an error counter, and the timestamp.

Two properties make it more durable than the error log:

  1. The table lives in msdb. Restarting the engine, recycling the error log with sp_cycle_errorlog, or applying a cumulative update does not clear it. A page that errored three months ago is still in the table this morning.
  2. The table records the precise page address. The error log says “error 824 on database X”. suspect_pages says database 7, file 1, page 4827301, bad checksum, four occurrences since the first event.

The tradeoff is that suspect_pages captures only one class of corruption: pages that hit an 823, 824, or torn-page error at the I/O layer. It is not a substitute for DBCC CHECKDB. Many corruption types CHECKDB detects (metadata corruption, allocation errors, structural index corruption) never produce an 823 or 824 and therefore never appear in suspect_pages.

How rows get written

SQL Server writes a row when it cannot validate a page after reading it from disk. The triggers are:

  • An 823 error: the OS returned an error from the read call. Hardware or filesystem level.
  • An 824 error: the OS returned the page, but SQL Server’s internal checks failed. Three subtypes are possible: bad checksum, torn page, or other logical consistency failure.
  • A torn page detected during recovery: only some of the sectors in the 8KB page made it to disk before the write was interrupted.

Pages are added when accessed. A page can be corrupt on disk but never read, in which case it never appears in suspect_pages. This is one reason a clean suspect_pages table is not proof of integrity.

The table is also written during Always On automatic page repair. When a replica encounters a corrupt page and obtains a good copy from its partner, the suspect_pages row is updated to reflect the resolution (see the event_type table below).

event_type values and the page lifecycle

The event_type column records what happened to the page. Three values mean active corruption; the others mean resolution. Investigate the first three, treat the rest as audit history.

event_typeMeaningWhat to do
1823 OS-level read error, or 824 logical consistency error that is not a bad checksum or torn pageInvestigate storage. Run DBCC CHECKDB on the database.
2824 bad checksumSame as above. The page was written with a checksum that no longer matches.
3824 torn pageSame as above. Typically a power event or storage crash during the write.
4Restored. The page was replaced from a backup or, during Always On automatic page repair, by a secondary pulling a good copy from the primary.Audit only. The corruption is resolved.
5Repaired. DBCC repaired the page (or the primary pulled a good copy from a secondary during AG automatic page repair).Audit only. Confirm with CHECKDB that the repair is sound.
7Deallocated by DBCC. DBCC could not repair and removed the page.Audit only, but expect data loss. Identify what was on the page.

event_type 5 records pages DBCC repaired (REPAIR_REBUILD or REPAIR_ALLOW_DATA_LOSS); event_type 7 records pages DBCC deallocated, which is a data-loss outcome of REPAIR_ALLOW_DATA_LOSS.

A page can move from event_type 1/2/3 to 4, 5, or 7 as it is restored, repaired, or deallocated. The original error entry is not deleted. SQL Server updates the existing row to reflect the latest state.

The diagram below shows the path a single page takes through the table:

flowchart LR
  A[Healthy 8KB page] --> B[823 or 824 on read]
  B --> C[Row added
event_type 1, 2, or 3] C --> D{Resolution} D --> E[Restore: event_type 4] D --> F[DBCC repair: event_type 5] D --> G[DBCC deallocate: event_type 7] E --> H[Row stays until DBA removes] F --> H G --> H

error_count and last_update_date

Two columns deserve attention beyond event_type:

  • error_count increments each time SQL Server records a failure for the same page. A value of 1 is consistent with a transient I/O blip. A value greater than 1 on the same page is a hardware escalation signal: multiple failures on the same sector indicate the medium is deteriorating, not that there was a one-off cable disconnect.
  • last_update_date is the timestamp of the most recent event for the page, including resolution events. Any change to a row updates this column.

Use both columns to triage. Sort by error_count DESC to find pages failing repeatedly; sort by last_update_date DESC to find pages that errored recently.

The 1000-row ceiling

suspect_pages has a hard limit of 1000 rows. When the table is full, SQL Server stops writing new entries. New 823, 824, and torn-page events still go to the SQL Server error log, but the durable record stops growing, and SQL Server raises error 4345 to flag the overflow.

Two operational consequences follow:

  1. Row count is a first-class signal. A table approaching 1000 rows means either a long history of unaddressed corruption or an active hardware failure generating many bad pages per minute. Either is urgent.
  2. The table must be pruned manually. SQL Server never removes rows on its own, even after they have been investigated or the underlying corruption has been resolved. The only automatic deletions occur when a database file is removed with ALTER DATABASE ... REMOVE FILE or when a database is dropped. Everything else requires a DBA to delete the row.

What suspect_pages does not catch

Three classes of corruption are invisible to suspect_pages:

  1. DBCC-only corruption. Allocation errors, metadata corruption, structural index problems, and many other forms CHECKDB detects do not trigger an 823 or 824. They never reach suspect_pages. A clean suspect_pages table alongside CHECKDB output that reports errors is normal, not contradictory.
  2. Never-accessed corruption. A page can be corrupt on disk but never read since the corruption occurred. If no query, backup, or CHECKDB touches the page, no row is written.
  3. Transient I/O errors that left no actual damage. A cable disconnect or transient checksum failure can write a row that subsequent CHECKDBs confirm is fine. The row is a true record of an event that happened, but it is not proof the page is still corrupt.

This is why suspect_pages pairs with a regular DBCC CHECKDB schedule. Weekly CHECKDB against production databases catches the corruption types that never produce an 823 or 824. suspect_pages catches the I/O-layer corruption that CHECKDB may take a week to discover, and it preserves it across restarts.

Reading the table

The standard query for active corruption:

-- Active corruption rows only
SELECT
    database_id,
    DB_NAME(database_id) AS database_name,
    file_id,
    page_id,
    event_type,
    error_count,
    last_update_date
FROM msdb.dbo.suspect_pages
WHERE event_type IN (1, 2, 3)
ORDER BY last_update_date DESC;

For full forensic context, including resolved rows that may still need investigation:

-- Full table with lifecycle interpretation
SELECT
    database_id,
    DB_NAME(database_id) AS database_name,
    file_id,
    page_id,
    event_type,
    CASE event_type
        WHEN 1 THEN '823 or 824 (other)'
        WHEN 2 THEN '824 bad checksum'
        WHEN 3 THEN '824 torn page'
        WHEN 4 THEN 'restored'
        WHEN 5 THEN 'repaired'
        WHEN 7 THEN 'deallocated by DBCC'
    END AS event_type_desc,
    error_count,
    last_update_date
FROM msdb.dbo.suspect_pages
ORDER BY database_id, file_id, page_id;

Row count and table fill ratio:

-- How close is the table to the 1000-row ceiling?
SELECT
    COUNT(*) AS row_count,
    1000 - COUNT(*) AS slots_remaining
FROM msdb.dbo.suspect_pages;

Anyone with access to msdb can read the data in suspect_pages. Modifying rows requires UPDATE permission on the table; members of the db_owner fixed database role in msdb or the sysadmin fixed server role can insert, update, or delete.

Operational workflow

A mature corruption-monitoring routine has three parts: alert on new entries, investigate them, and prune the table.

Alert on any new row with event_type 1, 2, or 3. SQL Server has no built-in alert for this. The standard implementation is a SQL Server Agent job that polls the table at a short interval (every 1 to 5 minutes), persists the maximum (database_id, file_id, page_id) or last_update_date it has seen, and fires when a new row appears. Investigate every new entry. A new event_type 1, 2, or 3 row means a page failed an I/O check. Pair the row with the SQL Server error log around last_update_date to find the matching 823 or 824 entry, which carries the OS error code and the operation that triggered the read. Run DBCC CHECKDB on the affected database. If the database is in an Availability Group, check whether automatic page repair has updated the row to event_type 4 or 5.

Prune investigated rows. Once an entry is understood and either resolved (event_type 4, 5, or 7) or confirmed as a transient blip, delete it from the table. This keeps the row count well below 1000 and ensures the next new-entry alert is meaningful. There is no built-in retention policy.

Warning: Only delete rows you have investigated and documented. Deleting a row does not repair the page on disk; it only removes the durable record. Always pair pruning with an incident record so the page_id remains traceable.

Pruning is a targeted DELETE, typically scoped to specific resolved rows:

-- Delete only rows you have investigated and recorded.
-- NEVER run a blanket DELETE without a WHERE clause.
DELETE FROM msdb.dbo.suspect_pages
WHERE database_id = <db_id>
  AND file_id = <file_id>
  AND page_id = <page_id>;

The checklist below summarizes the recurring cycle:

  • New-row alert on event_type 1, 2, or 3. Active storage degradation. Investigate now.
  • Row count approaching 1000. Table is about to silently stop recording. Investigate immediately and prune.
  • error_count greater than 1 on a single page. Same page failing repeatedly. Treat as hardware escalation.
  • Investigated rows with event_type 4, 5, or 7. Audit history. Safe to delete after the incident is closed.
  • Weekly DBCC CHECKDB. Catches corruption types that never hit suspect_pages.

Pairing with errors 823, 824, and 825

suspect_pages rows correlate with three error log signatures. The related guides on error 823/824 and error 825 cover these in detail; the relevant point here is what each one contributes to suspect_pages:

  • Error 823. The OS returned an error from the read. SQL Server writes a suspect_pages row with event_type 1.
  • Error 824. The OS returned the page but SQL Server’s internal validation failed. event_type is 1 (other), 2 (bad checksum), or 3 (torn page) depending on the specific failure.
  • Error 825. A read failed, SQL Server retried, and the retry succeeded. This is the disk-failing canary. The read succeeded, so no suspect_pages row is written because no corruption was ultimately detected. Watch error 825 in the error log, not in suspect_pages.

The asymmetry matters. Error 825 indicates a failing disk that has not yet produced a persistently corrupt page. suspect_pages indicates a corrupt page that has been observed. Both signals are needed; neither is sufficient on its own.

How Netdata helps

suspect_pages is a point-in-time table, not a streaming metric. The value of monitoring it is in correlating new rows with the surrounding storage and engine signals.

  • Row count trended over time. Plotting suspect_pages row count per interval exposes both the steady accumulation that signals an unhealthy table and the sudden jump that signals a hardware event in progress.
  • New-row events correlated with I/O stall. When a new suspect_pages row appears, the same time window in sys.dm_io_virtual_file_stats typically shows latency spikes on the affected file. The two together narrow the failure from “the instance has corruption” to “this volume is producing bad pages”.
  • Correlation with errors 823, 824, and 825. Error log signals land on the same time axis as suspect_pages changes. A burst of 825 retries that precedes a new suspect_pages row is the signature of a disk that was already failing.
  • Database state transitions. A database moving from ONLINE to SUSPECT often coincides with a flood of new suspect_pages entries. Alerting on both together reduces false positives.
  • Per-second granularity. Storage faults can produce many bad pages per minute. Per-second sampling captures the slope that minute-level polling flattens.

Netdata’s Microsoft SQL Server monitoring brings these signals together with per-second metrics and anomaly detection alongside CHECKDB scheduling and error log collection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.