The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / microsoft-sql-server / microsoft-sql-server-error-825-read-retry

Operations Guides

SQL Server Error 825: read-retry succeeded and the disk is failing

Error 825 is what SQL Server writes to the error log when a disk read failed on the first attempt but succeeded on a retry (attempt 2, 3, or 4). The query completes. The application sees no failure. But the storage underneath just told you it is failing.

Most monitoring setups never surface Error 825. It is a severity-10 informational message, and typical SQL Server Agent alert configurations target severity 20 and above. The error sits quietly in the log until something harder arrives: an 823 (hard I/O error) or an 824 (logical consistency error). By then the page may already be unreadable.

Treat every Error 825 as a page. The read worked this time, but the same sector, controller, or path will eventually return a hard failure. The job now is to identify which file, which volume, and which underlying component is deteriorating before the next read does not recover.

What this means

When SQL Server issues a read, the SQLOS I/O subsystem hands it to the operating system. If the OS returns a failure (bad sector, controller timeout, transport error), SQL Server does not immediately surface it. The engine retries the read up to four times before declaring a hard 823 or 824. If any retry succeeds, the page is returned to the buffer pool, the query continues, and Error 825 is written to the error log.

The message text looks like this:

Error: 825, Severity: 10, State: N.
A read of the file '<path>' at offset <n> succeeded after failing <N> time(s) with error: <OS error text>.
Additional messages in the SQL Server error log and system event log may provide more detail.
This error condition threatens database integrity and must be corrected.
Complete a full database consistency check (DBCC CHECKDB).

Severity 10 is the key gotcha. Microsoft classifies severities 0 through 10 as informational. They do not raise exceptions that flow into TRY…CATCH blocks, and they do not trigger default SQL Server Agent alerts. Without an explicit message_id-based alert on 825, the warning is invisible to operations.

This behavior has not changed since Error 825 was introduced in SQL Server 2005. The same read-retry path, the same severity, and the same alerting gap apply to every currently supported version through SQL Server 2022.

Common causes

CauseWhat it looks likeFirst thing to check
Failing disk sector825 on one file at one offset; OS event log reports bad blockssp_readerrorlog 0, 1, 'Error: 825' and the OS system event log
SAN or storage controller degradation825 across multiple files on the same LUN; latency spikes on the affected filesPer-file I/O stall deltas during the 825 window
HBA, fabric, or cabling fault825 entries clustered in time, possibly with link resets in the OS event logOS event log around the 825 timestamp
Storage firmware bug or MPIO path thrashing825 spikes during array rebalance, path failover, or firmware updatesStorage array logs; correlate with maintenance windows
Cloud disk throttling or transient fault825 on cloud VM storage during IOPS or throughput cap burstsCloud provider throttle metrics for the disk tier
Corrupt page on otherwise healthy media825 followed by an 824 on the same page idmsdb.dbo.suspect_pages for that database and page

Quick checks

Run these in order. All are read-only.

-- 1. Find every 825 in the current error log
EXEC sp_readerrorlog 0, 1, 'Error: 825';

-- 2. Search the most recent archived log if the event rolled over
EXEC sp_readerrorlog 1, 1, 'Error: 825';

-- 3. Check for harder I/O errors in the same window
EXEC sp_readerrorlog 0, 1, 'Error: 823';
EXEC sp_readerrorlog 0, 1, 'Error: 824';

-- 4. Check the suspect_pages table for active corruption records
SELECT database_id, DB_NAME(database_id) AS database_name,
       file_id, page_id, event_type, error_count, last_update_date
FROM msdb.dbo.suspect_pages
WHERE event_type IN (1, 2, 3)
ORDER BY last_update_date DESC;

-- 5. Map each database file to its volume and per-file I/O stall
SELECT
    DB_NAME(vfs.database_id) AS db,
    mf.name AS file_name,
    mf.type_desc,
    mf.physical_name,
    vs.volume_mount_point,
    vfs.io_stall_read_ms,
    vfs.num_of_reads,
    CASE WHEN vfs.num_of_reads > 0
         THEN vfs.io_stall_read_ms * 1.0 / vfs.num_of_reads
         ELSE 0 END AS avg_read_latency_ms,
    vs.available_bytes * 1.0 / vs.total_bytes AS pct_free
FROM sys.dm_io_virtual_file_stats(NULL, NULL) vfs
JOIN sys.master_files mf
    ON vfs.database_id = mf.database_id AND vfs.file_id = mf.file_id
CROSS APPLY sys.dm_os_volume_stats(mf.database_id, mf.file_id) vs
ORDER BY vfs.io_stall_read_ms + vfs.io_stall_write_ms DESC;

How to diagnose it

Goal: identify which file, which volume, and which storage component produced the 825, and confirm whether any page is already damaged.

  1. Capture the 825 entries with timestamps and offsets. Run sp_readerrorlog 0, 1, 'Error: 825'. Each row includes the physical file path, the byte offset, the retry count, and the OS error text. Record the file, the offset, and the time window.

  2. Correlate with sys.dm_io_virtual_file_stats. Snapshot the DMV, wait a measured interval, snapshot again, compute deltas per file. The file named in the 825 should show disproportionate io_stall_read_ms growth relative to its read count. Compare average read latency against known thresholds: under 10ms excellent, 10-20ms acceptable, over 20ms degraded, over 50ms severe.

  3. Check msdb.dbo.suspect_pages. The table records pages that hit an 823 or 824. Error 825 is a success-after-retry, so a clean retry may not add a row. An empty suspect_pages does not rule out active storage degradation. Any new row is a page-level alarm on its own. The table holds a maximum of 1000 rows; when it fills, new errors are not logged until rows are removed, so clear or archive resolved entries.

  4. Run DBCC CHECKDB against the database named in the 825. CHECKDB is I/O-intensive and creates an internal database snapshot that consumes TempDB space, so schedule it carefully. CHECKDB can return clean even when 825 is firing, because the retry path may deliver a correct page on a later attempt. A clean CHECKDB does not contradict an 825. Run it anyway: if a page is already torn or checksum-failed, this is where it surfaces.

  5. Pull the OS system event log for the same window. Storage subsystems usually log the underlying error before SQL Server does. Look for bad-block events, controller resets, MPIO path failovers, or disk firmware messages.

  6. Check for 825 clustering. A single 825 is a warning. A rising rate, several per hour, then several per minute, is a failing component. Trend the count over hours and days. Increasing frequency toward a hard 823 or 824 is the classic signature of SAN controller failure or deteriorating cabling.

  7. Correlate with backup health. 825 events frequently appear alongside backup failures (Msg 3203, Msg 2013) when the storage cannot deliver a clean read of a data file. If backups started failing around the same time, treat the storage subsystem as the prime suspect.

flowchart TD
    A[OS read fails] --> B[SQLOS retries up to 4 times]
    B -->|retry ok| C[Error 825 written
query completes] B -->|all retries fail| D[Error 823 or 824
page unreadable] C --> E[Invisible to default
SQL Agent alerts] E --> F[Investigate file, offset,
OS error, suspect_pages] F --> G{I/O stalls and
suspect_pages correlate?} G -->|yes| H[Storage degradation
confirmed] G -->|no| I[Watch frequency
trend]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Error 825 in error logEarliest reliable indicator of failing storageAny nonzero count in a window
Error 823 (hard I/O error)Read failed all retries; page may be unreadableAny occurrence is PAGE
Error 824 (logical consistency)Checksum or torn page detected after readAny occurrence is PAGE
msdb.dbo.suspect_pages rowsDurable record of pages that hit I/O errorsNew row since last check
Per-file io_stall_read_ms deltaSQL Server’s direct view of storage latencySustained average read latency above 20ms
Per-file io_stall_write_ms delta (log files)Log write latency directly impacts every commitSustained log write latency above 5ms
PAGEIOLATCH_* wait timeWorkers blocked waiting for physical readsRising share of total wait time
Backup job failuresStorage cannot deliver clean reads for backupMsg 3203 in job output
OS storage event logUnderlying controller, disk, and path errorsBad-block, path-reset, firmware events
Error 825 frequency trendPredicts escalation to a hard 823 or 824Count increasing over consecutive windows

Fixes

Error 825 is not a SQL Server configuration problem. The engine is reporting a storage problem. The fixes live below SQL Server.

Contain the risk immediately

  1. Take a full backup before the failing read becomes an unreadable read. A backup taken now is the best recovery point you will have. Verify it with RESTORE VERIFYONLY. If the backup fails with bad-block errors, you have just confirmed a hard storage fault.
  2. Identify and isolate the affected file and volume. If only one database file is on the deteriorating volume, consider whether you can fail over to a secondary (AlwaysOn) or restore the database to a different volume. Do not wait for the next 823.
  3. Engage the storage or hypervisor team. Provide them with the file path, offset, OS error text, and timestamp from the 825 entry. They have visibility into array, controller, and disk health that SQL Server does not.

Address the storage component

  1. Run vendor storage diagnostics against the LUN. Most arrays have a read-only surface scan or SMART-style check that can identify the failing sector without disrupting I/O.
  2. Replace the failing component. Disk, HBA, cable, or controller. Do not assume a single bad sector is isolated; on enterprise storage, a bad sector often indicates a failing disk surface or a degraded controller cache battery.
  3. On cloud storage, reattach the disk to a different host or upgrade the disk tier. Throttling-induced 825 errors on cloud VMs indicate you have hit an IOPS or throughput cap for too long. Moving to a higher tier or spreading I/O across additional disks is the structural fix.

Establish a recovery safety net

  1. Pre-stage a tested restore path. The only way to know your backups are valid is to restore them. Schedule regular restore tests, not just RESTORE VERIFYONLY.
  2. Maintain a baseline of per-file I/O latency. Without a baseline, you cannot tell whether a 15ms average read is normal for that volume or a recent regression.
  3. Confirm PAGE_VERIFY CHECKSUM is enabled on every user database. When the storage does deliver a corrupt page, checksums let SQL Server detect it via Error 824 instead of silently returning wrong data.

Prevention

  • Alert on message_id 825 directly. The default SQL Server Agent alert configuration uses severity thresholds that will never fire on a severity-10 message. Add an alert scoped to @message_id = 825, @severity = 0. The standard reference for this is Brent Ozar’s SQL Server alert script, which includes 825 alongside 823 and 824.
  • Treat any 825 as a page. Severity 10 means informational in the documentation, but operationally it means the storage just failed a read and got lucky. Page on the first occurrence, not the tenth.
  • Keep CHECKDB on a weekly schedule for production databases. Some corruption modes are only detected by CHECKDB. Error 825 may precede them by days or weeks.
  • Track suspect_pages over time. Investigate every new row. Manually remove resolved entries once investigated so new ones are obvious. The 1000-row cap means that once the table is full, new errors are not recorded until rows are removed.
  • Snapshot sys.dm_io_virtual_file_stats externally. The DMV is cumulative since instance startup. Without periodic external sampling, you have no delta to compare against when 825 fires.

How Netdata helps

  • Per-second error log scraping surfaces the first 825 within the same minute it is written, before the next read becomes an unrecoverable 823. Default severity-based alerting misses it; explicit message_id tracking does not.
  • Per-file I/O stall correlation (io_stall_read_ms and io_stall_write_ms deltas) lets you confirm that the file named in the 825 is also the file with climbing latency, ruling out random transient reads.
  • PAGEIOLATCH_* wait time trends show whether storage degradation is starting to block workers, which is the path from soft warning to users reporting the database is slow.
  • suspect_pages row count monitoring catches the moment a soft 825 escalates to a hard 824 and a row lands in msdb.
  • Backup job and duration trending flags the backup failures that often accompany 825 events, before someone needs a restore that does not exist.

Netdata’s Microsoft SQL Server monitoring brings these signals together with per-second metrics and ML anomaly detection.

The Netdata solution

Microsoft SQL Server monitoring with Netdata

Netdata monitors SQL Server with per-second metrics, pre-built dashboards, and ML-powered anomaly detection. Correlate wait statistics, blocking chains, transaction-log and TempDB pressure, Page Life Expectancy, memory grants, per-file I/O stalls, and AlwaysOn send/redo queues against the rest of your stack so you catch the incidents in these runbooks before they page anyone.