The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / vmware-vsphere / vmware-vsphere-host-hardware-health

Operations Guides

vSphere ESXi hardware health: ECC errors, fan failure, and thermal throttling

Hardware faults do not surface through hypervisor performance counters. They appear either as a Purple Screen of Death (PSOD) that kills every VM on the host instantly, or as slow degradation that looks like software misconfiguration until someone checks temperature sensors. Every DIMM error, failed fan, and degraded disk is the hypervisor’s problem.

Three failure classes dominate production incidents on ESXi: ECC memory errors that progress from correctable to uncorrectable, fan failures that trigger thermal throttling and silently cap CPU frequency, and predictive disk failures that kick off RAID rebuilds consuming storage I/O. All three share a common operational trap: without the vendor CIM provider VIB installed, most sensors report “unknown” or are absent entirely, leaving you blind to the root cause.

The CIM stack is also changing. As of ESXi 8.0, CIM and SLP services are deprecated, with informational VOBs warning that removal is planned for the next major release. The sfcbd CIM server has been disabled by default on fresh installs since ESXi 6.5 and only starts automatically when a third-party CIM VIB is present. If you are running ESXi 8.0 Update 3 or later and relying on CIM sensor data, you need a plan for what replaces it.

What this means

Fault classEarly signalEscalation pathEnd state
ECC memory (correctable)MCE entries in vmkernel.log with “CE Poll” on a specific channelCE rate climbs past 10/hour on same DIMMPSOD: “NMI IPI: Panic requested by another PCPU”
Fan failure / thermalFan RPM drops or reads 0 in IPMI SDRCPU temperature climbs past thresholdSilent MHz reduction, VMs slow, host CPU util looks high
Predictive disk failureSMART attributes flag degraded diskRAID controller marks drive predictive failedRebuild consumes IOPS, datastore latency spikes across all VMs

The critical pattern with ECC errors is progression. Correctable errors (CE) are the leading indicator. Broadcom KB 414711 documents a three-stage failure: corrected memory errors appear in vmkernel.log, then CPU heartbeat failures occur (“PCPU didn’t have a heartbeat for X seconds”), then a PSOD follows with “NMI IPI: Panic requested by another PCPU.” The window between stage 1 and stage 3 can be minutes to hours. When you see 10 or more CE errors within a few minutes on the same memory channel, that DIMM is about to fail.

Thermal throttling is the stealthiest of the three. ESXi does not log CPU throttling events. The CPU simply runs at a reduced frequency. The only indirect indicators are temperature sensors reading high and elevated host CPU utilization without corresponding VM throughput. Teams chase CPU contention, check ready time, inspect limits, and find nothing wrong because the problem is physical, not scheduling.

flowchart TD
    A[Physical hardware fault] --> B{Sensor class}
    B -->|DIMM ECC| C[MCE in vmkernel.log]
    B -->|Fan failure| D[CPU temperature rising]
    B -->|Disk SMART| E[Predictive failure flag]
    C --> F{CE rate >10/hr same channel?}
    F -->|Yes| G[Heartbeat loss then PSOD]
    D --> H[Thermal throttling active]
    H --> I[Silent MHz cap, VMs slow]
    E --> J[RAID rebuild starts]
    J --> K[Datastore latency spike]

Common causes

CauseWhat it looks likeFirst thing to check
Failing DIMMCE errors in vmkernel.log on specific channel, escalating rategrep -i "MCE|MCA" /var/log/vmkernel.log
Fan or PSU failureIPMI SDR shows fan RPM 0 or PSU status non-OKesxcli hardware ipmi sdr list
Blocked airflow / dustTemperature sensors high, no individual fan failedPhysical inspection, BMC/iDRAC/iLO fan map
Degraded diskSMART Reallocated_Sector_Ct or similar attribute climbingesxcli storage core device smart get -d <device>
Missing vendor CIM VIBAll sensors show “unknown” in Hardware Healthesxcli software vib list | grep -i cim
Unsupported BIOS after upgradeSensors flap between red and green after ESXi 8.x upgradeCheck HCL BIOS version, update firmware
Power management throttlingCPU frequency reduced, no thermal alarmVerify “High Performance” power policy on host

Quick checks

All commands are safe and read-only. Run them over SSH or via the ESXi shell.

# Read all IPMI sensor values: temperatures, fan speeds, voltages, PSU status
esxcli hardware ipmi sdr list

# Check memory ECC error summary
esxcli hardware memory get

# Search vmkernel.log for Machine Check Exception entries
grep -i "MCE\|MCA\|Machine Check" /var/log/vmkernel.log | tail -50

# Search for PCPU heartbeat loss (PSOD precursor)
grep -i "heartbeat" /var/log/vmkernel.log | tail -20

# Get SMART data for a specific disk device
esxcli storage core device smart get -d naa.600508b1001c3d0

# Get hardware platform summary (vendor, model, BIOS version)
esxcli hardware platform get

# Check whether the CIM broker service is running
/etc/init.d/sfcbd-watchdog status

# List installed vendor CIM provider VIBs
esxcli software vib list | grep -i "cim\|openmanage\|ssa\|hp-health\|cimc"

If the IPMI SDR list returns nothing or reports “unknown” for every sensor, you are likely missing the vendor CIM provider VIB. Generic IPMI sensors (CPU temp, fan speed) may still appear via the /dev/ipmi driver that hostd uses directly, but storage controller health and vendor-specific sensors require the VIB.

How to diagnose it

1. Confirm sensor visibility first. Before investigating any hardware fault, verify that your host actually reports sensor data. Run esxcli hardware ipmi sdr list and check whether fan speeds, temperatures, and PSU status populate. If they do not, install the vendor CIM provider before proceeding. Without sensors, you are guessing. Also confirm the sfcbd-watchdog service status, since the CIM broker must be running for provider data to flow.

2. For suspected ECC memory issues: Search vmkernel.log for MCE entries. Correctable errors appear with a format like MCA: [PCPU]: CE Poll G[status] B[bank] S[status] A[address] M[misc] Memory Controller Read Error on Channel [X]. Count occurrences per channel. If you see 10 or more CE entries within a few minutes on the same channel, that DIMM is failing. Look for the progression to heartbeat failures, which is the immediate PSOD precursor.

3. For suspected thermal throttling: Check IPMI temperature sensors. If CPU temperature is near or above the vendor threshold, check fan speeds. A single failed fan with redundant cooling is a P3; schedule replacement. All fans degraded, or a single fan failure on a non-redundant platform, is an emergency. Reduce host load immediately. Cross-reference with host CPU utilization: if utilization is high but VM throughput is low and ready time is normal, thermal throttling is the likely cause. Also verify the host power policy is set to “High Performance” rather than a balanced or low-power policy, which can independently cap CPU frequency and mimic throttling symptoms. On builds where the shell reports the package frequency, compare esxcli hardware cpu global get Current Speed with Max Speed; verify it against BMC temperatures and fans.

4. For suspected disk degradation: Pull SMART data with esxcli storage core device smart get -d <device-id>. Look for climbing reallocated sector counts, pending sectors, or media errors. If a disk is flagged predictive failed by the RAID controller, a rebuild may already be running. Check storage latency during this period, as rebuild I/O competes with VM workloads.

5. Rule out false positives. On ESXi 6.5 and 6.7, some IPMI sensors intermittently flip from green to red, generating false hardware health alarms. This was resolved in ESXi 6.5 U3 and 6.7 U2. If you are on an affected version, use esxcfg-advcfg -s <sensor-ID>[,<sensor-ID>...] /UserVars/HardwareHealthIgnoredSensors to ignore specific sensors, using the Node-Sensor ID suffix from esxcli hardware ipmi sdr list. Separately, after upgrading to ESXi 8.x, sensor flapping between red and green is often caused by an unsupported BIOS version sending corrupted IPMI data. The fix is a BIOS update to the version listed in the VMware HCL.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
ECC CE error count per channelCorrectable errors are the PSOD precursor>10/hour on same channel sustained
PCPU heartbeat failures in vmkernel.logImmediate PSOD precursorAny “PCPU didn’t have a heartbeat” entry
CPU temperature (IPMI)High temp triggers thermal throttlingTrending toward vendor TCC threshold
Fan RPM per fan (IPMI)Fan failure removes cooling capacityRPM at 0 or below minimum threshold
PSU status (IPMI)Lost redundancy reduces failure toleranceNon-OK status on any PSU
SMART reallocated/pending sectorsPredictive disk failure indicatorCount increasing over time
Host CPU utilization vs VM throughputThrottling shows high util, low throughputUtilization climbing without ready time or workload change
sfcbd-watchdog memory usageCIM broker memory leak can trigger host swappingRSS growing unbounded over days
Sensor state changes per hourSensor flapping indicates firmware or BMC issue>5 state changes/hour on stable hardware

Fixes

Correctable ECC errors climbing toward uncorrectable

When CE errors appear in vmkernel.log, identify the failing DIMM from the MCE address and channel information. The bank and address fields in the MCE entry map to a physical DIMM slot per your server vendor’s documentation. Schedule a DIMM replacement during a maintenance window.

If the CE rate exceeds 10 per hour on the same channel, evacuate VMs off the host immediately via vMotion and place the host in maintenance mode before the PSOD occurs. Once a PSOD happens, every VM on the host crashes simultaneously. Evacuating before the uncorrectable error is strictly better than recovering after.

Fan failure causing thermal throttling

Check the BMC directly (iDRAC, iLO, IMM, XCC) for fan status and any thermal alerts the ESXi CIM layer may not surface. If a single fan has failed on a platform with redundant cooling, you have time but should schedule replacement. If cooling is non-redundant or multiple fans are degraded, reduce host load immediately.

Thermal throttling is not something you fix in software. The CPU is protecting itself by reducing frequency to stay below the thermal limit. The fix is physical: replace the failed fan, clean dust from heatsinks and airflow paths, verify the chassis is not in an environment exceeding its intake temperature spec, and reseat heatsinks if thermal paste has degraded.

Predictive disk failure and RAID rebuild impact

If SMART data shows a disk heading toward failure, proactively replace it before the RAID controller marks it failed and initiates a rebuild. A rebuild on a busy datastore saturates IOPS and spikes latency for every VM on that LUN or disk group.

If a rebuild is already running, expect elevated storage latency (DAVG) across all VMs on the affected datastore until it completes. Consider scheduling the rebuild during off-peak hours if the controller supports it. Monitor datastore latency and queue depth during the rebuild window. See vSphere datastore latency high: reading GAVG, DAVG, and KAVG for latency decomposition.

Missing sensor visibility (no vendor CIM VIB)

Install the vendor-specific CIM provider VIB for your hardware. Dell OpenManage, HP SSACLI or hp-health, and Cisco CIMC are the common providers. Without the appropriate VIB, most sensors show “unknown” and you lose all hardware health visibility.

Note that Dell dropped CIM provider support for LSI-based PERC controllers in their ESXi 6.7 image. VMware also blocked the LSI CIM provider from loading. If you are running ESXi 6.7 or later on Dell hardware with local RAID, disk health monitoring for the RAID controller may be broken regardless of VIB installation. In that case, monitor disk health through the iDRAC or OpenManage Server Administrator instead of through ESXi CIM.

Prevention

  • Install vendor CIM VIBs on every host at deploy time. Without them, sensor data is absent and hardware faults are invisible until PSOD.
  • Monitor ECC CE error rate per channel. Alert when any channel exceeds 10 errors per hour. This gives you hours of warning before uncorrectable failure.
  • Track CPU temperature trends, not just thresholds. A gradual upward trend over weeks indicates degrading cooling (dust, paste, fan wear) before any single threshold trips.
  • Verify host power policy is “High Performance.” Balanced or low-power policies cap CPU frequency independently of thermal events and mimic throttling symptoms.
  • Forward vmkernel.log to a central syslog. MCE entries and heartbeat failures are the earliest hardware failure signals. If logs live only on the host, a PSOD may erase them.
  • Plan for CIM deprecation on ESXi 8.0+. CIM and SLP are deprecated with removal planned for the next major release. Evaluate vendor-specific monitoring paths (iDRAC, iLO, XCC) before CIM disappears.
  • Patch BMC firmware proactively. A known BMC issue caused multiple hardware sensor alert floods on ESXi 8.0 U3; the ESXi-side alert-flood issue was resolved in 8.0 U3g, and update the BMC firmware as well. BMC firmware bugs produce sensor noise that erodes alert trust.

How Netdata helps

  • Correlates hardware sensor data with VM performance counters. When CPU temperature spikes and host utilization climbs but VM throughput drops, Netdata’s per-second metrics let you see the thermal throttling signature in a single view rather than stitching together separate tools.
  • Surfaces ECC error rate trends. Monitoring correctable error counts per channel over time catches the progression toward uncorrectable failure hours before a PSOD, giving you time to evacuate VMs.
  • Detects anomalous CPU utilization patterns. ML-based anomaly detection flags the specific pattern of high utilization without corresponding throughput that indicates thermal or power-management throttling, distinguishing it from genuine CPU contention.
  • Correlates predictive disk failures with storage latency. When a SMART attribute starts climbing and datastore latency spikes simultaneously, Netdata shows both signals aligned in time, confirming the rebuild I/O impact across all VMs on the datastore.
  • Tracks sfcbd-watchdog resource usage. The CIM broker memory leak is a known cause of host swapping. Monitoring its RSS catches the leak before it degrades management services.
The Netdata solution

VMware vSphere monitoring with Netdata

Netdata auto-discovers vCenter, ESXi hosts, VMs, and datastores through the vSphere API and collects them per second with ML-powered anomaly detection. Correlate CPU ready and co-stop, ballooning and host swap, datastore latency, and snapshot growth against the host and guest signals behind them, so you catch the incidents in these runbooks before they page anyone.