The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-server-flapping-up-down

Operations Guides

HAProxy server flapping: rise, fall, and health checks oscillating UP and DOWN

A backend server that alternates between UP and DOWN is worse than one that is cleanly down. Every transition redistributes traffic: when the server goes DOWN its connections move to the survivors, when it comes back UP it absorbs a fresh share of load before it has warmed up, and every cycle costs retries, redispatches, and connection churn. If the flap rate is high enough, the oscillation itself becomes the incident.

The signature is easy to spot in the stats: chkdown keeps incrementing, lastchg (seconds since the last status change) never grows large, and the status field cycles through UP, DOWN 1/3, DOWN, UP 1/3 and back. Meanwhile chkfail climbs even while the server still shows UP, because individual checks are failing below the fall threshold.

This guide covers how to confirm flapping, how to read the health check counters to find the cause, how to stop the churn safely while you investigate, and how to tune rise, fall, inter, and the check timeout so a borderline server stops oscillating.

What this means

HAProxy’s health check engine runs independently of traffic. Each server is probed every inter milliseconds (default 2000). The state transitions are threshold-based:

  • A DOWN server must pass rise consecutive checks (default 2) before it returns to service.
  • An UP server must fail fall consecutive checks (default 3) before it is marked DOWN.

The transitional states appear in the stats as UP 1/3 or DOWN 2/3 style values: the fraction shows progress through the rise/fall counter. Flapping means the server lives on the boundary: it fails fall checks, goes DOWN, the load lifts off it (or the transient fault clears), it passes rise checks, comes back UP, receives traffic again, and fails again. Each cycle is a few inter periods long, so with defaults a full flap can complete in under 20 seconds.

flowchart TD
  UP["UP (serving traffic)"]
  TDN["DOWN n/fall (transitional)"]
  DOWN["DOWN (no traffic)"]
  TUP["UP n/rise (transitional)"]
  UP -->|"fall consecutive check failures"| TDN
  TDN -->|"threshold reached"| DOWN
  TDN -->|"a check passes"| UP
  DOWN -->|"checks pass"| TUP
  TUP -->|"rise consecutive successes"| UP
  TUP -->|"a check fails"| DOWN

Why this matters operationally:

  • Traffic churn. Every DOWN transition redistributes the server’s share of traffic to survivors. If the survivors are near capacity, the extra load pushes their response times up, which can make their own checks slower. This is how one flapping server starts a collapse cascade.
  • Retry noise. Requests in flight when the server drops produce wretr and wredis events and latency outliers, even when users never see an error.
  • Counter semantics hide the real failure rate. chkfail only increments for checks that fail while the server is UP. Once the server is DOWN, further failed checks do not move chkfail. A flapping server can therefore look less broken than it is if you only watch chkfail. chkdown is the counter that tells the truth about oscillation.

Common causes

CauseWhat it looks likeFirst thing to check
Check timeout near actual check latencycheck_duration close to the effective check timeout; last_chk alternating between OK and timeout codesCompare check_duration to timeout check (or inter if timeout check is not set)
rise/fall too tight for a jittery serverServer passes most checks but occasionally fails 2-3 in a row under loadCount check failures vs successes in logs; check whether failures cluster at traffic peaks
Server on the edge of overloadFlapping correlates with traffic peaks; rtime and scur/slim on the server rise before each DOWNPer-server rtime, scur, and application-level CPU/memory on the server
Intermittent packet loss or network instabilitylast_chk shows L4TOUT or L4CON without the server being loaded; other servers on the same path may flap tooctime trend on the backend; retransmits on the HAProxy host (ss -ti), network device counters
Application-level intermittent failurelast_chk shows L7STS or L7RSP (wrong status or wrong response body) while the process stays aliveThe health endpoint itself: does it depend on a downstream (database, cache) that is intermittently slow?
Health endpoint slower than real trafficChecks time out under load while the app still serves requestsWhat the health endpoint actually does; whether it shares a thread pool or connection pool with real requests
Passive checks (observe layer4/layer7) with aggressive on-errorFlapping with real traffic errors feeding back into health stateobserve, error-limit, and on-error settings on the server line

Quick checks

All of these are read-only against the runtime socket. Column numbers below match the standard show stat CSV field order; the first line of show stat output is a header row you can use to verify positions on your version.

# Status, check counters, and last change time for every server
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {print $1"/"$2": status="$18" chkfail="$22" chkdown="$23" lastchg="$24"s last_chk="$57}'

# Detailed per-server state (operational vs configured state)
echo "show servers state" | socat unix-connect:/var/run/haproxy.sock stdio

# Check latency: is check_duration near the check timeout?
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {print $1"/"$2": check_duration="$39"ms"}'

# Is the flapping server also slow or saturated on real traffic?
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {print $1"/"$2": scur="$5" slim="$7" rtime="$61"ms"}'

# Retry and redispatch pressure caused by the flapping
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": wretr="$16" wredis="$17}'

Two things to look at immediately in the first command’s output:

  • lastchg resetting repeatedly. If lastchg never climbs past a few tens of seconds, the server keeps changing state. A stable server has a lastchg that grows for hours or days.
  • last_chk result codes. These tell you the failure mode directly: L4TOUT (TCP connect timeout), L4CON (connection refused), L6TOUT (TLS timeout), L7TOUT (response timeout), L7STS (wrong HTTP status), L7RSP (unexpected response content), L7OK (passing). The flap between an OK code and a specific failure code is your root-cause pointer. If you want just the bare code, check_status (column 37) carries it without the longer description.

How to diagnose it

  1. Confirm it is flapping, not just down. Take two samples of status, chkfail, chkdown, and lastchg a minute apart. chkdown incrementing between samples confirms UP to DOWN transitions are happening right now. chkfail incrementing while status stays UP means the server is failing individual checks below the fall threshold and is one bad stretch away from another transition.

  2. Read last_chk and classify the failure. Timeout codes (L4TOUT, L6TOUT, L7TOUT) point at latency: the check itself, the network path, or an overloaded server. L4CON points at the server refusing connections (process down or at a connection limit). L7STS/L7RSP point at the application answering but answering wrong.

  3. Compare check_duration to the effective check timeout. If timeout check is set, that is the read timeout for checks. If it is not set, inter doubles as the whole check timeout, so a server whose checks occasionally take longer than inter will produce spurious failures. A check_duration hovering near the timeout is the classic borderline case.

  4. Correlate flap timing with load. Look at the server’s rtime and scur/slim around the transitions. If the server flaps only during traffic peaks and its real-request latency climbs right before each DOWN, it is overloaded, and the health check is correctly detecting that. The fix is capacity or per-server maxconn, not check tuning.

  5. Rule out the network. If last_chk shows L4TOUT but the server’s own metrics look idle, check ctime on the backend (rising connect time suggests congestion), and look for retransmits on the HAProxy host with ss -ti. If several servers behind the same network path flap together, suspect the path, not the servers. Simultaneous multi-server flapping is almost always network or a shared dependency.

  6. Inspect the health endpoint itself. For HTTP checks, hit the exact check URL from the HAProxy host repeatedly and time it. If the endpoint touches the database or a cache, it can fail while the application still serves cached or simple requests fine. If it is a trivial static response, the opposite problem applies: the server can pass checks while real requests fail (see Health Check Green, App Broken).

  7. Stop the churn before it hurts the survivors. If the flapping is causing visible traffic churn (rising wredis, survivor scur climbing), put the server in maintenance mode while you investigate:

# Disruptive for that server: stops all new traffic to it. Existing sessions finish.
echo "set server <backend>/<server> state maint" | socat unix-connect:/var/run/haproxy.sock stdio

MAINT stops the redistribution oscillation completely. The survivors get a stable traffic share and you get a quiet system to debug on. Do not forget the server is in MAINT; a forgotten MAINT server is silent capacity loss. Bring it back with state ready when done.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
status (per server)The routing decision; transitional values (UP 1/3, DOWN 2/3) show rise/fall progressStatus changing at all on a normally stable server
chkdownCounts UP to DOWN transitions; the definitive flapping counterAny increment on a server that should be stable; increments on multiple servers at once (escalate)
chkfailIndividual failed checks while UP; leads transitionsClimbing while status still UP: the server is borderline
lastchgSeconds since last state change; a flapping server never lets it growRepeatedly resetting below a few hundred seconds
last_chkThe actual failure code; classifies the root causeOscillation between an OK code and a timeout/refusal/status code
check_durationLatency of the check itselfApproaching timeout check (or inter if unset)
wretr / wredisHAProxy masking the instability from usersRising in step with the flap cycle
Per-server scur/slim, rtimeTells you whether the server is actually overloaded when it failsSaturation or latency spike right before each DOWN
Survivor scur, backend qcurMeasures the blast radius of each redistributionSurvivors approaching their limits during flaps

Fixes

Stop the churn first (MAINT)

If the flapping server is actively destabilizing the backend, put it in MAINT via the runtime API as shown above. This is a traffic-affecting action for that server, so do it deliberately, and track how long it stays in MAINT. It converts an oscillating problem into a stable one you can reason about.

Check timeout too tight: set timeout check explicitly

If check_duration is close to the timeout, the common misconfiguration is a low inter set for fast detection with no separate timeout check. Because inter then serves as the entire check timeout, you have made the detection interval and the failure threshold the same number: any check that merely runs slow counts as failed. Set timeout check to a value comfortably above the p99 of check_duration but below inter, so a slow check fails but a normal one never does. Tradeoff: longer check timeouts slow down detection of genuine failures slightly.

Rise/fall too tight: widen the thresholds

Defaults (rise 2, fall 3, inter 2000) mean a server goes DOWN after roughly 6 seconds of consecutive failure and returns after roughly 4 seconds of success. For a server with jittery but recoverable behavior, that is a hair trigger. Raising fall to 5-10 (or spacing checks with a larger inter) makes a DOWN decision require sustained failure, and raising rise makes re-admission require sustained health. Tradeoff: you accept slower detection of real failures and traffic hitting a genuinely dead server for longer. Tune against your actual failure-detection budget, not against a desire to make the flapping counter quiet.

If many servers share the same check schedule and hardware, spread-checks in the global section adds jitter to check timing and avoids resonance effects where checks synchronize and pile up.

Server genuinely overloaded: fix capacity, not the check

If the server flaps because it is overloaded under its traffic share, health check tuning only masks it. Options: give the server a per-server maxconn so HAProxy queues instead of overwhelming it (see HAProxy maxconn hierarchy), add capacity, or fix whatever makes it slow (rtime decomposition, application profiling). In this case flapping was the health check working as designed.

Intermittent network: fix the path

L4TOUT/L4CON with idle servers and rising ctime means the path between HAProxy and the backend is dropping or delaying packets. Check retransmits on the HAProxy host (ss -ti), the interface error counters, and any middlebox (firewall connection tracking tables are a classic culprit for intermittent refusal). No amount of rise/fall tuning fixes a lossy path; it only changes how loudly HAProxy complains about it.

Passive checks amplifying the problem

If the server uses observe layer4 or observe layer7, real traffic errors feed into the health state via error-limit (default 10) and on-error. The default on-error fail-check simulates a failed check and forces the fast interval, which can accelerate a flap under a burst of real errors. If passive checks are causing oscillation, consider relaxing error-limit or removing observe until the underlying instability is fixed.

Prevention

  • Set timeout check explicitly on every checked server so check timeout is decoupled from inter. Never let a fast inter silently become a short timeout.
  • Choose fall and rise from your detection budget, and make rise strict enough that a server is demonstrably stable before it gets traffic back. A server that just recovered should prove it.
  • Baseline chkdown and lastchg per server and alert on any chkdown increment on servers that are normally stable for weeks. Flapping alone warrants an alert; it becomes urgent when survivors approach saturation.
  • Alert on check_duration approaching the check timeout. That is the earliest possible warning: the server is borderline before a single transition happens.
  • Watch the blast radius. During any flap, track survivor scur/slim and backend qcur. If one server’s DOWN transitions push survivors over 80%, the real fix is headroom, and you are one flap away from a cascade (see HAProxy backend queue building).
  • Make health endpoints check what real traffic needs, but keep them fast. A health endpoint that depends on a slow downstream will flap when the downstream does.

How Netdata helps

  • Per-server chkfail, chkdown, and lastchg as time series turns flapping from a log-grep exercise into a visible oscillation pattern: you see the flap frequency and whether it is accelerating.
  • check_duration trended against the configured timeout surfaces the borderline server days before the first DOWN transition, which is the cheapest time to fix it.
  • Correlating chkdown with wretr/wredis and survivor scur shows the cost of each flap: whether the redistribution is being absorbed cleanly or pushing the backend toward queuing.
  • lastchg never growing is a simple, robust alert condition for oscillation that does not depend on thresholds you have to tune per service.
  • Backend-level correlation (qcur, rtime, 5xx) against flap events distinguishes “server is broken” from “server is the first victim of a shared dependency problem”.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.