The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-backend-losing-servers-capacity

Operations Guides

HAProxy backend losing servers: active server count and cascade risk

A backend can lose half its servers and HAProxy will still report the BACKEND aggregate row as UP. As long as one server is available, the rollup status says traffic is being served, and status-based dashboards stay green. Meanwhile the surviving servers are absorbing redistributed load, their session counts are climbing, and you are one health check failure away from a full 503 outage for that backend.

This is the leading edge of the backend collapse cascade: initial server failure, traffic redistribution, survivor overload, survivor health check failures, more redistribution, total collapse. The operators who catch it early are watching the active server count and the load on survivors, not the aggregate status.

This article covers how to read the act/bck counts and per-state server rows, how to judge whether a partial loss is dangerous for your specific pool, and what to do before the cascade completes.

What this means

On the BACKEND row of show stat, two fields describe how much of the pool is actually serving traffic:

  • act (CSV field 19): number of active servers currently UP in the backend.
  • bck (CSV field 20): number of backup servers currently UP.

Both only count servers HAProxy considers available. A server in DOWN, MAINT, or DRAIN does not appear in either number. To get the full picture, including how many servers are DOWN versus intentionally removed, you have to count the individual server rows by their status field. That distinction matters operationally: a backend with 6 configured servers showing act=3 is a very different situation if the 3 missing servers are DOWN (failure) versus MAINT (planned work someone forgot to close out).

The cascade mechanism works like this: servers go DOWN, HAProxy redistributes their traffic to survivors, survivors slow down or hit their per-server maxconn, their health checks start timing out or the queue grows, more servers get marked DOWN, and the loop accelerates. The BACKEND status only flips to DOWN when the last available server goes, at which point every request to that backend gets a 503 (or a TCP reset in TCP mode). Everything before that flip is invisible if you only alert on status.

flowchart TD
  A[Server fails health check] --> B[HAProxy marks it DOWN]
  B --> C[act decreases, traffic redistributes]
  C --> D[Survivor scur rises, rtime increases]
  D --> E{Survivors overloaded?}
  E -- no --> F[Degraded but stable: fix the failed server]
  E -- yes --> G[qcur grows, survivor checks flap]
  G --> H[More servers marked DOWN]
  H --> C
  H --> I[act reaches 0: BACKEND DOWN, 503 for all requests]

The key decision point is “survivors overloaded?”. Losing 50% of a pool that was running at 20% utilization is a ticket. Losing 50% of a pool that was running at 70% utilization is the first move of an outage.

Common causes

CauseWhat it looks likeFirst thing to check
Application failure on individual serversOne or two servers DOWN, check_status shows L7STS or L7TOUT, others healthyPer-server status and check_status; app logs on the failed servers
Network partition to a subset of serversMultiple servers DOWN nearly simultaneously, check_status shows L4TOUT or L4CON, chkdown increments togetherWhether the DOWN servers share a rack, subnet, or dependency; econ on those rows
Slow dependency causing survivor overload, then check failuresFirst one server DOWN for a real reason, then others follow as load concentrates; rtime and qcur rising on survivorsShared dependency health (database, cache); per-server rtime comparison
Rolling deploy with a broken releaseServers fail in sequence as they receive the new version; lastchg recent and staggeredDeploy timeline; whether a rollout is in progress
Servers in MAINT or DRAIN reducing capacity silentlyact low but no DOWN rows; servers sitting in MAINT/DRAIN from old maintenancePer-state server counts; how long MAINT servers have been in that state
Health check misconfigurationHealthy application, but checks failing; chkfail rising, check_status showing unexpected codesCheck endpoint and expected status; see the L4TOUT/L4CON and L7STS guides

Quick checks

All of these are read-only and safe to run during an incident. They assume the admin stats socket at /var/run/haproxy.sock; adjust the path for your deployment. The field numbers below use HAProxy’s 0-indexed CSV positions, so act (field 19) is $20 in awk.

# Active and backup server counts per backend
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "BACKEND" {print $1": active="$20" backup="$21}'

This is the fastest single check. Compare active against how many servers the backend should have. If you do not know the expected count, the next command gives you the full breakdown.

# Per-state server counts per backend (UP, DOWN, MAINT, DRAIN, transitions)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {backends[$1]++; states[$1","$18]++} END {for (b in backends) {printf "%s: total=%d", b, backends[b]; for (k in states) {split(k,a,","); if(a[1]==b) printf " %s=%d", a[2], states[k]}; print ""}}'
# Individual server status, check failures, and last check result
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {print $1"/"$2": status="$18" chkfail="$22" chkdown="$23" lastchg="$24"s check_status="$37}'

check_status tells you why a server is DOWN: L4TOUT (TCP timeout), L4CON (connection refused), L7STS (wrong HTTP status), L7TOUT (response timeout). lastchg is seconds since the last status change, which tells you whether this is a fresh failure or a long-standing one.

# Are the survivors overloaded? Per-server session utilization
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {if($7>0) print $1"/"$2": scur="$5" slim="$7" pct="int($5/$7*100)"%"}'
# Queue and latency on the backend
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "BACKEND" {print $1": qcur="$3" qmax="$4}'

# Retries and redispatches: is HAProxy masking instability?
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "BACKEND" {print $1": wretr="$16" wredis="$17}'

How to diagnose it

  1. Establish expected versus actual server count. Run the act/bck check and the per-state count. You need three numbers: total configured servers, servers in each state, and how many are actually serving. If the gap is MAINT or DRAIN, find out who put them there and when, because forgotten maintenance is a capacity risk, not a failure.

  2. Classify the failure from check_status. L4TOUT/L4CON means HAProxy cannot reach the server at the TCP level (network, dead process, firewall). L7STS means the server is reachable but the health endpoint returns the wrong status (application-level). L7TOUT means the server accepts connections but responds too slowly, often the first sign of overload rather than a hard failure.

  3. Check the timing pattern. Simultaneous chkdown across servers with recent lastchg points at a shared cause: network partition, shared dependency, or a deploy. Staggered failures in sequence point at a cascade already in progress.

  4. Measure survivor load. Pull per-server scur/slim for the remaining UP servers, plus backend qcur and qtime. This is the step that determines severity. If survivors are above roughly 80% of their session limit, or qcur is nonzero and growing, the backend is at capacity and the next failure is the outage.

  5. Check whether HAProxy is already compensating. Rising wretr/wredis means connections to servers are failing and being retried or redispatched. Users may see nothing yet, but the retry counters tell you the pool is unstable before errors become visible. High retries with rising 5xx means compensation is failing.

  6. Find the first failure. In a cascade, the first server that went DOWN usually holds the root cause. The later failures are often just overload. Compare lastchg values to identify which server failed first, and investigate that server’s application and its dependencies.

  7. Confirm or rule out the shared dependency. If rtime was climbing on all servers before the first failure, suspect the database, cache, or downstream API rather than the servers themselves. Servers drowning in slow dependency responses will fail health checks (L7TOUT) even though their own process is healthy.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
act on BACKEND rowsDirect measure of serving capacity; the count the aggregate status hidesBelow 50% of non-MAINT/DRAIN servers, or trending down
Per-state server counts (UP/DOWN/MAINT/DRAIN)Distinguishes failures from intentional removal; MAINT forgotten for weeks is latent capacity lossAny DOWN; MAINT/DRAIN older than the planned window
Survivor scur/slimTells you whether the remaining pool can absorb redistributed loadSurvivors above 80% of their session limit
qcur / qtime on the backendFirst sign the surviving pool is saturated; appears before errorsAny sustained nonzero value
rtime on survivorsOverloaded survivors slow down before they fail checksClimbing, or rising after each server loss
wretr / wredisHAProxy masking instability; retries precede visible errorsSustained nonzero, especially with rising 5xx
chkfail / chkdown / lastchgFlapping detection and failure timelinechkfail incrementing on UP servers; recent lastchg
Frontend hrsp_5xx minus backend hrsp_5xxSeparates HAProxy-generated 503s from backend-returned errorsDelta rising: HAProxy running out of servers to route to
check_status per serverThe actual failure reasonL4TOUT/L4CON (reachability), L7STS (app status), L7TOUT (overload)

The escalation rule of thumb: fewer than 50% of non-MAINT/DRAIN servers active is high cascade risk if survivors are loaded, and one server remaining means any single failure is a total outage. But the 50% line is not universal. An overprovisioned pool running at 15% utilization can lose half its servers and be boring. Always pair the count with survivor scur/slim before deciding severity.

Fixes

Restore the failed servers

The correct fix is almost always to bring the failed servers back, not to route around their absence. Work from the check_status classification: TCP-level failures get network and process investigation on the server; L7STS gets application investigation. If a deploy is in progress and servers are failing as they receive new code, roll back first and diagnose after.

Stop flapping servers from churning the pool

If a server is oscillating UP/DOWN, each transition redistributes traffic and each recovery dumps load back onto a server that may not be ready. Put genuinely broken servers into MAINT via the runtime API (set server <backend>/<server> state maint) so the pool stabilizes while you investigate. This is a deliberate, reversible operator action that takes the server out of rotation immediately; document who did it and why, because MAINT servers are invisible to DOWN alerts. Restore with set server <backend>/<server> state ready.

Relieve the survivors while you work

If survivors are near saturation, reduce their load rather than adding more: shed nonessential traffic at the frontend if you have ACL-based controls, or shift traffic to another backend or site if one exists. Do not raise per-server maxconn as a reflex; the per-server limit exists to protect the application, and raising it on already-overloaded servers converts queuing at HAProxy into collapse at the application.

Use backup servers and capacity-aware routing

If your backends define backup servers, HAProxy uses them only when all active servers are DOWN by default; option allbackups changes that to use all backups simultaneously. For earlier, capacity-based failover, the nbsrv(<backend>) sample fetch returns the number of available servers and can drive an ACL that switches traffic before the pool empties, for example routing to an overflow backend when nbsrv falls to 2 or below. Both patterns are described in the official HAProxy documentation and blog; test them before you need them, and note that stick-table persistence can keep sessions pinned to a backup server even after the active servers recover.

The all-down hazard

When every server in a backend is DOWN, HAProxy returns 503s, but the frontend keeps accepting connections and sessions can pile up in queues and connection state. A known issue (haproxy/haproxy#1946) documents severe CPU spikes under high connection volume in exactly this state, recovering as soon as one server comes back. If a backend is fully down under load, consider draining traffic away from the frontend, not just waiting for servers to return.

Prevention

  • Alert on the count, not the status. BACKEND status UP/DOWN alone misses everything before total outage. Alert when act drops below a defined floor per backend, and page when the composite condition holds: more than 50% of non-MAINT/DRAIN servers DOWN, survivors overloaded (scur/slim > 80% or qcur > 0), persisting for more than a couple of minutes.
  • Size for N+1 at minimum. The capacity question to answer per backend: with one server removed, can the rest serve peak traffic without queuing? If not, your redundancy is nominal.
  • Set per-server maxconn deliberately. It is what turns survivor overload into visible queuing (qcur) instead of silent application collapse, and it makes the cascade detectable in HAProxy’s own metrics.
  • Make health checks verify the real application. TCP-only checks or trivial /health endpoints let broken servers stay in rotation, which makes server count changes abrupt when they finally happen. Checks that exercise real dependencies fail earlier and more informatively.
  • Track MAINT and DRAIN duration. Servers parked in MAINT for weeks silently shrink every pool they belong to. Audit them on a schedule.
  • Watch retries as the early-warning channel. A backend whose wretr/wredis counters climb for days before a server finally goes DOWN gave you the warning; most teams just never look.

How Netdata helps

  • Netdata collects the HAProxy stats CSV continuously, so act/bck per backend and per-server UP/DOWN state become time series rather than point-in-time show stat snapshots. You can see the pool eroding, not just discover it after the fact.
  • Per-server scur/slim, qcur, and rtime are charted alongside the server counts, which makes the key correlation (count down, survivor load up) a single glance instead of two manual queries during an incident.
  • Retry and redispatch counters (wretr/wredis) are surfaced as rates, so the “HAProxy is masking instability” phase of a cascade is visible in a dashboard rather than only in a runtime socket dump.
  • Health check transition signals (chkdown, chkfail, status changes) let you reconstruct the failure sequence during post-incident review: which server went first, and how long the cascade took.
  • Alerting on ratios (survivors above a percentage of their session limit, active servers below a fraction of the pool) matches how cascade risk actually behaves, and Netdata’s per-second collection catches the fast transitions that minute-resolution polling misses.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.