The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-health-check-l7sts

Operations Guides

HAProxy health check L7STS: server DOWN on the wrong HTTP status

A backend server flips to DOWN in the stats page, last_chk shows L7STS, and the log line says something like Layer7 wrong status, code: 503, info: "Service Unavailable". The server process is running. The port is open. You can curl it by hand and get a response. But HAProxy has pulled it out of rotation and traffic is concentrating on the survivors.

L7STS means: the TCP connection worked, the HTTP response arrived in time, and the status code was not the one the health check expects. The failure is at the application layer, not the network layer. That narrows the search space considerably, but only if you know what status HAProxy was actually expecting.

This page covers what L7STS means, how it differs from the other last_chk codes, why the server did not go DOWN on the first failed check, and the small set of causes behind almost every L7STS incident in practice. For the wider health-checking picture, see how HAProxy actually works in production.

What this means

HAProxy’s health check engine runs independently of traffic. It probes each server on a schedule and classifies the result into a short code visible in the last_chk stats field, in the stats web UI, and in log lines when a state change happens. The layer-7 codes you will see most often:

CodeMeaningWhat actually happened
L7OKLayer 7 check passedValid HTTP response with an expected status code
L7STSLayer 7 wrong statusValid HTTP response, but the status code was not expected (for example 500, 404, 403)
L7RSPLayer 7 invalid responseSomething arrived, but HAProxy could not parse it as a valid HTTP response
L7TOUTLayer 7 response timeoutConnected, sent the request, no response within the check timeout
L7OKCCheck conditionally passedNon-success status treated as passing by configuration, for example 404 with disable-on-404

L7STS is the interesting one operationally because both ends did their jobs at the protocol level. The kernel accepted the connection, the application read the request and wrote a complete HTTP response, and HAProxy parsed it successfully. The application simply answered with a status the check is configured to reject. Compare:

  • L7RSP points at a protocol mismatch or a broken intermediary: the server answered, but not with parseable HTTP.
  • L7TOUT points at a hung application or an overloaded event loop: the request went in, nothing came back before the timeout.
  • L4TOUT / L4CON never got to HTTP at all: TCP timeout or connection refused.

The key fact about L7STS is that the “expected” status is a matter of configuration, and the default is broader than most operators assume. When no http-check expect directive is configured, HAProxy treats any 2xx or 3xx response as healthy. Anything else, including 401, 403, 404, 405, and every 5xx, produces L7STS.

Common causes

CauseWhat it looks likeFirst thing to check
Application genuinely failingL7STS/500 or L7STS/503, usually on real request paths tooPer-server hrsp_5xx in the stats CSV; if it is elevated, the check is telling the truth
Trivial /health endpoint broke while the app worksL7STS/404 or L7STS/500 on checks, but backend hrsp_5xx is flatHit the exact check URI by hand from the HAProxy host
Check hits the wrong URI or method after a deploySudden L7STS/404 or L7STS/405 on one or all servers, timed with a releaseApplication access log: does the check request line appear, and with what method and path
Auth or host-header gating on the check endpointL7STS/401 or L7STS/403 while browser/manual tests passCompare the exact request HAProxy sends (method, Host header, version) with what the app requires
HTTP version mismatchL7STS/400 or L7STS/426 on backends that reject HTTP/1.0Whether the backend requires HTTP/1.1; the check request version sent by option httpchk
Overly strict http-check expectServer returns a legitimate status (redirect, 204) that the expectation does not coverThe http-check expect line in the running config vs the status the endpoint actually returns

Two of these deserve emphasis because they generate the most confusion:

A 403 on the health endpoint. The server is up, the endpoint exists, but something in the request is missing: a Host header the virtual host routing needs, a token, a source-IP allowlist entry. HAProxy sends a minimal request by default. If your application or an upstream filter gates on any of that, you get L7STS/403 while manual tests with full headers pass.

The default check request itself. When you configure option httpchk with no method or URI, the documented defaults are an OPTIONS request to /. If your application does not answer OPTIONS / with a 2xx or 3xx, for example because it returns 405 for unsupported methods, the check fails on a healthy application. Explicitly configuring the method and URI (option httpchk GET /healthz) removes this whole class of surprise.

Quick checks

All of these are read-only.

# Per-server status, check counters, and last check result in one pass
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" && $2 != "BACKEND" {print $1"/"$2": status="$18" chkfail="$22" chkdown="$23" lastchg="$24"s last_chk="$57}'

The last_chk field carries the detail: L7STS/500 is not the same investigation as L7STS/404. The log line emitted at the state change also includes the code and the reason string.

# Reproduce the check by hand from the HAProxy host
curl -sv -o /dev/null -w "%{http_code}\n" http://<server-ip>:<port>/<check-uri>

Run this exactly as the check runs it: same URI, and if your config sets a Host header or specific method in the check, replicate those. A mismatch between what curl sends and what HAProxy sends is the whole story in a large fraction of L7STS incidents.

# Watch the check arrive from the application side
# (web server access log on the backend)
tail -f /var/log/nginx/access.log | grep healthz

If the check request never appears in the application log, something between HAProxy and the app is answering or dropping it. If it appears with an unexpected method or path, the check configuration is not what you think it is.

# Confirm what HAProxy is actually configured to send
haproxy -c -f /etc/haproxy/haproxy.cfg && grep -E "httpchk|http-check|option.*check" /etc/haproxy/haproxy.cfg

Look for http-check expect lines. If none exist, the expected set is 2xx and 3xx by default.

# Is the application failing real traffic too?
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": 5xx="$44}'

Per-server hrsp_5xx is the cross-check. Elevated 5xx on the same server means the health check is accurately reporting a broken application. Flat 5xx with L7STS means the check endpoint or the check request is the problem, not the app.

How to diagnose it

  1. Read the exact code. Get last_chk for the affected server and note the status code after the slash. L7STS/503 and L7STS/404 are different incidents.

  2. Check the transition state. If the server shows DOWN 2/3 or UP 1/3, it is mid-transition under the rise/fall counters, not fully committed to a state. More on this below.

  3. Correlate with real traffic. Compare per-server hrsp_5xx and frontend versus backend 5xx. If backend 5xx is elevated on the same server, the application is broken and the health check is doing its job. If frontend and backend 5xx are both quiet, the check is failing in isolation.

  4. Reproduce the request manually. From the HAProxy host, send the exact request the check sends: same method, URI, HTTP version, and any configured headers. The status you get back should match the L7STS code. If it does not, the difference between your manual request and HAProxy’s request is the bug.

  5. Check the application log for the check request. Confirm the method, path, and response the server logged. This catches config drift: a deploy that moved the health endpoint, added auth, or started rejecting the check’s HTTP version.

  6. Check timing. A server that alternates between L7OK and L7STS with small lastchg values is flapping: intermittently returning a bad status, usually under load. That points at resource exhaustion or a dependency that fails intermittently, not at a config error.

flowchart TD
  A[Health check fires] --> B{TCP connect OK?}
  B -- no --> C[L4TOUT / L4CON]
  B -- yes --> D{Response before timeout?}
  D -- no --> E[L7TOUT]
  D -- yes --> F{Parseable HTTP?}
  F -- no --> G[L7RSP]
  F -- yes --> H{Status expected?}
  H -- yes --> I[L7OK]
  H -- no --> J[L7STS: check the status code after the slash]

Why the server did not go DOWN immediately

HAProxy does not mark a server DOWN on a single failed check. The fall parameter sets how many consecutive failed checks are required before the status flips; the usual default is fall 3. Likewise rise sets consecutive successes needed to come back UP. While the counters are in progress, the stats show transitional states: DOWN 2/3 means two of the three required failures have happened; UP 1/3 means one of the required successes on the way back.

This has two operational consequences:

  • Detection delay. With fall 3 and a 2-second check interval, a hard failure takes roughly 4 to 6 seconds to pull the server. During that window, traffic still flows to the failing server. If your checks run on a slower interval, the window is proportionally longer.
  • chkfail without DOWN. chkfail increments on every failed check, but the server only transitions after fall consecutive failures. A server with a steadily rising chkfail and a stable UP status is failing intermittently below the threshold. That is flapping territory and worth investigating before it tips over. chkdown counts the actual transitions to DOWN, and lastchg tells you how long the current state has held.

A fast-moving L7STS incident therefore shows up in counters before it shows up in status. Watching chkfail rate and last_chk content gives you earlier warning than watching UP/DOWN alone.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
last_chk per serverThe actual check result with the status code; the fastest way to classify the failureAny persistent value other than L7OK
chkfail rateFailed checks accumulate before status changesIncrementing faster than occasional blips, roughly more than 1 per minute on one server
chkdownCount of real transitions to DOWNIncreasing on one server (flapping) or several at once (shared cause)
lastchgSeconds since the last status changeSmall values recurring on the same server indicate flapping
check_durationHow long the last check tookApproaching timeout check; slow checks delay detection
Per-server hrsp_5xxTells you whether the app is failing real traffic alongside the checkElevated 5xx on the same server showing L7STS
Frontend minus backend 5xx deltaSeparates HAProxy-generated errors from application errorsDelta near zero while a server shows L7STS: the app is the problem
Active server count per backendL7STS servers leaving rotation reduces capacity and can cascade onto survivorsBelow 50 percent of non-MAINT servers UP, or survivors approaching their session limits

The cascade angle matters: every server pulled by L7STS pushes its share of traffic onto the survivors. If survivors were already warm, their scur climbs toward slim, queuing starts, and their response times degrade, which can tip their own health checks over. One bad check configuration can take down a whole backend this way. See HAProxy backend queue building (qcur) and scur approaching slim for the follow-on signals.

Fixes

If the application is genuinely broken. The check is working as designed. Fix the application or its failing dependency; HAProxy will bring the server back automatically after rise consecutive passes. Do not loosen the check to silence the alarm. If flapping is oscillating the server in and out of rotation and making things worse, put it in MAINT via the runtime API while you work, and watch survivor load.

If the check endpoint or request is wrong. Make the check explicit instead of relying on defaults. Configure option httpchk with a specific method and URI that the application actually serves with a 2xx, and align http-check expect with what that endpoint returns. If the backend requires a Host header or a specific HTTP version, set it in the check. The general shape:

backend app
    option httpchk GET /healthz
    http-check expect status 200
    server app1 10.0.0.11:8080 check inter 2s fall 3 rise 2

Validate with haproxy -c -f before reloading, and reproduce the exact request with curl first so you know what status to expect.

If the expectation is too strict. An endpoint that legitimately returns a redirect or a 204 will fail a naive http-check expect status 200. Either widen the expectation to match what the endpoint returns by design, or point the check at an endpoint whose semantics you control.

If the check is fine but the endpoint is trivial. This is the deeper design problem. A /health endpoint that returns a static 200 proves almost nothing: the app can be returning 500 on every real request while the check stays green. The inverse failure also exists, where the trivial endpoint breaks while the app is fine. The durable fix is a health endpoint that exercises the same critical dependencies real requests use, at a cost the check interval can sustain. This “health check green, app broken” blind spot is one of the most common postmortem findings. It is also why an HTTP check beats a bare TCP check: TCP only proves the port accepts connections, which is exactly the failure L7STS exists to catch. The tradeoff is that a deeper check costs more per probe and can fail for dependency reasons that do not affect all request paths equally.

Prevention

  • Make check requests explicit. Always configure method, URI, and expected status. Never rely on the default request shape against an application you do not fully control.
  • Alert on chkfail rate and last_chk content, not just UP/DOWN. You get minutes of warning instead of learning about it when the server leaves rotation.
  • Track per-server hrsp_5xx next to health status. This pair distinguishes “check lies” from “app broken” in one glance.
  • Log the status code on state changes and keep those lines. L7STS/404 in the log at 03:12 is the whole diagnosis.
  • Review rise/fall and inter deliberately. Faster detection means less traffic sent to a failing server, but also more sensitivity to transient blips. Know your detection window: interval times fall.
  • Test the check after every deploy that touches routing, auth, or the health endpoint. A one-line curl from the HAProxy host, replicating the check request, catches most L7STS regressions before the counters do.

How Netdata helps

Netdata collects the HAProxy stats CSV continuously, so the signals above are already correlated on one dashboard instead of requiring manual socket queries during an incident:

  • Per-server health status and check counters (status, chkfail, chkdown, lastchg) plotted over time, so flapping and transition patterns are visible rather than inferred from snapshots.
  • Per-server hrsp_5xx alongside health state, which is the fastest way to separate a genuine application failure from a check misconfiguration.
  • Frontend versus backend 5xx breakdown, exposing the HAProxy-generated error delta when servers leave rotation and 503s start.
  • Survivor saturation signals (scur/slim, qcur) on the same timeline as server DOWN events, so you can see a cascade forming while there is still time to act.
  • Check-related counters at per-second granularity, catching short flaps that minute-resolution polling averages away.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.