The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-response-errors-eresp

Operations Guides

HAProxy response errors (eresp): truncated and invalid backend responses

The eresp counter is climbing on a backend, and clients are seeing broken pages, truncated downloads, or intermittent 502s. The backend servers report healthy. Health checks are green. The 5xx rate, if it moved at all, does not explain the volume of user complaints.

This is the failure mode eresp exists to capture: HAProxy connected to the backend successfully, sent the request, and then received a response it could not treat as valid. Either the bytes were not parseable HTTP, or the server closed the connection before the response finished. That is a protocol-level failure, not an application returning an error status code, and it needs a different investigation path than a 5xx spike.

What this means

eresp is field 15 of the CSV stats (1-indexed, as awk counts columns). It is a cumulative counter on BACKEND and SERVER rows and increments when:

  • The server sent bytes HAProxy could not parse as valid HTTP (junk headers, malformed status line, protocol garbage).
  • The server closed the connection mid-response: a Content-Length body that ends early, a chunked response that never terminates, or a connection drop partway through transfer.

Two semantic traps to know up front:

  • eresp is not a 5xx. A backend that returns a well-formed HTTP 500 increments hrsp_5xx, not eresp. eresp fires when the response never made it through the HTTP layer intact. Depending on when the failure happens, the client may see a 502 from HAProxy, a truncated response, or a dropped connection.
  • eresp is backend/server-scoped. FRONTEND rows do not carry it; request-phase errors on the frontend side are counted under ereq instead. When you are chasing truncated backend responses, look only at BACKEND and SERVER rows.

In TCP mode (mode tcp), there is no HTTP to parse, but eresp can still increment on abnormal server-side closes after a successful connect. The official docs describe eresp as “response errors” and state that srv_abrt (server-aborted data transfers) is counted within it; in TCP mode, a server aborting a transfer falls into this path.

The most common monitoring mistake with this signal is treating it as interchangeable with econ. They are different failure modes: econ means HAProxy never got a connection to the server (refused, timed out, reset during connect), while eresp means the connection worked and the response broke. Lumping them together, or watching only one, hides which side of the failure you are on.

flowchart TD
  A[Request proxied to backend] --> B{Connection established?}
  B -- No --> C[econ increments]
  B -- Yes --> D{Response received intact?}
  D -- Yes, valid HTTP --> E[Normal path - hrsp_ counters]
  D -- Invalid HTTP bytes --> F[eresp increments]
  D -- Closed mid-response --> F
  F --> G{When did it break?}
  G -- Before/early in headers --> H[Client usually sees 502]
  G -- During body transfer --> I[Client sees truncated response]

Common causes

CauseWhat it looks likeFirst thing to check
Backend crash or OOM kill mid-responseeresp and srv_abrt rise together; rtime may drop (fast failures); server restart evidence on the hostBackend logs and dmesg on the server for OOM killer activity
timeout server too low for slow responseseresp correlates with rtime approaching the configured timeout; affects specific slow endpointsCompare rtime distribution against timeout server in the config
Protocol mismatcheresp immediately after a config or backend change; affects all or most requests to one backendConfirm the backend speaks the protocol HAProxy expects (HTTP vs plain TCP, HTTP/1.1 vs HTTP/2 settings)
Network device truncating or corrupting responseseresp spread across servers on a shared path; no backend-side errors; often large responses affectedTest the same request from HAProxy host directly to the server, bypassing middleboxes
Server closing keep-alive connections earlyIntermittent eresp, often on reused connections; may coincide with a drop in connection reuseCompare connect vs reuse counters and check backend keep-alive timeout vs HAProxy’s
TCP mode: unexpected server-side closeeresp in a mode tcp backend; server closes sessions that should stay openServer-side logs for why the session ended

Quick checks

All read-only. Adjust the socket path if yours differs.

# Per-server and per-backend eresp (field 15, 1-indexed)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": eresp="$15}'
# Correlate with econ, wretr, and wredis on the same rows
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": econ="$14" eresp="$15" wretr="$16" wredis="$17}'
# Server aborts: srv_abrt tracks transfers aborted by the server
# and is counted within eresp. cli_abrt is the client-side twin.
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": cli_abrt="$50" srv_abrt="$51}'
# Captured protocol error samples: the single most useful command
# for eresp. Shows recent invalid requests/responses with the
# actual bytes HAProxy choked on.
echo "show errors" | socat unix-connect:/var/run/haproxy.sock stdio
# Is this eresp or a 5xx problem? Compare frontend vs backend 5xx
# (hrsp_5xx is field 44)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '{print $1"/"$2": 5xx="$44}'
# Check whether rtime is approaching timeout server (timeout hypothesis)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "BACKEND" {print $1": rtime="$61"ms qcur="$3}'

Also grep HAProxy logs for the termination state field. Flags like sD (server closed during data transfer), SH/SD (server timeout during headers/data), and -- (clean close) on sessions to the affected backend tell you which side terminated and when. Per-request log timing is more sensitive than the aggregate counters for catching this pattern.

How to diagnose it

  1. Confirm scope and rate. Pull per-server eresp twice, a minute apart, and compute deltas. A counter that went up once and stopped is a past event; a counter incrementing steadily is a live failure. Note whether one server is responsible or all servers in the backend.

  2. Rule out the easy confusion: is it eresp or 5xx? If the frontend 5xx rate tracks the backend 5xx rate, the backends are returning well-formed error responses and eresp is not your problem. If eresp is rising while backend 5xx stays flat, you have a genuine protocol-level failure.

  3. Pull show errors. This is the decisive step. HAProxy keeps a ring buffer of recent protocol violations, including the raw bytes that failed to parse. If the captured response is junk (non-HTTP bytes, binary garbage, an error page from a middlebox), you have a protocol mismatch or a corrupting middlebox. If show errors is empty but eresp climbs, the failures are mid-response closes rather than unparseable bytes, which points at crashes or timeouts.

  4. Check srv_abrt alongside. Server aborts rising in step with eresp points at the server terminating transfers: crashes, OOM kills, application restarts. Check backend application logs and host-level evidence (dmesg for the OOM killer, service restart timestamps).

  5. Test the timeout hypothesis. If rtime (rolling average, last 1024 requests) sits near your timeout server, slow responses are being cut off. Identify whether specific endpoints (large reports, exports, streaming responses) correlate with the eresp timing. A request that legitimately needs 45s against a timeout server of 30s will produce exactly this signature.

  6. Bypass HAProxy to isolate the layer. From the HAProxy host, issue the same request directly to the backend server with curl -v. If the response is complete and valid direct to the server but broken through HAProxy, look at what sits between: middleboxes, and version-specific HAProxy behavior. If the response is truncated even directly, the backend is at fault.

  7. Check connection reuse. If eresp events cluster on what should be reused connections, compare connect vs reuse counters on the backend. A backend that silently closes idle keep-alive connections earlier than HAProxy expects can race with request dispatch and produce mid-response-looking failures. A recent drop in the reuse ratio supports this.

  8. Mind the reload caveat. All cumulative counters reset on reload. If Uptime_sec is low, short-window deltas may mislead. Confirm the process has been up long enough for the trend to be real.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
eresp (per server, delta rate)The core counter for broken/truncated responsesAny sustained nonzero rate
srv_abrtServer-initiated aborts, counted within eresp; closest companion signalRising in step with eresp
econDistinguishes “can’t connect” from “response broke”Rising with eresp suggests deeper backend failure, not just truncation
wretr / wredisHAProxy masking instability via retries before errors surfaceNonzero retries preceding an eresp spike
rtime vs timeout serverTimeout-induced truncationrtime consistently above ~50% of timeout server
hrsp_5xx frontend vs backendSeparates HAProxy-generated errors from backend errorsFrontend 5xx rising while backend 5xx flat
connect vs reuseDetects keep-alive breakdown correlated with mid-response closesSudden drop in reuse ratio
Log termination states (sD, SH, SD)Per-request ground truth on who closed and whenServer-side closes or timeouts clustering on one backend or endpoint

Fixes

Backend crashing or OOM-killed mid-response

Fix the backend, not HAProxy. Find the crash evidence (application logs, OOM killer in dmesg, restart timestamps) and address memory limits or the crash bug. Tradeoff to be aware of: a crash that closes the TCP connection cleanly (FIN) at a response boundary may not increment eresp at all; it shows up as a truncated response to the client with little HAProxy-side evidence. That is why per-request logs matter.

timeout server too low

Raise timeout server to accommodate legitimately slow responses, or move the slow endpoints (exports, reports) behind a separate backend with a longer timeout. Tradeoff: longer timeouts hold backend connection slots longer, which increases queuing risk under load (qcur). If you raise the timeout, watch queue depth. The better long-term fix is usually making the slow operation asynchronous.

Protocol mismatch

Align the modes. If the backend speaks plain TCP, the HAProxy backend must be mode tcp. If a backend was recently changed (new framework, HTTP/2 enabled, TLS added or removed), that change is your suspect. show errors output makes this diagnosis nearly instant because the captured bytes are visibly not HTTP.

Network truncation or corruption

Identify and fix the middlebox, or remove it from the path. This is the rarest cause but the hardest to see from either endpoint in isolation; the bypass test in step 6 is how you prove it.

Keep-alive race on reused connections

Tune the idle timeouts so the side with the shorter patience is the client side of the pair: HAProxy’s backend keep-alive timeout should be shorter than the server’s, so HAProxy closes idle connections before the server can. Verify the effect by watching the reuse ratio and eresp after the change.

Masking with retries

HAProxy’s retries directive and option redispatch can absorb some transient failures before clients see them, and a rising wretr/wredis alongside eresp means that compensation is already happening. Do not treat retries as a fix for eresp: they add latency, they cannot safely retry non-idempotent requests, and a truncated response partway through the body is not a retryable condition. Retries buy time; they do not repair a broken backend.

Prevention

  • Alert on eresp and econ separately. They are different failure modes with different responses. Ratios, not absolutes: eresp as a percentage of requests to that backend, sustained, is the alertable condition.
  • Watch srv_abrt as a leading companion. It often moves before or with eresp when backends are unstable.
  • Keep show errors in your incident runbook. It is the fastest path from “eresp is rising” to “here are the actual bytes that broke.”
  • Size timeout server against measured response times, including the slow legitimate endpoints, and revisit when the application changes.
  • Correlate eresp with deploys and config changes. Protocol mismatches and keep-alive races almost always trace back to a recent change on one side of the connection.
  • Log termination states and per-request timing. The aggregate counters tell you something broke; the logs tell you which requests, which endpoints, and who closed first.

How Netdata helps

  • Netdata collects the HAProxy CSV stats per server and per backend, so eresp, econ, srv_abrt, and wretr/wredis are charted on the same timeline, which makes the “connect failure vs response failure” split visible at a glance.
  • Counter resets on reload are handled automatically, so delta-based eresp rates stay meaningful across config reloads that would otherwise produce phantom spikes.
  • Per-server breakdowns let you see immediately whether eresp is isolated to one backend server or spread across the pool, which is the first fork in the diagnostic path.
  • Correlating eresp with rtime and qcur on the same dashboard surfaces the timeout-cascade pattern (slow responses being cut off, then queuing) without manual log digging.
  • Anomaly detection on a normally-zero counter like eresp catches the low, steady drip of mid-response failures that threshold alerts miss until users complain.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.