The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-monitoring-checklist

Operations Guides

HAProxy monitoring checklist: the signals every production proxy needs

Most HAProxy outages are visible in its own stats long before users notice. The problem is that the stats socket exposes dozens of fields, and teams usually wire up the five metrics their dashboard template shipped with and stop there. Then a reload storm, a retry storm, or a silent stick-table overflow teaches them what they were missing.

This checklist organizes the signals that catch production failures into four maturity levels: survival, operational, mature, and expert. Each level builds on the previous one. You do not need to reach expert on day one, but you should know which level you are at, because that determines which failure modes are invisible to you.

Everything here is collectable from the runtime API (stats socket), the CSV stats, or the host OS. No agents inside HAProxy are required.

How to use this checklist

Work down the levels in order. For each signal, confirm three things: you are collecting it, you are alerting on the right condition, and you have verified the alert fires (reload counters reset, thresholds tuned as ratios, not absolutes).

The canonical collection pattern:

# Process and event-loop liveness
echo "show info" | socat unix-connect:/var/run/haproxy.sock stdio | head -1

# CSV stats: per-proxy rows (FRONTEND, BACKEND, and per-server)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio

# Loaded certificate list (HAProxy does not expose expiry; check cert files on disk)
echo "show ssl cert" | socat unix-connect:/var/run/haproxy.sock stdio

# FD usage vs limit for the running worker
HPID=$(pgrep -x haproxy | head -1)
echo "FDs used: $(ls /proc/$HPID/fd 2>/dev/null | wc -l)"
grep 'Max open files' /proc/$HPID/limits

The stats socket path varies by deployment (/var/run/haproxy.sock is a common default, not a guarantee). All socket commands above are read-only and safe on a production process. Do not run runtime commands that change state (disable server, set maxconn, shutdown session) as part of a monitoring check.

flowchart TD
  L1["Level 1 - Survival: is it up and serving?"] --> L2["Level 2 - Operational: is it healthy?"]
  L2 --> L3["Level 3 - Mature: is it degrading?"]
  L3 --> L4["Level 4 - Expert: what is it hiding?"]

Level 1: survival

The absolute minimum. If any of these fire, users are already affected.

SignalSourceWhy it mattersWarning sign
Process and stats-socket livenesspgrep -x haproxy plus show info over the stats socketA PID check alone lies: the process can exist while the event loop is stuck or drainingNo active worker, or no socket response for > 15s with Uptime_sec > 30
Backend aggregate statusCSV status on BACKEND rowsA DOWN backend means 100% of requests to it get 503 (or TCP reset in TCP mode)Any production backend DOWN
Frontend session rateCSV rate on FRONTEND rows, SessRate in show infoProves traffic is actually arriving. Zero rate on a production frontend means something upstream broke or HAProxy is rejectingRate at zero or far below the time-of-day baseline
Frontend 5xx rateCSV hrsp_5xx on FRONTEND rows, delta-basedThe user-visible error rateSustained 5xx above ~1% of req_rate
Current sessions vs limitCurrConns / Maxconn in show info, scur / slim on CSV rowsAt 100% of maxconn, new connections queue or are rejected. This is the single most common HAProxy incidentscur sustained above 90% of slim

Two liveness subtleties that bite in practice:

  • Master-worker mode. The master process stays alive even while workers restart. Monitor the worker, not the master. After a reload, multiple PIDs are normal; only one is the active worker.
  • Cold start gating. Gate availability alerts on Uptime_sec so restarts and reloads do not page you. A backend that is DOWN during the first minutes of process life may just be waiting for health checks to pass their rise threshold.

Level 2: operational

What a competent production team should have. This is where you start detecting problems before users do.

SignalSourceWhy it mattersWarning sign
Per-server UP/DOWN statusCSV status on server rowsThis is the routing table. Each DOWN server concentrates load on survivorsAny server DOWN
Active/backup server countCSV act / bck on BACKEND rows, or count server rows by statusA backend can be “UP” with one of five servers left. Aggregate status hides capacity lossFewer than 50% of non-MAINT/DRAIN servers active
Queue depth and queue timeCSV qcur (gauge), qtime (rolling average over last 1024 requests)Queuing is the earliest indicator of backend saturation, appearing before errors and before timeoutsAny sustained nonzero qcur; qtime consuming a meaningful fraction of timeout server
Response timeCSV rtime on BACKEND and server rowsThe backend application’s processing time as seen by HAProxySustained rtime beyond 2x baseline, or above 50% of timeout server
Connection errorsCSV econ on BACKEND and server rowsHAProxy tried to reach a backend and failed: refused, timed out, or resetSustained econ rate above ~1% of connection attempts
Retries and redispatchesCSV wretr / wredisThe canary most teams miss. HAProxy is masking backend instability from users right nowAny sustained nonzero rate; above 1% of requests is significant instability
Bytes in / outCSV bin / boutBandwidth and anomaly detection. bout/bin inverting on a web app is worth a lookSustained approach to NIC capacity

Three things to internalize at this level:

  • wretr/wredis are pre-failure signals. High retries with zero user-visible errors means the compensation layer is working. High retries with rising errors means it just stopped working. The outage feels sudden only if you were not watching the retries.
  • qtime is an average; qcur is now. A brief queue burst can be diluted across the 1024-sample window and never show in qtime. Alert on qcur for real-time visibility, use qtime for latency budgeting.
  • maxconn is a hierarchy, not a number. Limits are enforced independently at global, frontend, backend, and per-server levels. A per-server slim can bottleneck a backend while the global limit has plenty of headroom. Monitor all four.

Level 3: mature

Full coverage of internals and the distinctions that determine where you look first during an incident.

SignalSourceWhy it mattersWarning sign
Frontend minus backend 5xx deltaComputed: frontend hrsp_5xx rate minus backend hrsp_5xx ratesSeparates “the app is broken” from “HAProxy is generating 503/504 itself.” These have completely different response pathsDelta rising: HAProxy-side capacity, connectivity, or config problem
Response errorsCSV eresp on BACKEND and server rowsProtocol-level failure: the server crashed mid-response or sent garbage. Distinct from a clean 5xxAny sustained nonzero rate
Connect timeCSV ctime on BACKEND and server rowsNetwork path and backend accept-queue health. With http-reuse, ctime near zero is normal; a sudden appearance of nonzero ctime means reuse brokectime climbing toward timeout connect, or > 2x baseline
Health check failuresCSV chkfail, chkdown, last_chk, check_durationchkfail rising while status stays UP means a borderline server flapping below the fall thresholdchkfail rate climbing; check_duration near timeout check
Idle_pctshow infoHAProxy’s own CPU saturation signal. Below 20% the event loop is stressed; below 5% latency will spikeSustained < 20%
FD headroom/proc/<pid>/fd count vs Max open files in /proc/<pid>/limits; Maxsock in show infoFD exhaustion is a cliff: accept() fails immediately, no graceful degradationUsage above 80% of limit
SSL frontend key rateSslFrontendKeyRate in show info (not in CSV)Full TLS handshakes per second. The dominant CPU cost in TLS-terminating proxiesRate trending toward tested capacity, with Idle_pct falling
SSL session cache effectivenessSslCacheLookups and SslCacheMisses in show info (hit rate must be computed; there is no SslCacheHits field)Low reuse means every connection pays full handshake costMiss rate > 50% sustained with meaningful lookup volume, excluding post-reload transients
Certificate expiryopenssl x509 -in <cert> -noout -enddate on cert files on disk; show ssl cert lists loaded files but does not expose expiryExpiry is total TLS failure for the affected frontend, and it is self-inflicted every timePage at 24h, ticket at 7 days, plan at 30 days
Soft-stop lingeringStopping in show info, HAProxy PID countDraining processes hold FDs and memory. Reload storms accumulate themA soft-stop process older than ~10 minutes; more than 2-3 HAProxy PIDs

The busy-polling caveat is mandatory reading at this level: when busy-polling is enabled, Idle_pct sits near zero by design because the event loop never sleeps. Alerting on Idle_pct in that mode pages you forever. Use show activity run-queue depth or system CPU instead. Similarly, Idle_pct is an average across threads; one hot thread can hide inside a healthy average, which is what show activity per-thread counters are for.

On certificates: expect a brief SslCacheMisses spike after every reload, because the in-process session cache is invalidated. Do not alert on the transient.

Level 4: expert

The signals teams add after learning the hard way.

  • Kernel accept queue overflow. ListenOverflows and ListenDrops in /proc/net/netstat (or nstat -az). When the kernel queue overflows, SYNs are dropped silently. HAProxy sees nothing because the connections never reached user space. Client timeouts with perfectly healthy HAProxy metrics is the signature. Check the ConnRate vs SessRate gap in show info as a secondary hint.
  • Per-thread skew. show activity exposes per-thread counters (introduced in HAProxy 1.9). With nbthread > 1, one saturated thread caps total throughput while average Idle_pct looks fine.
  • Connection pool utilization. CSV connect vs reuse on BACKEND rows, plus idle_conn_cur. A sudden drop in the reuse ratio means something broke pooling (backend started sending Connection: close, HTTP version change). It shows up as rising ctime and SslBackendKeyRate before anyone notices the cause.
  • Stick-table utilization. show table (not in CSV). Eviction under default settings is silent; with nopurge, new entries are rejected. Either way, rate limiting and persistence quietly stop working for affected clients. Above 80% used deserves a ticket.
  • DNS resolver health. show resolvers for server-template / service-discovery deployments. Resolver failures silently route to stale IPs while server status may still read UP on cached results.
  • Peer replication health. show peers. Broken peer sync means stick tables diverge: rate limiting works on one node and not the other.
  • Allocator pressure. PoolFailed in show info, detail via show pools. Normally zero; any nonzero value means HAProxy failed to allocate internal objects and dropped work.
  • Dropped logs. DroppedLogs in show info. Losing logs during an incident cripples the investigation. UDP syslog drops are silent at the sender; this counter is the only local signal.
  • Tail latency from logs. qtime/ctime/rtime/ttime are averages over 1024 samples. P99 requires parsing per-request timing fields from HAProxy logs. A backend failing fast shows low rtime with rising 5xx, which looks great on a latency dashboard.

Rules that apply at every level

  • Alert on ratios, not absolutes. scur > 1000 is meaningless without knowing slim. scur > 80% of slim works on every instance regardless of size. Same for 5xx: express it as a fraction of req_rate.
  • Handle counter resets. Every cumulative CSV counter (hrsp_5xx, econ, wretr, bin…) resets to zero on reload. Use Uptime_sec drops or the Stopping flag to detect reloads, and make sure your monitoring treats the reset as a reload, not an incident.
  • Health checks are not traffic. A server can pass health checks while returning 500 on every real request (the health endpoint is fine, the app is not). Always pair server status with per-server hrsp_5xx. If frontend 5xx approximately equals backend 5xx while all servers are UP, the application is broken, not HAProxy.
  • Do not confuse session rate with request rate. With keep-alive and HTTP/2 multiplexing, one session carries many requests. req_rate / rate is the reuse ratio; a drop toward 1:1 means every request now pays full connection setup cost.

How Netdata helps

Netdata’s HAProxy collector scrapes the stats socket and CSV stats at per-second granularity, which matters specifically for the signals in this checklist:

  • Per-proxy and per-server charts for scur/slim, qcur, and status, so the maxconn hierarchy and per-server saturation are visible instead of averaged away.
  • Frontend vs backend 5xx side by side, making the HAProxy-generated error delta a visual comparison rather than a manual computation.
  • wretr/wredis and econ/eresp as first-class series, so retry-masked instability and protocol-level failures show up before they become user-visible errors.
  • Idle_pct trending next to SslFrontendKeyRate, which is the correlation that confirms (or rules out) TLS as the CPU consumer.
  • Counter-reset-aware rate computation, so reloads do not produce phantom 5xx drops or spikes.
  • Host-level correlation with FD usage, /proc/net/netstat listen overflows, and process counts, covering the kernel-side and reload-storm signals the CSV stats cannot see.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.