The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-server-template-dynamic-backends

Operations Guides

HAProxy server-template and dynamic backends: monitoring service-discovery routing

With a static HAProxy config, the server list lives in the config file. Capacity review is config review: count the server lines and you know what the backend can absorb. With server-template, the server list lives in DNS. HAProxy pre-provisions a pool of empty server slots, and its internal resolver fills and empties them as A or SRV records change. This is the mechanism behind HAProxy fronting Kubernetes Ingress, Consul, and most autoscaled service-discovery setups.

That shift moves the failure domain. The resolver becomes a runtime dependency whose failure does not look like an outage: servers stay UP on cached IPs until those IPs stop working. The act count becomes a discovery signal, not just a health signal. And depending on how your control plane pushes changes, reloads can go from rare to constant, with real resource costs.

This article covers how the mechanism works, what changes operationally versus static backends, and which signals tell you the discovery pipeline is healthy.

What server-template is and why it matters

server-template declares a named pool of server slots in one line:

backend web
  balance roundrobin
  server-template web 10 _web._tcp.service.consul resolvers mydns check init-addr none resolve-opts allow-dup-ip resolve-prefer ipv4

The arguments are: a name prefix (web), a slot count (10), and a DNS name to resolve. For SRV records, the port comes from the record itself, so it is omitted from the template line. For plain A records, the port is given in the template and the record supplies only IPs. At runtime, HAProxy materializes slots as web1, web2, and so on, and the resolver section (resolvers mydns) keeps them populated.

Every property you took for granted from static config is now dynamic. Server identity, count, and addresses can all change without a config edit or a reload. Monitoring that assumes a fixed inventory will misread this environment.

How it works

flowchart LR
  DNS["DNS A or SRV records"] -->|queried on TTL cadence| RES["resolvers section"]
  RES -->|IP and port answers| POOL["server-template slot pool"]
  POOL -->|per-slot probing| HC["health check engine"]
  HC -->|UP / DOWN| ACT["backend act count"]
  ACT -->|serves traffic| APP["live endpoints"]
  RES -.->|failure| STALE["stale cached IPs"]
  STALE -.->|connect failures| ECON["econ rises"]

The lifecycle of a slot:

  1. Template expansion. At startup, HAProxy creates the full slot pool as named servers (web1 through web10). With init-addr none, HAProxy boots even if DNS is unreachable; slots start unresolved and are resolved at runtime. This is the standard pattern for container environments.
  2. Resolution. The resolver queries the configured nameservers on a TTL-driven cadence. SRV answers carry both IP and port; A answers carry IPs only.
  3. Slot assignment. Answers map into slots in order. As records appear and disappear, slots are reassigned. A slot’s IP can change over its lifetime, which has a monitoring consequence covered below.
  4. Health checking. The health check engine probes each populated slot independently of resolution. A resolved slot that fails checks leaves the active set until it recovers.
  5. Routing. The backend’s act count is the number of active-role slots currently UP. HAProxy load-balances only across those.

Two details drive most of the operational surprises. First, when DNS records are removed, some HAProxy versions do not immediately clear the stale address from the slot; the health check is the backstop that takes the dead endpoint out of rotation. Second, when the resolver itself fails, HAProxy keeps using its cached results until they expire, so a resolver outage surfaces later as connection errors to dead IPs, not as an immediate error anywhere.

Where dynamic backends show up

  • Kubernetes Ingress (HAProxy Ingress Controller). Endpoints churn with pod scheduling, rollouts, and autoscaling. Slot counts must cover peak replica counts, and endpoint churn during rollouts is constant background noise.
  • Consul service discovery. Services registered in Consul are consumed through Consul’s DNS interface, typically via SRV records.
  • Any autoscaled group behind DNS. Anything where the backend inventory changes faster than you want to edit config files.

In all of these, DNS TTL controls how quickly topology changes propagate into HAProxy. A long TTL means HAProxy reacts slowly to scale-down events; a short TTL increases query load on the resolver and the DNS server.

Operational differences from static config

The resolver is a single point of silent failure. In a static config, DNS is a startup-time concern. Here it is a runtime dependency. A resolver that stops answering does not down any server; HAProxy serves from cache until entries expire, then keeps routing to whatever it last knew. show resolvers (sent queries, received answers, timeouts, errors) is the only place this degradation surfaces before traffic breaks.

act is now a discovery metric. With static servers, a dropping act count means health failures. Here it can also mean DNS returned fewer endpoints, records were deleted, or slots are exhausted. Alert on act relative to the expected endpoint count for the service, not just on act = 0. Note that act and bck only count servers currently UP; for the full picture (DOWN, MAINT, DRAIN), count individual server rows by their status field.

Slots cap how many endpoints receive traffic. If DNS returns more endpoints than the template has slots, the extras get no traffic and nothing errors. Size the pool for peak replica count, and watch act approaching the slot count as a capacity signal on the discovery side.

A slot name is not a machine. web3 is a slot, not an identity. Its IP can change when records churn. Per-server time series will show a slot’s behavior jump when it remaps, and per-server econ can spike on a slot that just received a stale assignment. Interpret per-server metrics accordingly; the backend aggregate is usually the more honest signal.

Endpoint churn shows up as econ and check failures. When a pod is rescheduled, its old IP goes stale. Connections to the slot fail (refused or timed out), econ increments, and health checks fail with last_chk showing L4CON or L4TOUT until the slot re-resolves. Some churn-driven econ during rollouts is expected; baseline it so you can tell rollout noise from a resolver outage.

Reloads may be driven by discovery events. Some control planes react to endpoint changes by regenerating config and reloading HAProxy. Frequent reloads produce the classic reload storm: old processes lingering in soft-stop (Stopping: 1), file descriptor and memory pressure, and counters resetting on every cycle. Where possible, prefer resolver-driven updates or Runtime API server add/delete operations, which change the live server list without a reload: add server and del server are documented in the management manual since HAProxy 2.4, set server ... state ready|drain|maint covers live state transitions, and wait srv-removable (drain-before-delete) is documented in current manuals alongside a stricter del server that refuses to delete a server with active or idle connections. If reloads are unavoidable, set hard-stop-after so draining processes cannot linger indefinitely on long-lived connections.

Counter resets corrupt naive alerting. Every reload zeroes cumulative counters (econ, hrsp_5xx, wretr). Rate-based alerts that do not handle resets will see phantom drops and spikes. Use Uptime_sec to tag resets, as covered in the counter-reset guide linked below.

Per-slot health checks have a cost. Every populated (and provisioned) slot carries health check scheduling overhead. Community reports describe high CPU usage with very large slot counts (on the order of a thousand or more). Keep slot counts close to what the service actually needs.

A few config-level traps are specific to this pattern. On Kubernetes, multiple records may resolve to the same IP, which HAProxy rejects by default; resolve-opts allow-dup-ip fixes this, but there is a documented case where the option is silently reset when placed on a default-server line while the template uses resolvers, so put it on the server-template line itself (upstream issue 2519). HAProxy prefers IPv6 answers when the host has an IPv6 stack; set resolve-prefer ipv4 explicitly if IPv6 is not actually routable in your environment. Finally, after a reload that restores server state from a state file, some versions do not re-resolve SRV names promptly, leaving slots with stale addresses; if you use load-server-state-from-file, test the reload path and verify act and slot addresses after every reload.

Reading the live state

These checks are safe and read-only:

# Live server list, including auto-generated slot names, addresses, and state
echo "show servers state" | socat unix-connect:/var/run/haproxy.sock stdio

# Active/backup counts per backend
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "BACKEND" {print $1": active="$20" backup="$21}'

# Resolver health: queries sent, answers received, timeouts, errors
echo "show resolvers" | socat unix-connect:/var/run/haproxy.sock stdio

# Reload and resolution-failure indicators
echo "show info" | socat unix-connect:/var/run/haproxy.sock stdio | \
  grep -E "^(Uptime_sec|Stopping|FailedResolutions):"
# Connection errors per backend and server (stale-IP detection)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 != "FRONTEND" {print $1"/"$2": econ="$14}'

# How many HAProxy processes are running (reload storm check)
pgrep -c haproxy

show servers state is the ground truth for “what does HAProxy think the service looks like right now.” Compare its addresses against a direct DNS lookup of the same name when you suspect staleness.

Signals to watch in production

SignalWhy it mattersWarning sign
show resolvers timeouts and errorsOnly direct view of resolver health before routing breaksAnswers lagging queries; timeout or error counts growing
FailedResolutions (show info)Cumulative resolution failuresAny sustained increments; should sit at zero in steady state
Backend act countMirrors discovered, healthy endpointsact below expected endpoint count, or dropping without a deploy
Per-server status and last_chkWhich slots fail and why (L4TOUT/L4CON point at dead IPs)Multiple slots failing with L4TOUT/L4CON after endpoint churn
econ rate per serverConnections attempted to stale or dead IPsecon climbing on slots whose addresses DNS no longer returns
wretr / wredisChurn being masked by retries and redispatchesRetry rate rising during rollouts and staying elevated afterward
Uptime_sec and HAProxy PID countReload frequency and reload-storm detectionUptime repeatedly resetting; more than 2-3 concurrent HAProxy processes
Stopping durationDrain behavior after reloadsSoft-stop lingering beyond roughly 10 minutes
act vs slot countDiscovery-side capacity ceilingact pinned at the template slot count while the service scales past it

The correlation that shortens most incidents in this architecture: resolver errors rising first, then econ, then servers flapping DOWN with L4TOUT. That ordering says “DNS broke, and now the cached topology is decaying,” which sends you to the resolver and the DNS server rather than to the backends.

How Netdata helps

  • Netdata’s HAProxy collector reads the stats CSV over the runtime socket or HTTP stats endpoint, so per-backend act counts, per-server status, econ/eresp, and retry counters are charted continuously rather than sampled by hand during incidents.
  • Historical act per backend makes “DNS started returning fewer endpoints” visible as a step change, distinguishable from gradual health-check-driven loss.
  • Per-server econ correlated with server status transitions pinpoints stale-IP churn during rollouts and separates it from application-level failures.
  • Reset-aware handling of cumulative counters (with Uptime_sec as the reset marker) keeps frequent reloads from rendering as phantom drops in error rates.
  • Comparing frontend versus backend 5xx rates alongside econ distinguishes HAProxy-generated 503s (no slots left UP) from errors the discovered endpoints returned themselves.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.