The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-counter-resets-on-reload

Operations Guides

HAProxy counter resets on reload: phantom drops and spikes in your metrics

Your HAProxy 5xx dashboard just showed a cliff-edge drop to zero, then a vertical spike an hour later. No user complaints. No backend incidents. No deploys. If this pattern repeats on a schedule, or right after config changes, you are almost certainly looking at counter resets from HAProxy reloads, not real traffic events.

Every time HAProxy reloads its configuration, a new worker process starts. That new process has its own in-memory statistics, and every cumulative counter starts at zero: hrsp_5xx, econ, eresp, bin, bout, wretr, wredis, chkfail, the SSL counters, all of them. A reload is a new process for stats purposes, not a continuation of the old one.

Any monitoring pipeline that computes rates from these counters without handling resets will produce garbage at reload boundaries: phantom negative rates, phantom error spikes, false alerts, or a real incident hidden inside the noise. This is a monitoring configuration problem, not an HAProxy bug. HAProxy has always worked this way.

This article covers what actually resets, how to confirm a reset caused a given graph artifact, and how to make collection, rate math, and alerting reset-aware.

What this means

HAProxy keeps its statistics in process memory. There is no persistence layer for the classic counters. When you run systemctl reload haproxy (or send the equivalent signal), the running process does not mutate its config in place. A new process starts with the new configuration, the old process enters soft-stop and drains its existing connections, and the new process begins serving with a completely fresh set of counters.

Two consequences:

  1. Counters go to zero on the new process. A monitoring system scraping hrsp_5xx sees something like: 15,432, 15,471, 15,509, 0, 3, 9. Naive delta math reads the 15,509 to 0 transition as either a huge negative rate (clamped or wrapped into a huge positive rate, depending on the tooling) or as a discontinuity that corrupts the rate window.

  2. The old process becomes invisible. After the new process takes over the stats socket, the old draining process’s counters stop updating in most collection setups. Any traffic still flowing through the old worker is unmeasured until it exits.

The artifacts are predictable:

  • Phantom drops. The cumulative error graph falls off a cliff to zero. If your alert logic is “5xx total decreased” or your SLO burn math assumes monotonic counters, it misfires.
  • Phantom spikes. Some rate computations treat a counter decrease as a wrap-around and compute an enormous positive delta. One reload can produce a single data point that dwarfs real traffic.
  • Masked real events. If a genuine error burst happens in the same scrape window as a reload, the reset math can swallow it entirely. The alert that should have fired did not.
  • Broken averages. Rolling-window averages (qtime, ctime, rtime, ttime) restart with no samples on the new process. Instantaneous rate fields (rate, req_rate) also restart. Baselines computed across a reload boundary are unreliable.

In environments with frequent automated reloads (Kubernetes ingress controllers, service discovery re-rendering config on every backend change), this can happen dozens of times a day, and every reload is a chance for a false page or a missed one.

Common causes

CauseWhat it looks likeFirst thing to check
Config reload (manual or systemctl reload)All cumulative counters drop to zero at one timestamp; Uptime_sec restartsshow info field Uptime_sec: is it small right after the artifact?
Automated reloads from service discovery or ingress controllerReset artifacts repeat many times per day, often correlated with backend scale eventsCount reloads: process start times, Uptime_sec history, or config-management logs
Full restart (deploy, crash, OOM kill)Same counter reset plus a traffic gap; old connections droppedSystem journal, dmesg for OOM killer, process start time
Reload storm: old workers lingering in soft-stopMultiple haproxy PIDs, rising FD and memory use, repeated resets close togetherpgrep -c haproxy, Stopping: 1 in show info on the draining process
Manual clear counters via Runtime APICounters zeroed with no reload: Uptime_sec keeps climbingCheck who ran Runtime API commands; distinguish from reloads by uptime continuity

The distinguishing test for most of these is Uptime_sec from show info. A reload or restart resets it. A manual clear counters does not.

Quick checks

All of these are read-only.

# 1. Current uptime of the active worker. A small value right after a
#    graph artifact confirms a reload or restart caused it.
echo "show info" | socat unix-connect:/var/run/haproxy.sock stdio | grep -E "^(Uptime_sec|Uptime):"

# 2. Is the process you are scraping in soft-stop (draining)?
#    Stopping: 0 = normal, 1 = soft-stop in progress.
echo "show info" | socat unix-connect:/var/run/haproxy.sock stdio | grep "^Stopping:"

# 3. How many haproxy processes are running? More than the expected
#    master plus one worker means a reload just happened or old
#    workers are lingering.
pgrep -c haproxy

# 4. Current cumulative counters for a suspect proxy. Note the values
#    now and compare across the next reload. Field positions are the
#    standard show stat CSV order.
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '{print $1"/"$2": 5xx="$44" econ="$14" bin="$9" bout="$10}'

# 5. When did the current process actually start?
ps -o pid,lstart,cmd -C haproxy

Check 1 is the single most useful one. If Uptime_sec is 45 seconds and your dashboard shows a “5xx collapse” 45 seconds ago, you have your answer.

How to diagnose it

When a graph shows an impossible drop or spike in an HAProxy cumulative metric, work through this sequence before treating it as a real event.

  1. Overlay Uptime_sec on the artifact. If the metric discontinuity lines up with an uptime drop to near zero, it was a reload or restart. Stop investigating the traffic; start investigating why the reload happened.

  2. Check whether the artifact shape matches reset math. A real error spike rises over several samples and decays. A reset artifact is a single-sample discontinuity: cliff to zero, or one absurdly large delta. If your graph shows exactly one anomalous point at the reload instant and normal values on both sides, that is reset math, not traffic.

  3. Correlate with other cumulative counters. A real 5xx burst shows up in hrsp_5xx but leaves bin/bout and req_tot growing smoothly. A reload zeroes all of them simultaneously. If every cumulative counter for the proxy broke at the same timestamp, it was a process event.

  4. Rule out a restart vs a reload. A graceful reload keeps existing connections alive on the draining worker. A restart drops them. Check the system journal and pgrep -c haproxy history. If connections were reset and there was a service gap, treat it as a restart incident and review why it restarted.

  5. Check for a reload storm. If resets are frequent, count processes and look for Stopping: 1 lingering on old workers. Frequent reloads from service discovery plus no hard-stop-after means old workers can hold FDs and memory while draining. That is a separate operational problem, but it also multiplies your counter-reset artifacts.

  6. Only then look at the traffic. If uptime is old, no reload occurred, and the counters still moved strangely, check for a manual clear counters (Runtime API audit) or a tooling bug in your collector.

flowchart TD
  A[Graph shows drop or spike in cumulative counter] --> B{Uptime_sec dropped at same time?}
  B -- Yes --> C{Multiple PIDs or Stopping: 1 lingering?}
  C -- Yes --> D[Reload storm: fix reload frequency and hard-stop-after]
  C -- No --> E[Single reload: fix monitoring rate math]
  B -- No --> F{All counters moved together?}
  F -- Yes --> G[Manual clear counters or collector bug]
  F -- No --> H[Real traffic event: investigate service]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Uptime_sec (show info)The reset detector. Every reload zeroes it along with the countersAny drop; small absolute value right after a graph artifact
Stopping (show info)Flags the draining old worker during a reloadStopping: 1 persisting more than a few minutes: stuck drain, missing hard-stop-after
haproxy process countReveals reload frequency and lingering old workersMore than expected master + worker processes
Cumulative error counters (hrsp_5xx, econ, eresp, wretr, wredis)The counters that produce the phantom artifactsAny rate alert that does not survive a counter reset is misconfigured
Byte counters (bin, bout, req_tot)Canary for resets: they move with every reloadSimultaneous discontinuity across all of them means process event, not traffic
SSL counters (SslFrontendKeyRate, SslCacheLookups, SslCacheMisses)Also reset on reload; cache-miss spikes after reload are expected brieflyPost-reload cache miss rate read as a real TLS problem
Instantaneous rates (rate, req_rate)Do not reset, but restart from an empty window on the new processFirst samples after reload are meaningless; do not alert on them

Scope note: everything in the stats CSV and show info that is cumulative is per-process. Instantaneous gauges (scur, qcur) survive conceptually but reflect only the new process after a reload, which is also a discontinuity if you were mid-incident.

Fixes

The fix is on the monitoring side. Grouped by approach.

Use reset-aware rate math

The core fix: compute rates with a function that treats a counter decrease as a reset, not as a negative delta or a wrap.

  • In Prometheus, rate() and increase() already handle counter resets by discarding the decreasing sample. If you hand-rolled delta queries, switch to these.
  • In custom collectors: when current < previous, count only current as the delta for that interval (assume the counter started at zero), or discard the interval entirely and mark it.
  • Never alert on the raw cumulative value of an HAProxy counter. Alert on reset-aware rates expressed as ratios (for example, 5xx rate as a percentage of req_rate), which are naturally bounded.

Tradeoff: reset-aware rate math discards or approximates the reload interval. If reloads are frequent and your scrape interval is long, real errors in the pre-reload window are undercounted. Shorten the scrape interval relative to your reload frequency, or gate alerts as below.

Detect resets explicitly and gate alerts

Even with correct rate math, the window around a reload is untrustworthy: first samples are empty, averages have no history, and the old worker’s traffic is invisible. Use the reset detector the platform gives you:

  • Track Uptime_sec as a metric. A drop between scrapes means a reload or restart happened. Suppress or annotate alerts whose evaluation window crosses that drop.
  • Use Uptime_sec as an alert precondition. The standard pattern is to require Uptime_sec above a threshold (for example, 600 seconds) before rate-based alerts can fire. This eliminates cold-start and post-reload false positives in one move.
  • Track the Stopping flag. It tells you a drain is in progress, which explains both the reset and any temporary measurement gap.

Reduce how often resets happen

Every reload you eliminate is a measurement discontinuity you do not have to handle.

  • Batch configuration changes instead of reloading on every service discovery event.
  • If you run a controller or templating system that reloads on every backend change, add debouncing or change aggregation.
  • Set hard-stop-after so draining workers actually exit; lingering soft-stop processes extend the window where two processes exist and one of them is invisible to your collector.

Tradeoff: less frequent reloads mean slower propagation of legitimate config and backend changes. For dynamic environments, find the batching interval that balances convergence speed against metric continuity.

Persist counters across reloads (HAProxy 3.0 and later)

HAProxy 3.0 added a counter persistence mechanism for the first time:

  • The Runtime API command dump stats-file writes current proxy counters to a file.
  • The global directive stats-file preloads those counters into the next process at startup, so the new worker resumes from the old values instead of zero.
  • Only proxy counters (frontends, backends, servers, listeners) are covered, and objects need a guid assigned for the state to match across processes.

This is best-effort, not exact. Connections still draining on the old worker when you dump will lose their subsequent counter increments, and there are known cases where preloading produces small counter decreases rather than a clean continuation, which poorly-behaved rate math can still misread. Test it against your monitoring pipeline before relying on it, and keep the Uptime_sec gating in place regardless. HAProxy 3.3 also added an experimental shared-memory variant (shm-stats-file <name>, tagged EXPERIMENTAL in the configuration manual) that preserves shareable counters in shared memory across reloads and releases them on a proper stop; treat it as experimental until it stabilizes for your version.

Fix the dashboards, not just the alerts

Even with alerting fixed, dashboards that render raw cumulative counters will keep showing cliffs at every reload, and someone will misread one during an incident. Convert panels to reset-aware rates, and mark reload events (from Uptime_sec drops) as annotations so the discontinuities have a visible explanation.

Prevention

  • Reset-aware rates everywhere. No HAProxy cumulative counter should ever feed a raw-value alert or a naive delta. Use non-negative derivative / reset-aware rate functions.
  • Uptime_sec gating on rate alerts. Require a minimum process age before rate-based alerts evaluate. This one condition removes most false pages from reloads, restarts, and cold starts.
  • Reload annotations. Emit an event or annotation on every Uptime_sec drop so graph discontinuities are explainable at a glance.
  • Reload hygiene. Batch config changes, debounce discovery-driven reloads, set hard-stop-after, and alert on process count and lingering Stopping: 1 to catch reload storms.
  • Ratio-based thresholds. Alert on 5xx as a percentage of request rate, not absolute counts. Ratios are bounded and survive traffic-scale differences, and they behave better across resets.
  • Scrape interval shorter than reload interval. If reloads happen every few minutes and you scrape every 60 seconds, your rate math is structurally unreliable. Fix the interval, the reload frequency, or both.

How Netdata helps

  • Netdata collects HAProxy stats through the socket and CSV interfaces and computes rates with counter-reset-aware math, so reload zeroing does not turn into phantom negative or wrapped spikes on dashboards.
  • Uptime_sec is collected alongside the counters, which makes it possible to see the exact reload moment overlaid with the discontinuity it caused.
  • Cumulative error counters (hrsp_5xx, econ, eresp, retries and redispatches) are charted as rates, so a reset appears as a gap or annotation rather than a fake traffic event.
  • The Stopping flag and process-level visibility surface lingering draining workers, which is the difference between a normal reload and a reload storm.
  • Per-second collection granularity shortens the corrupted window around each reload compared to minute-level scraping, so less real signal is lost at the boundary.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.