The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-show-pools-memory-breakdown

Operations Guides

HAProxy show pools: finding which internal pool is consuming memory

HAProxy’s memory footprint is not one number. The process holds dozens of internal object pools: buffers, connections, sessions, tasks, stick-table entries, SSL state. When RSS climbs, when PoolFailed goes nonzero, or when you are sizing a memory budget for a new instance, the question is never “how much memory is HAProxy using” but “which pool is using it”. show pools is the runtime API command that answers that.

PoolFailed in show info tells you an allocation failed. It does not tell you what failed to allocate. show pools is the diagnostic companion: it lists every internal pool with used versus allocated counts, per-pool failure counters, and average demand, so you can distinguish a connection-buffer blowup from stick-table growth from SSL-context memory.

The command is available since HAProxy 1.9 (confirmed present in the v1.9.0 tag’s management documentation). and is safe to run on a production process: it is read-only and does not flush or reclaim anything.

What show pools gives you

Run it through the stats socket like any other runtime command:

# Dump all internal memory pools
echo "show pools" | socat unix-connect:/var/run/haproxy.sock stdio

Each line is one pool. The output includes the pool name, the item size in bytes, the allocated count and total allocated bytes, the used count, a needed_avg field, a failures counter, and a [SHARED] flag on pools shared across threads. Item sizes are rounded up to the nearest multiple of 16 bytes, so small objects look larger than their struct definitions suggest.

Two things in the output trip people up:

  • Used includes thread caches. Since 2.9 the output annotates (~N by thread caches). Objects recently freed by a thread sit in a per-thread cache and still count as used. A pool can show a high used count with zero active connections. That is normal, not a leak.
  • Allocated is high-water behavior, not current demand. Pools grow to satisfy peak demand and do not aggressively return memory to the OS. Allocated greater than used is expected.

Pair show pools with the summary fields from show info:

# Pool summary and memory limit
echo "show info" | socat unix-connect:/var/run/haproxy.sock stdio | \
  grep -E "^(Memmax_MB|PoolAlloc_MB|PoolUsed_MB|PoolFailed|CurrConns):"
FieldMeaning
PoolAlloc_MBTotal memory allocated from the OS for pools (the caching high-water mark)
PoolUsed_MBMemory actually holding live objects right now
PoolFailedCumulative count of failed pool allocations since the worker started. Should always be zero
Memmax_MBPer-process memory limit, 0 means unlimited
CurrConnsCurrent connections, the main driver of buffer and connection pools

A diagnostic flow

flowchart TD
  A[RSS growth or PoolFailed nonzero] --> B[show info: PoolUsed_MB vs PoolAlloc_MB vs Memmax_MB]
  B --> C[show pools byusage]
  C --> D{Largest growing pool?}
  D --> E[Buffers: CurrConns x 2 x tune.bufsize]
  D --> F[Stick tables: show table used vs size]
  D --> G[SSL pools: handshake rate, Lua, cert storage]
  D --> H[Monotonic growth across reloads: suspect leak, check version]

Sorting and filtering the output

On a busy instance the unfiltered dump is long. Since HAProxy 2.7 the command accepts sorting and filtering options:

# Sort by total usage, largest consumers first (2.7+)
echo "show pools byusage" | socat unix-connect:/var/run/haproxy.sock stdio

# Sort by item size
echo "show pools bysize" | socat unix-connect:/var/run/haproxy.sock stdio

# Only pools whose name starts with a prefix, limit to 10 entries
echo "show pools match ssl 10" | socat unix-connect:/var/run/haproxy.sock stdio

byusage is the one you want during an incident: it sorts by total usage in reverse order and only shows pools with used entries, so the top of the list is your answer. byname sorts alphabetically, useful for diffing two captures.

Since HAProxy 3.2, show pools detailed shows extra per-pool information, including which pools have been merged. Pool merging combines pools whose item sizes are close together to reduce pool count and memory overhead, so a single reported pool may back more than one object type (confirmed: the detailed option is present in v3.2.0 management docs but absent in v3.1.0; the 3.2 changelog documents pool merging). If you are comparing output across versions, expect the pool list on 3.2 to be shorter than on older releases.

Reading the fields that matter

FieldWhat it tells youWhat to look for
Item sizeBytes per object, rounded to 16Confirms what a full pool costs; buffers track tune.bufsize
AllocatedObjects currently allocated from the OSHigh-water behavior; compare to used and to needed_avg
UsedObjects live or held in thread cachesSubtract the (~N by thread caches) annotation mentally before panicking
needed_avgAverage demand over timeAn internal heuristic, not a threshold. Useful only relative to allocated, tracked over time
FailuresAllocation failures from this poolNonzero means this pool hit a limit. This is the per-pool counterpart of PoolFailed

The exact relationship between needed_avg and when a pool grows is not formally documented. Treat it as a trend signal. If needed_avg sits far below allocated for weeks, the pool was sized for a peak that is not coming back. If it tracks close to allocated, the pool is under real pressure.

Mapping pools to subsystems

The pool names map onto HAProxy’s internal subsystems. Knowing which subsystem owns the growing pool is what turns show pools output into an action.

  • Buffer pools. Two buffers per connection (request and response), each tune.bufsize bytes, 16384 by default. Add roughly 20 KB of connection metadata and the rule of thumb from the HAProxy documentation is about 33 kB per connection. If the buffer pool dominates, the driver is CurrConns. Reduce concurrency or reduce tune.bufsize; both trade against real capacity.
  • Connection and session pools. Scale with concurrent sessions and with idle backend connections kept alive for reuse (http-reuse). Growth here tracks CurrConns and connection pool sizing.
  • Task pools. Proportional to active connections plus health checks. Rarely the problem on their own; if tasks grow, connections grew first.
  • Stick-table entries. Proportional to configured table sizes and entry data types. Cross-check with show table, which reports used versus size per table. A table near capacity is both a memory consumer and a functional risk: eviction is silent, and with nopurge new entries are rejected, which breaks rate limiting and persistence for new clients.
  • SSL pools. Session cache, OCSP stapling state, certificates in memory, and per-connection SSL context. Large certificate bundles cost megabytes before a single connection arrives. High new-handshake rates (SslFrontendKeyRate) churn the context pools.

These have completely different fixes. Buffer pressure means tuning tune.bufsize or maxconn. Stick-table growth means resizing tables or shortening expire. SSL pool growth means reviewing session cache sizing, certificate count, or what is loaded into the process.

Distinguishing pressure from a leak

Most show pools investigations end in “this is demand, not a leak”. Before concluding otherwise, eliminate the normal explanations:

  1. Thread caches. On multi-threaded instances (nbthread > 1), the used count includes per-thread caches. Quiet instance, high used count, large (~N by thread caches) annotation: normal.
  2. High-water allocation. PoolAlloc_MB above PoolUsed_MB after a traffic peak is the allocator caching, not leaking.
  3. Reload residue. After a reload, the old worker retains its pools until it finishes draining. Total HAProxy memory is the sum of all running processes. Check pgrep -c haproxy and Stopping in show info on the old process before blaming the new one. Frequent reloads with long-lived connections (WebSocket, streaming) turn this into real growth; hard-stop-after bounds it.
  4. Connection growth. Recompute the expected buffer footprint: CurrConns x (2 x tune.bufsize + ~20 KB metadata). If the pool numbers match the formula, the “leak” is traffic.

A genuine leak has a specific signature: one or two pools growing monotonically across hours and across reloads of the same binary, uncorrelated with CurrConns, with needed_avg climbing alongside used. There is precedent. In HAProxy 2.4.x, a non-POSIX realloc behavior in the Lua allocator caused the h2s and ssl_sock_ct pools to grow unboundedly when Lua scripts were loaded (fixed in later 2.4.x patches). If you see that shape, check your version’s changelog before building a mitigation, and capture show pools twice with a known interval so the growth rate is in evidence.

One operational note: show pools does not flush pools. Per the management guide, sending SIGQUIT to a foreground process dumps and flushes pool state, but sending signals to a production worker is disruptive and can terminate it. For diagnosis, stick to the read-only socket command.

Sizing a memory budget

show pools is also the sizing tool. When you are setting a budget for a new instance or a new maxconn, work from the drivers:

  • Connections. maxconn x ~33 kB at default tune.bufsize is the floor for buffer and connection memory at full saturation. Raising tune.bufsize to 32768 doubles the buffer term; lower maxconn proportionally or accept the larger footprint.
  • Stick tables. Sum configured table sizes. Large tables with many stored data types can dominate pool memory on low-traffic instances.
  • TLS. Session cache size plus in-memory certificates. Multiple large SAN bundles are not free.
  • Headroom. Peaks, reload overlap (two processes briefly), and thread caches all sit on top.

Then enforce it. Memmax_MB in show info reports the per-process memory limit (0 means unlimited); the limit itself is set with the -m command-line option. Without a limit, HAProxy allocates until the OOM killer decides for you. That is a cliff edge, not a degradation curve. On a shared host, set the limit and alert on PoolUsed_MB / Memmax_MB crossing 80%.

Any nonzero PoolFailed means allocations have already failed and work has been dropped. Use show pools to find which pool’s failures counter is incrementing, then decide whether the fix is more budget, less demand, or a smaller tune.bufsize.

Signals to watch in production

SignalWhy it mattersWarning sign
PoolFailedAllocations failing; connections or requests being droppedAny nonzero value
PoolUsed_MB / Memmax_MBPressure against the configured limitSustained above 80%
PoolUsed_MB trend vs CurrConns trendSeparates demand-driven growth from leaksMemory climbing while connections are flat
Per-pool used/allocated from periodic show pools byusageIdentifies the specific resource under pressureOne pool growing monotonically across days
Stick table used vs sizeTable memory and functional capacityUtilization above 80%
HAProxy process countReload residue holds old poolsMore than 2-3 PIDs persisting

Collect show pools byusage on a slow schedule (minutes, not seconds) so that when PoolFailed fires or RSS alarms, you already have the per-pool history to compare against.

How Netdata helps

  • Netdata collects the show info memory fields (PoolAlloc_MB, PoolUsed_MB, PoolFailed) as time series, so the high-water-vs-used relationship and any failure events are visible without socket scripts.
  • Plotting pool memory against CurrConns and session rate makes the demand-vs-leak distinction a visual one: correlated curves are traffic, diverging curves are a leak.
  • Alerting on PoolFailed transitioning from zero catches allocation failures at the first occurrence rather than after the OOM killer acts.
  • Per-backend and per-frontend session saturation metrics (scur/slim) explain sudden buffer-pool growth when a slow backend starts holding connections longer.
  • Process-count and RSS tracking per HAProxy PID exposes reload residue, the most common false positive in pool-growth investigations.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.