The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / haproxy / haproxy-rate-limiting-per-source-ip

Operations Guides

HAProxy per-source-IP rate limiting: stick tables, conn_rate, and http_req_rate

Per-source-IP rate limiting in HAProxy is built on stick tables: in-memory key-value stores that track counters per client key (usually the source IP) and expose those counters to ACLs. When it breaks, it breaks silently. A full table, a mismatched expire, or an IPv6 client population can disable your limiting without a single error in the logs, and most teams find out only after an abuse incident.

This is an operational reference: what to track, how to configure it, how to inspect the table at runtime, how to verify the policy is actually denying traffic, and the failure modes that make limiting stop working.

What a stick table gives you

A stick table is a fixed-size, in-memory store inside each HAProxy process. Each entry is keyed by something you choose (source IP, a header value, a cookie) and stores typed counters that HAProxy updates as traffic passes. For rate limiting, three data types matter:

Data typeWhat it countsWhat it catches
conn_rate(<period>)TCP connection rate per key over the periodConnection floods, SYN-level abuse, rapid reconnect loops
http_req_rate(<period>)HTTP request rate per key over the periodScraping, API abuse, credential stuffing at the request layer
http_err_rate(<period>)HTTP error response rate per key over the periodDirectory scanning (404s), brute force (401/403s), fuzzing (400s)

All three are sliding-window averages, not instantaneous counters. A client that bursts to 200 req/s in one second and then stops will read as over-threshold for roughly the length of the window before the rate decays. Plan thresholds with that lag in mind.

The counters answer different questions. conn_rate sees connection-level abuse that never reaches HTTP. http_req_rate sees application-layer abuse riding on a few keep-alive connections, where conn_rate looks normal. http_err_rate separates chatty-but-legitimate clients from clients whose traffic is mostly failing, the classic scanner and credential-stuffing signature. Tracking only one leaves blind spots.

Reference configuration

An enforcement setup has three parts: the table definition, the tracking directive, and the ACLs that act on the counters.

frontend ft_http
    bind :80
    # Track per-source-IP connection, request, and error rates
    stick-table type ip size 1m expire 60s store conn_rate(10s),http_req_rate(10s),http_err_rate(10s)

    # Track every client from connection accept onward.
    # Must happen here, not in http-request rules: tcp-request connection
    # rules are evaluated before any http-request directive runs.
    tcp-request connection track-sc0 src

    # Enforce at the connection layer for connection floods
    tcp-request connection reject if { sc_conn_rate(0) gt 50 }

    # Enforce at the request layer
    http-request deny deny_status 429 if { sc_http_req_rate(0) gt 100 }

    default_backend bk_app

Points that matter operationally:

  • Evaluation order is the silent trap. If you track with http-request track-sc0 src and then read sc_conn_rate(0) in a tcp-request connection rule, the connection rule runs before tracking has happened, the fetch returns 0, and the reject never fires. No error, no log line. Track at connection level when connection rules depend on the counter.
  • expire must outlive the rate window. If expire is shorter than the tracking window, entries vanish mid-window and counters reset, weakening enforcement and churning memory.
  • The sticky counter slot must match. If you track on sc0 but read with sc_http_req_rate(1), the fetch returns 0 and the rule never fires. One of the most common silent misconfigurations.
  • Connection-layer and request-layer rules are complementary. tcp-request connection reject acts before HTTP processing and increments dcon/dses; http-request deny acts per request and increments dreq. You will use these counters later to verify enforcement.

The enforcement flow, and where the observability hooks sit:

flowchart LR
    C[Client source IP] --> F[Frontend]
    F --> T["track-sc0 src
stick table lookup"] T --> A{ACL rate check} A -->|under threshold| B[Backend] A -->|conn_rate over| R1["reject
dcon / dses"] A -->|req_rate over| R2["deny 429
dreq"] T -.->|table full| X["entry rejected or evicted
limiting silently off"]

Inspecting the table at runtime

Stick tables are not in the CSV stats. You inspect them through the runtime API. This is the only way to see who is tracked and at what rate.

# List all tables with capacity and current entry count
echo "show table" | socat unix-connect:/var/run/haproxy.sock stdio

# Dump entries for a specific table (shows key, use, exp, and stored counters)
echo "show table ft_http" | socat unix-connect:/var/run/haproxy.sock stdio

# Top talkers: request rate above a threshold
echo "show table ft_http data.http_req_rate gt 50" | socat unix-connect:/var/run/haproxy.sock stdio

# Connection floods
echo "show table ft_http data.conn_rate gt 100" | socat unix-connect:/var/run/haproxy.sock stdio

# Likely scanners / brute force
echo "show table ft_http data.http_err_rate gt 10" | socat unix-connect:/var/run/haproxy.sock stdio

Entry lines show the key plus tracked values (for example http_req_rate(10000)=37), along with exp (time until expiry). Two things to check first during an incident:

  • Table header: size vs used. If used is at or near size, you are in overflow territory (see below) and limiting may already be degraded for new sources.
  • The shape of the top entries. A handful of IPs far above everyone else suggests targeted abuse. Thousands of IPs each slightly over threshold suggests either a distributed attack or a threshold set below your legitimate per-user rate. NAT and corporate proxies aggregate many users behind one IP, so their legitimate rate can be very high.

Verifying the policy is actually blocking

A rate limit that never fires is indistinguishable from a rate limit that is silently broken, unless you watch the denial counters. Pair table inspection with the frontend denial metrics:

# TCP-layer denials (connection rejects)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "FRONTEND" {print $1": dcon="$82" dses="$83}'

# Request-layer denials (ACL denies)
echo "show stat" | socat unix-connect:/var/run/haproxy.sock stdio | \
  awk -F, '$2 == "FRONTEND" {print $1": dreq="$11}'

How to read these:

  • dreq rising while top table entries sit above your threshold: enforcement is working.
  • Table shows abusive entries but dreq/dcon are flat: the ACL is not matching. Check the sticky counter slot, the fetch name, the tracking directive’s evaluation point, and the threshold direction.
  • dreq spikes after a config change: your ACLs may now be denying legitimate traffic. Denied requests surface to clients as 403/429, so correlate with hrsp_4xx on the frontend.
  • First enabling enforcement: expect a brief 4xx bump as previously-flowing abusive traffic starts getting denied. That is the policy working, not a regression.

These are cumulative counters that reset on every reload, so alert on delta rates, not absolute values.

Gotchas that silently disable limiting

None of these produce an error message.

Table overflow disables limiting for new IPs. The table is fixed size. When it fills, the default behavior is to evict expired (and then oldest) entries, silently, with no counter or log line. Under a flood of unique source IPs, entries are evicted before their natural expiry, their counters reset to zero, and new abusers cycle through the table faster than the window can convict them. With nopurge configured, eviction is disabled and new entries are rejected outright, so brand-new clients are never tracked or limited at all. Either way, limiting degrades exactly when you need it most. Size the table for peak concurrent unique keys plus headroom, and monitor utilization.

Expire/window mismatch resets counters. If expire is too short relative to the rate window, entries expire between requests, the counter restarts, and a steady abuser never accumulates a convicting rate. Keep expire comfortably longer than the tracking window.

Reloads wipe the table. Stick tables live in process memory. Every reload starts a new process with an empty table unless peers replicate state, and peer sync itself takes time. A reload during an active attack reopens the window. Frequent automated reloads in dynamic environments make this worse; it is the same counter-reset problem covered in the reload guide, with a security consequence.

IPv6 makes per-IP limiting much less effective. An attacker with a single /64 has effectively unlimited source addresses and can rotate through them, defeating per-IP tracking while flooding your table toward overflow. Two mitigations, used together: key on a prefix rather than a full address using the native ipmask converter (e.g. src,ipmask(24,64) to mask IPv4 to /24 and IPv6 to /64), and combine IP tracking with a second key such as an API token or cookie where your application has one. Note that type ip tables store only IPv4 addresses and type ipv6 tables store only IPv6; a type ip table never tracks IPv6 clients, so on dual-stack services you need type ipv6 with ipmask masking or a second table.

Aggregated clients share one key. CDNs, NAT gateways, and corporate egress proxies concentrate many users behind one source IP. A per-IP threshold reasonable for direct clients will false-fire on these. Whitelist known aggregators, set a much higher threshold for them, or track a better key (verified X-Forwarded-For, an auth token) for that traffic class.

Sliding-window lag cuts both ways. After a burst ends, the reported rate takes roughly one window to decay. If automation bans clients on the current rate, a client that already stopped can still get banned seconds later; if you unban manually, the client stays blocked until the window clears.

Per-process tables diverge in multi-instance deployments. Each HAProxy instance has its own table. Without peers replication, an abuser is limited per instance, not globally, so the effective per-client budget is your threshold multiplied by instance count. With peers, broken replication (check show peers) silently produces the same divergence.

Sizing and tuning decisions

  1. Estimate peak unique keys. Take unique source IPs per minute at peak, multiply by (expire in minutes), add headroom for bursts. A DDoS with rotating sources will try to exhaust whatever number you pick, which is why utilization monitoring matters more than the initial guess.
  2. Pick the window to match the abuse pattern. Short windows (10s) catch sharp bursts but let slow-and-low scrapers through. Longer windows (60s+) catch slow abuse but react slowly. Many deployments run two tables or track two windows on one.
  3. Choose thresholds as ratios of legitimate behavior. Look at your real top talkers in show table before setting limits. A threshold below the 99th percentile of legitimate per-source rate will deny real users; one set at 10x the legitimate p99 catches only the egregious.
  4. Decide overflow policy explicitly. Default eviction keeps the table accepting new entries at the cost of tracking continuity. nopurge protects already-tracked clients at the cost of ignoring new ones. Neither is wrong; picking one deliberately beats discovering the default during an attack.
  5. Memory is proportional to size times stored data types. Three rate counters per entry cost more than one. Stick-table memory is part of HAProxy’s pool usage, so account for it alongside connection buffers when capacity planning.

Signals to monitor

SignalWhy it mattersWarning sign
Stick table used vs size (show table)Overflow silently disables limiting for new sourcesused > 80% of size sustained
dreq / dcon / dses (frontend CSV)Proves the policy is denying; absence of denials during a known attack means broken enforcementFlat during confirmed abuse, or sudden spike after config change
Top-entry rates (show table <t> data.X gt N)Identifies scrapers, bots, DDoS sources while the attack is underwayEntries far above baseline, or thousands of marginal entries
hrsp_4xx on the frontendDenied requests surface as 403/429; also catches scanner noise (400/404) and Slowloris (408)Spike correlated with enforcement change, or uncorrelated spike meaning attack
Peer sync state (show peers)Diverged tables mean inconsistent limiting across instancesPeer disconnected or update counts frozen
New entry churn vs expiryEntry creation rate persistently exceeding expiry rate predicts overflow before it happensused climbing toward size over hours

Severity guidance: per-source abuse signals are TICKET-level unless automated mitigation is wired in, at which point enforcement is self-acting and the alert becomes informational. The exception is table overflow on a security-critical table, which deserves urgent attention because your protection is down and you have no other way to know.

How Netdata helps

  • Frontend denial counters over time. Netdata collects the HAProxy CSV stats, including dreq, dcon, and dses, so you see enforcement activity as a time series instead of polling the socket by hand during an incident.
  • Denials correlated with 4xx responses. Overlaying dreq against hrsp_4xx separates policy denies (your rate limiter working) from client errors and scanner noise.
  • Saturation context during attacks. Session rate, scur/slim, and request rate alongside denial counters tell you whether limiting is holding the line or connections are still piling up.
  • Reload detection. Counters reset on reload and stick tables empty with them; Netdata’s counter-reset handling and uptime tracking help you spot when a reload reopened the rate-limit window.
  • Anomaly flagging on request and error rates. ML-based anomaly detection on frontend rates surfaces the early shape of a scrape or credential-stuffing run before it crosses your static threshold.
The Netdata solution

HAProxy load balancer monitoring with Netdata

Netdata monitors HAProxy with per-second frontend, backend, and queue metrics plus ML-powered anomaly detection. Correlate maxconn saturation, queue buildup, health-check cascades, 5xx attribution, and file-descriptor exhaustion against the backend and host signals behind them, so you catch the incidents in these runbooks before they page anyone.