The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nginx / nginx-upstream-keepalive

Operations Guides

NGINX upstream keepalive: eliminating per-request TCP and TLS overhead

Every proxied request that opens a fresh TCP connection burns latency on the handshake and, if the upstream uses HTTPS, on TLS negotiation. At low volume this cost is invisible. At production throughput it becomes a measurable tax on every request, and at extreme scale it can exhaust the kernel’s ephemeral port range and bury the host in TIME_WAIT sockets. For HTTPS upstreams, the CPU cost of TLS handshakes across thousands of requests per second consumes worker cycles that could be spent proxying traffic.

The keepalive directive inside an upstream block caches idle backend connections per worker process. When it works, the next request reuses an existing socket and $upstream_connect_time logs as 0.000. When it is missing or misconfigured, you pay the full connection tax every time, and the symptoms look like upstream latency or connect() failed (99: Cannot assign requested address) errors.

What it is and why it matters

The keepalive directive in an upstream block configures a per-worker cache of idle connections to backend servers. Without it, nginx closes the upstream connection after each response. The next request must open a new TCP connection, complete the three-way handshake, and, for HTTPS upstreams, perform a full TLS negotiation. On a local network the TCP handshake adds roughly one millisecond; a TLS handshake can add tens of milliseconds and significant CPU load. Under high request rates the kernel holds each closed socket in TIME_WAIT state for twice the maximum segment lifetime, typically 60 seconds. This is the primary cause of ephemeral port exhaustion on busy reverse proxies.

Nginx in reverse proxy mode uses two connections per proxied request: one from the client and one to the upstream. Reusing the upstream socket removes per-request overhead and prevents the kernel from churning through the ephemeral port range.

How it works

The keepalive mechanism is straightforward in concept but has specific prerequisites and version-dependent behavior.

Inside an upstream block, the keepalive directive sets the maximum number of idle connections that each worker process will cache for that upstream group. When a worker finishes proxying a request, it keeps the upstream connection open and places it in a pool keyed by the upstream server address. The next time that worker needs to reach the same backend, it pulls the idle connection instead of opening a new one.

The pool is per-worker, not global. If you configure worker_processes 8 and keepalive 32, nginx may hold up to 256 idle connections against the upstream in total. This is important for capacity planning on the backend: the upstream server must tolerate the sum of idle pools from all workers plus any active connections currently in flight.

Since nginx 1.29.7, keepalive to HTTP upstreams is enabled by default with a value of keepalive 32 local. The local keyword means cached connections are isolated to the location block that acquired them; they are not shared across different location blocks that reference the same upstream. Without local, any matching idle connection to the same upstream server is reused regardless of which location triggered it.

Versions prior to 1.29.7 required explicit configuration to enable upstream keepalive. You needed to set proxy_http_version 1.1 and proxy_set_header Connection "" in the location block to prevent nginx from sending Connection: close to the backend. Since 1.29.7, the default proxy_http_version is 1.1 and the Connection header is no longer sent by default, so those two lines are unnecessary for basic keepalive. If you need to force HTTP/1.0 to a specific backend, use proxy_http_version 1.0 and proxy_set_header Connection "close".

Several directives control connection lifecycle:

  • keepalive_requests sets the maximum number of requests that nginx will send over a single cached connection before closing it. The default is 1000 since nginx 1.19.10; before that it was 100.
  • keepalive_timeout sets how long an idle connection stays in the pool. The default is 60 seconds since nginx 1.15.3.
  • keepalive_time sets the maximum time a connection may be kept alive regardless of activity. The default is 1 hour since nginx 1.19.10.

When a request arrives, nginx checks the worker’s pool for an idle connection to the target upstream server. If one exists and has not exceeded these limits, nginx reuses it. The access log variable $upstream_connect_time will show 0.000 because no new TCP or TLS handshake occurred. If no idle connection is available, nginx opens a new one and $upstream_connect_time will be nonzero. Note that DNS resolution time is not included in this metric.

flowchart TD
    Request[New upstream request] --> Check{Idle connection available?}
    Check -->|Yes| Reuse[Reuse connection]
    Check -->|No| Handshake[New TCP plus TLS handshake]
    Reuse --> Proxy[Proxy request]
    Handshake --> Proxy
    Proxy --> Return[Return response]
    Return --> Keepalive{Below keepalive_requests and keepalive_time?}
    Keepalive -->|Yes| Pool[Add to idle pool]
    Keepalive -->|No| Close[Close connection]
    Pool --> Timeout{Idle exceeds keepalive_timeout?}
    Timeout -->|Yes| Close
    Timeout -->|No| Check

FastCGI upstreams require fastcgi_keep_conn on in the location block to use keepalive. SCGI and uwsgi do not support keepalive connections. Memcached upstreams use standard keepalive without extra header manipulation.

Where it shows up in production

The most obvious symptom of missing or broken upstream keepalive is elevated $upstream_connect_time. If your backends are on the same datacenter network, new TCP connections should complete in under a millisecond. Seeing a consistent P95 of 2-10 ms on every request often means you are paying for TLS handshakes that could have been avoided.

At higher scale, the kernel symptom is more severe. Without connection reuse, every proxied request opens a new outbound socket. When the response finishes, that socket enters TIME_WAIT for roughly 60 seconds. If your request rate exceeds the drain rate, TIME_WAIT sockets accumulate until they exhaust the ephemeral port range. The error log fills with connect() failed (99: Cannot assign requested address) and clients receive 500-series errors. Confirm this by checking ss -tan state time-wait | wc -l. The fix is enabling upstream keepalive, not merely widening net.ipv4.ip_local_port_range.

A subtler symptom appears during nginx restarts. After a restart or reload, each worker’s keepalive pool is empty. The first requests to each upstream incur the full handshake penalty, causing a brief latency spike until the pools warm. This is normal, but if you never see $upstream_connect_time drop to 0.000 after the warm-up period, your pool is either too small or the backend is closing connections prematurely.

Keepalive also affects cascade failure patterns. When backends slow down, active upstream connections stay busy longer. If the pool is small and traffic is high, new requests cannot find idle connections and must open new ones. This increases load on the already struggling backend and accelerates the consumption of ephemeral ports.

Tradeoffs and when to use it

For HTTP upstreams, keepalive should be the default. The exceptions are specific and worth understanding.

The keepalive directive caps only idle connections, not total active connections. A setting of keepalive 32 does not mean “at most 32 connections to this upstream.” It means “cache up to 32 idle connections per worker.” Active connections in flight are not counted against this limit. You must size the upstream server to handle active load plus the sum of all worker idle pools.

Because pools are per-worker, load balancers with many workers can open a surprising number of idle connections against a small backend. If your upstream is resource-constrained, you may need to reduce the keepalive count or use the max_conns parameter on the server directive to limit total active connections.

Persistent HTTP/1.1 connections are a prerequisite for HTTP request smuggling. If your backend parses chunked encoding or Content-Length headers differently than nginx, reusing connections across requests can enable desync attacks. In high-security environments with untrusted or ambiguous backends, you may choose to disable keepalive or ensure the backend strictly follows RFC 7230.

The local parameter, now default since 1.29.7, trades efficiency for isolation. Connections are not shared across location blocks, which prevents certain cross-location state issues but increases total idle connections against the upstream. If you have many location blocks hitting the same upstream and backend capacity is tight, you may want to test removing local.

SCGI and uwsgi upstreams cannot use keepalive. FastCGI requires an extra directive. Memcached upstreams use standard keepalive without header manipulation.

Signals to watch in production

SignalWhy it mattersWarning sign
$upstream_connect_timeZero means keepalive reuse; nonzero means new handshakeP95 consistently above 0 for local upstreams
Upstream keepalive hit rateFraction of requests reusing a cached connectionBelow 70 percent indicates significant overhead
TIME_WAIT socket countSockets held after close; accumulate without keepaliveThousands of entries targeting upstream ports
Ephemeral port utilizationExhaustion blocks new upstream connectionsApproaching 80 percent of ip_local_port_range
Active connection utilizationProxy mode uses two slots per requestSustained above 80 percent of worker_connections times worker_processes divided by 2
$upstream_response_timeCascade failures start when backends slow and pools emptyP95 trending up while connect time stays flat

To approximate keepalive hit rate from access logs, compare $upstream_connect_time values. A value of 0.000 strongly indicates reuse. Very fast local connections under one millisecond may occasionally be misclassified, so treat this as an approximation rather than an exact ratio.

How Netdata helps

Netdata collects nginx stub_status metrics and system-level TCP metrics. Use them to detect keepalive problems before they become outages:

  • Correlate nginx active connections with $upstream_connect_time from access logs. Flat or rising connect time while active connections are stable means your keepalive pool is ineffective.
  • Monitor kernel TIME_WAIT sockets and ephemeral port usage on the nginx node. A climbing TIME_WAIT count with steady traffic is a leading indicator that upstream connections are not being reused.
  • Watch the accepts-handled gap from stub_status alongside connection slot utilization. If slots are filling and connect time is nonzero, you may be hitting the proxy connection multiplier limit while also failing to reuse upstream sockets.
  • Alert on sustained nonzero $upstream_connect_time for upstreams that should be local and fast.
The Netdata solution

Web server monitoring with Netdata

Netdata monitors NGINX with per-second request, connection, and latency metrics plus ML anomaly detection. Correlate connection and file-descriptor exhaustion, upstream cascade failures, buffer spill, and TLS CPU with the host signals behind them.