The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / nginx / nginx-proxy-buffer-tuning

Operations Guides

NGINX proxy buffer tuning: proxy_buffers, proxy_buffer_size, and busy buffers

When proxy_buffering is on (the default), NGINX absorbs the upstream response in memory before sending it to the client. This shields backends from slow clients and enables compression, but three directives control it: proxy_buffer_size for headers, proxy_buffers for the body, and proxy_busy_buffers_size for the in-flight flush window. Misconfiguration causes 502s, silent disk spills, and reload failures.

The defaults are modest: eight body buffers of one memory page each, and one header page (typically 4K or 8K). That works for static sites and small JSON, but it fails for modern workloads: APIs with large JWT tokens in headers, bulk exports returning multi-megabyte JSON, and Server-Sent Events streams. Undersized body buffers spill to disk. Undersized header buffers return 502. Invalid proxy_busy_buffers_size values prevent NGINX from starting or reloading.

What it is and why it matters

proxy_buffer_size sets the buffer for the upstream response header. The default is one memory page. If headers exceed this value, NGINX logs “upstream sent too big header while reading response header” and returns 502. API workloads with large Set-Cookie headers, Content-Security-Policy directives, or traced request metadata often need more.

proxy_buffers sets the number and size of body buffers per connection. The default is 8 buffers of one memory page. These are allocated per connection in each worker. If the full response fits, NGINX buffers it in memory and flushes to the client at its own pace. If the response exceeds the total size, NGINX writes the remainder to temporary files under proxy_temp_path.

proxy_busy_buffers_size controls how much data can be in the “busy” state: buffers actively being sent to the client while the upstream response is still arriving. Its default is twice the size of either proxy_buffer_size or a single proxy_buffers buffer, whichever is larger. It must be at least as large as the bigger of those two, and no larger than the total proxy_buffers size minus one buffer. Violating this range causes nginx -t to fail with an [emerg] error and the configuration will not load.

proxy_buffering enables or disables the entire mechanism. The default is on. When on, NGINX fills proxy buffers before sending to the client. When off, the response streams synchronously as received.

How it works

Per response, NGINX reads the upstream status line and headers into the single proxy_buffer_size buffer. If they do not fit, the request fails immediately with 502. There is no fallback.

If the headers fit, NGINX reads the body into the cyclic proxy_buffers pool. As each buffer fills, it can be marked busy and flushed to the client, provided the total busy volume does not exceed proxy_busy_buffers_size. This is the high-water mark for data allowed in flight to the client while the upstream is still transmitting.

If the upstream sends data faster than the client receives it, the buffers fill up. Once all proxy_buffers are full and the response is not yet complete, NGINX writes chunks to temporary files on disk under proxy_temp_path. The response is then served from a mix of memory buffers and disk files. The spill is silent: there is no error log entry. Evidence is elevated latency and disk I/O on the temporary file partition. Because the worker writes to disk while still reading from the upstream, a slow partition adds back-pressure to the upstream connection.

When proxy_buffering is off, none of this happens. NGINX streams the response to the client synchronously as it arrives. This is required for Server-Sent Events and long-lived streams, because buffering would stall the connection indefinitely. The trade-off is that the upstream must wait for the client to receive each chunk, and NGINX cannot efficiently apply compression or retry on timeout. With buffering off, $upstream_response_time includes time spent sending to the client, not just time waiting for the upstream.

flowchart LR
    U[Upstream response] --> H{Headers}
    H -->|fit| HS[proxy_buffer_size]
    H -->|exceeds| ERR[502 invalid header]
    HS --> B{Body}
    B -->|fill| PB[proxy_buffers]
    PB -->|flush window| PBB[proxy_busy_buffers_size]
    PBB --> C[Client]
    PB -->|exceeds capacity| DISK[proxy_temp_path]
    DISK --> C
    U -.->|proxy_buffering off| C

Where it shows up in production

The most common symptom of undersized body buffers is silent disk I/O on the proxy_temp_path partition. NGINX does not log a warning when it spills to disk. The signals are elevated $request_time values that far exceed $upstream_response_time, plus disk latency on the temporary file partition. If your application returns large JSON blobs or CSV exports and you have not tuned proxy_buffers, you are likely serving responses from disk. Use iostat -x 1 on that partition during peak load to confirm disk-bound spills.

Undersized header buffers appear as intermittent 502s that correlate with specific endpoints, not upstream downtime. APIs that inject large trace headers or authentication cookies can exceed the default header buffer and trigger the “upstream sent too big header while reading response header” error. Raising proxy_buffer_size to 16K or 32K usually resolves this.

Streaming endpoints fail when proxy_buffering is left on for Server-Sent Events or long-lived HTTP stream locations. The client receives nothing until the upstream closes the connection, which for an event stream may be never. The connection hangs. The fix is disabling buffering for that location, not adjusting buffer sizes. Compression also interferes with streaming: gzip or brotli forces NGINX to accumulate chunks before compressing, which stalls the stream.

ingress-nginx controller v1.13.0 and v1.13.1 shipped a hardcoded proxy-busy-buffers-size default of 8k that conflicted with larger user-defined proxy-buffer-size values, causing nginx -t to fail with [emerg] "proxy_busy_buffers_size" must be less than the size of all "proxy_buffers" minus one buffer (kubernetes/ingress-nginx issue #13598). The controller removed that default in PR #13780 (merged August 2025) so nginx derives the busy-buffer limit from proxy_buffer_size and proxy_buffers again; the fix ships in controller v1.13.2. If you run ingress-nginx, verify your controller version and check that buffer directives align.

Tradeoffs and when to use it

Larger proxy_buffers keep data in RAM and eliminate disk I/O for large responses, but each worker allocates that memory per connection. Sixteen buffers of 256K consume up to 4MB per connection. A worker handling one hundred concurrent proxy connections could allocate 400MB just for proxy buffers. Under high concurrency, this can exhaust memory or push workers toward OOM. Size buffers for your P99 response body, not the maximum theoretical payload.

Total proxy buffer memory per worker roughly equals (proxy_buffers count x size) + proxy_buffer_size + proxy_busy_buffers_size. Calculate this upper bound before raising values in high-concurrency environments.

For API-heavy workloads, raise proxy_buffer_size independently of proxy_buffers. Headers and bodies scale differently. A 16K header buffer with the default 8 x 8K body buffer is often the right starting point for modern APIs.

For Server-Sent Events and long-lived HTTP streams, disable buffering entirely:

location /events {
    proxy_pass http://backend;
    proxy_buffering off;
    proxy_cache off;
    proxy_http_version 1.1;
    proxy_set_header Connection '';
    proxy_read_timeout 86400s;
}

For WebSocket proxying, also add proxy_set_header Upgrade $http_upgrade; and proxy_set_header Connection 'upgrade';.

Compression must also be disabled for these locations, because gzip forces NGINX to accumulate chunks before compressing, defeating the stream.

For proxy_busy_buffers_size, the safest approach is to leave it at the default unless you have a specific reason to change it. If you do change it, remember the enforced constraint. For example, if proxy_buffers is set to 16 256k, the total buffer space is 4MB. Valid proxy_busy_buffers_size must then be at least 256K and at most (16 - 1) * 256k, or 3.75MB. Setting it to 4MB would cause a config test failure. This constraint ensures that at least one full buffer remains available for reading from the upstream while others are busy flushing.

Signals to watch in production

SignalWhy it mattersWarning sign
502 rate with upstream header errorsproxy_buffer_size too small for response headersError log: “upstream sent too big header while reading response header”
$request_time minus $upstream_response_timeLarge gap indicates NGINX overhead, client slowness, or disk flush from buffer spillGap grows on large responses while upstream time stays flat
Disk I/O on proxy_temp_path partitionBuffer overflow writes bodies to diskLatency spikes correlating with large response sizes
Worker RSS memoryEach connection allocates proxy_buffers in fullRSS per worker exceeds baseline for current connection count
Active connections in Writing stateIncludes time flushing buffers to clientWriting dominates with normal upstream latency, indicating output-side delay

How Netdata helps

  • Netdata correlates $request_time and $upstream_response_time from access logs, surfacing the gap that reveals disk spilling or slow client flushes.
  • Disk latency and utilization charts for the proxy_temp_path partition catch silent buffer overflows.
  • Per-worker RSS memory tracking alerts when buffer bloat causes memory growth beyond what connection count explains.
  • Error log monitoring detects “upstream sent too big header” and [emerg] config test failures.
  • The active connection state breakdown (Reading / Writing / Waiting) isolates whether high Writing counts stem from upstream delays or from NGINX flushing to slow clients.
The Netdata solution

Web server monitoring with Netdata

Netdata monitors NGINX with per-second request, connection, and latency metrics plus ML anomaly detection. Correlate connection and file-descriptor exhaustion, upstream cascade failures, buffer spill, and TLS CPU with the host signals behind them.