The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / rabbitmq / rabbitmq-disk-free-limit-default-50mb

Operations Guides

RabbitMQ disk_free_limit at the default 50MB: the production footgun

RabbitMQ ships with disk_free_limit set to 50MB. When free disk space on the node’s data partition drops below that limit, the broker raises a disk alarm and blocks every publisher on every node in the cluster. Consumers keep draining, but message ingestion stops until free space climbs back above the threshold.

The 50MB default exists so development installs on tiny machines do not silently eat a laptop’s disk. On a production server with terabytes of storage, it is not a safety margin; it is a tripwire one inch off the floor. A single log rotation, an Erlang crash dump, or a burst of persistent messages can consume 50MB in seconds, and the resulting incident looks exactly like a broker outage: publishers frozen, applications timing out, queue depth flat at zero growth while upstream systems back up.

What the disk alarm actually does

The disk alarm is one of RabbitMQ’s two cluster-wide resource alarms (the other is the memory watermark). The mechanics matter because they are more aggressive than most operators expect:

  • It is cluster-wide. The alarm fires on one node, but the block propagates. All publishers on all nodes stop. Connections that attempted to publish show state: blocked; connections with an active alarm that have not yet published show state: blocking.
  • Consumers are not blocked. They keep receiving and acknowledging messages, which is the intended recovery path: drain queues, free space, alarm clears.
  • It is a cliff edge. There is no gradual degradation. Free space crosses the threshold and publishing halts instantly.
  • Free space is checked periodically, not continuously. The broker checks disk space roughly every 10 seconds, increasing the frequency as the limit is approached. A fast-filling disk can overshoot the limit between checks, and after you free space it can take until the next check for the alarm to clear.

That last point has a sharp edge the official documentation calls out: if the limit is set too low and messages are being paged out rapidly, RabbitMQ can run out of disk and crash in the gap between checks. A low limit does not just fail to protect you; it can fail to protect you while also blocking your publishers.

Why 50MB is dangerous on a real server

The failure is almost never “the message store grew to fill a terabyte disk.” It is something mundane landing on the data partition:

  • Log rotation or a burst of log volume. RabbitMQ logs, plus anything else writing to the same partition, can blow through 50MB during a single noisy event (a connection storm, an auth failure flood).
  • An Erlang crash dump. When the BEAM VM crashes it writes erl_crash.dump, which on a large node can be gigabytes. One crash can both fill the partition and block the restart path.
  • Quorum queue WAL growth. Quorum queues append every operation to a write-ahead log. Under sustained throughput, or when a follower falls behind and compaction stalls, WAL segments accumulate fast. This is exactly the workload where you least want a 50MB margin.
  • A burst of persistent messages or aggressive paging to disk under memory pressure.
  • Backups, snapshots, or co-located processes writing to the same partition.

There is also a nastier variant: the disk alarm death spiral. When publishing stops, some of the work that would eventually reclaim disk (compaction, cleanup) can stall or fall further behind, so the node does not recover on its own and you are doing manual cleanup at 3 a.m. A 50MB limit turns a non-event into that scenario.

flowchart TD
  A[Log rotation, crash dump, WAL burst, or backup writes to data partition] --> B[disk_free drops below 50MB default]
  B --> C[Disk alarm fires on one node]
  C --> D[All publishers blocked cluster-wide]
  D --> E[Publish rate drops to zero, upstream systems back up]
  C --> F[Consumers keep draining]
  F --> G{Enough space freed?}
  G -->|Yes| H[Alarm clears on next disk check]
  G -->|No, or crash between checks| I[Manual cleanup or node crash]
  I --> J[erl_crash.dump fills partition further]

Choosing a sane value

The goal is a limit large enough that routine, non-broker events can never trip it, and that gives you real runway to react before the disk is actually full.

Relative to RAM. Set disk_free_limit.relative to 1.0-2.0, meaning 1-2x system RAM. The reasoning: under memory pressure RabbitMQ pages queue contents to disk, so the disk should be able to absorb on the order of the broker’s full memory footprint plus headroom.

Absolute. Set a fixed value of at least 2GB, sized up for your workload and partition. Absolute values are predictable: you know exactly what will trip the alarm, regardless of what hardware the node lands on.

A word of caution on relative in containers. The relative limit is computed from the memory RabbitMQ believes it has. In containerized deployments, the memory visible via cgroups depends on the Kubernetes resource configuration, and there are cases where the cgroup reports the entire VM’s memory rather than the container’s limit. That means relative = 1.0 can silently become “1x the host’s RAM,” which may be far more disk than the partition even has, or a number you never intended. If you run RabbitMQ in containers, prefer an absolute limit unless you have verified what memory the broker actually detects.

Precedence when both are set. If disk_free_limit.absolute and disk_free_limit.relative are both present, the absolute value wins. This was not always true: older releases had a bug where the relative value overrode the absolute one, which bit operators whose base images or Helm charts shipped a relative default (some charts set relative = 1.0) that silently overrode their configured absolute value. The fix landed in 3.11.5. If you are on anything older, check for a stray relative setting in your config chain before trusting your absolute value. Relatedly, some vendor images load config files in an order where a later file’s relative setting overrides your absolute one; if your absolute limit appears to be ignored, inspect the effective config rather than assuming RabbitMQ is broken.

The Kubernetes Cluster Operator changed its default to 2GB years ago, so operator-managed clusters are less exposed. The 50MB default persists in the core broker itself, including 4.x.

Setting it

In rabbitmq.conf:

# Pick one. Absolute is preferred in containers.
disk_free_limit.absolute = 2GB

# Or relative to detected RAM (1.0 = 1x RAM):
# disk_free_limit.relative = 1.0

You can also change it at runtime without a restart:

# Runtime change: absolute
rabbitmqctl set_disk_free_limit 2GB

# Runtime change: relative to RAM
rabbitmqctl set_disk_free_limit mem_relative 1.0

The runtime form is useful for emergencies (you need publishers unblocked now and cannot free space immediately), but it does not survive a broker restart. Treat it as a bridge and put the real value in rabbitmq.conf before the next restart finds you unprotected again.

Verify what the broker is actually enforcing, not what you think you configured:

# Check effective limit, free space, and alarm state per node
curl -s -u guest:guest http://localhost:15672/api/nodes | \
  jq '.[] | {name, disk_free, disk_free_limit, disk_free_alarm}'

# Cross-check at the OS level against RabbitMQ's view
df -h /var/lib/rabbitmq

If disk_free_limit in the API does not match your config file, suspect the precedence issue above: something else in the config chain set a relative or absolute value after yours.

Alerting: do not inherit the footgun into your thresholds

Raising the limit is half the job. The other half is making sure your alerting floor is not anchored to the same broken default.

A common pattern is to alert when free space drops below some multiple of the limit, for example warning at 3x disk_free_limit. With the 50MB default, “3x the limit” is 150MB. On a modern server that alert is noise-adjacent and still far too late to give you meaningful runway.

Use a floor that combines the relative margin with an absolute minimum:

  • Warn when disk_free < max(3 * disk_free_limit, 1GB)
  • Page when disk_free_alarm is true, sustained for more than a minute, with evidence of traffic impact (publish rate was non-zero, or connections are in blocked/blocking state). The sustain and traffic conditions filter out cold starts and idle instances where the alarm fires but nobody is affected.

Two more things worth tracking:

  • Rate of consumption, not just the level. Runway is (disk_free - disk_free_limit) / consumption_rate. A partition losing 100MB/minute is a very different incident from one losing 1MB/minute at the same absolute free space.
  • The quorum WAL directory specifically. Aggregate free space hides which consumer is growing. Quorum queue WAL segments live in their own subdirectory and can grow much faster than classic queue storage.

Signals to watch in production

SignalWhy it mattersWarning sign
disk_free and disk_free_limit per node (API)Shows actual runway against the enforced limit, not the configured onedisk_free < max(3x limit, 1GB), or limit reads 50MB on a prod node
disk_free_alarm (API, boolean)The cluster-wide publishing halt itselftrue sustained >60s with publish traffic present
Connection states (blocked/blocking)Confirms publishers are actually frozen, distinguishing real impact from an idle-instance alarmAny non-zero count alongside an active alarm
Publish rate vs deliver_get rateDuring a disk halt, publish drops to zero while consumer drain continuesPublish = 0 with stable connection count and non-zero ack rate
Disk consumption rate on the data partitionConverts a static level into a time-to-alarm estimateRunway under your response time at current burn rate
Quorum queue WAL directory sizeThe fastest-growing disk consumer on quorum-heavy nodesGrowth outpacing snapshot/compaction, especially with a lagging follower

The publish-rate pattern is how you distinguish a disk halt from a publisher-side failure: in a disk alarm, connections stay up, publish rate goes to zero, and deliver/ack rates keep moving. The RabbitMQ mental model guide covers how this composite differs from the memory wall, which looks similar but blocks on mem_alarm instead.

Prevention checklist

  • Set the limit explicitly. Absolute 2GB minimum, or relative 1.0-2.0 on bare metal/VMs where detected RAM is trustworthy. Never inherit the 50MB default into production.
  • Verify the effective value via the API after every config change and after upgrades, especially on older versions or vendor images with layered config files.
  • Prefer absolute in containers. Cgroup memory detection makes relative limits unpredictable under Kubernetes.
  • Isolate the data partition. Keep RabbitMQ’s data directory on its own partition so co-located logs, backups, and crash dumps from other software cannot trip the broker’s alarm.
  • Alert on the floor, not the limit. disk_free < max(3 * disk_free_limit, 1GB) as a warning, the alarm itself as a page with sustain and traffic conditions.
  • Know your runtime override. rabbitmqctl set_disk_free_limit buys you breathing room mid-incident; pair it with a permanent config change.

How Netdata helps

  • Netdata collects per-node disk_free, disk_free_limit, and disk_free_alarm from the management API, so you can see runway and alarm state per node in one view instead of querying each broker by hand.
  • Correlating the alarm flag with publish, deliver, and ack rates makes the disk-halt signature obvious: publish flat at zero while consumer drain continues.
  • Connection state tracking surfaces the blocked/blocking population, confirming whether an active alarm is actually impacting traffic.
  • OS-level disk metrics for the data partition, alongside broker-reported free space, catch the divergence cases where RabbitMQ’s view and the filesystem’s view disagree, and let you compute burn rate and time-to-alarm.
  • Historical retention lets you see whether a limit was silently reset, for example after a restart dropped a runtime-only set_disk_free_limit change.

For the full signal inventory this fits into, see the RabbitMQ monitoring checklist and the monitoring maturity model.