The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / redis / redis-swapping-fragmentation-below-one

Operations Guides

Redis mem_fragmentation_ratio below 1.0: detecting swap death

A mem_fragmentation_ratio of 0.72 on a substantial dataset means the operating system has swapped out Redis memory pages. Redis stays alive and responds to PING, but every command touching a swapped key blocks the single event loop on disk I/O. Because Redis logs nothing about swap, the resulting latency catastrophe looks like a mystery.

mem_fragmentation_ratio equals used_memory_rss / used_memory. When the ratio drops below 1.0, the resident set size is smaller than the memory Redis requested from its allocator. The missing bytes are in swap. Operators typically watch for ratios above 1.5, so a low ratio is misread as “good fragmentation” when it is actually the worst memory-related failure mode short of an OOM kill.

The operational threshold is roughly 0.8 on instances with used_memory above 100 MB. Smaller instances can show ratio noise from process overhead, but on a production dataset a sustained value below 0.8 means swap death.

What this means

used_memory tracks what Redis requested from its allocator. used_memory_rss tracks what the OS kept in RAM. When RSS falls below allocated memory, the kernel paged some of Redis’s anonymous memory out to disk.

Redis has no awareness that its pages are on swap. The event loop continues accepting commands, but when it touches a swapped key, value, or internal structure, the thread blocks waiting for the storage layer to page data back in. Because command execution is single-threaded, one swapped access stalls every other client. The latency is a wall, not a slope.

Swapped pages persist even after host memory pressure subsides. The kernel does not bring them back automatically when free RAM becomes available; they stay on disk until Redis touches them again. The latency hit can outlast the original pressure by hours.

Common causes

CauseWhat it looks likeFirst thing to check
Host memory overcommit or competing processesRatio drops as other processes allocate RAM; system swap usage grows; latency climbs silentlyfree -h and /proc/swaps
Fork COW pressure from persistenceRatio drops during or immediately after BGSAVE or AOF rewrite; latest_fork_usec spikedrdb_bgsave_in_progress, aof_rewrite_in_progress, and host RAM headroom
vm.swappiness too highGradual ratio decline under normal load; no sudden memory spike in Rediscat /proc/sys/vm/swappiness
Container memory limits without host headroomRatio drops inside a container while the host shows free RAM; OOM kills may followContainer cgroup memory limits versus used_memory_rss
NUMA misconfigurationIntermittent ratio drops on large bare-metal hosts under loadnumactl --hardware or /sys/devices/system/node/

Quick checks

# Confirm ratio and instance size
redis-cli INFO memory | grep -E "mem_fragmentation_ratio|used_memory:"
# Check system swap usage and available RAM
free -h && swapon --show
# Check kernel tendency to swap anonymous pages
cat /proc/sys/vm/swappiness
# Check per-process swap consumption for the Redis process
cat /proc/$(pidof redis-server | awk '{print $1}')/status | grep VmSwap
# Check allocator metrics to rule out non-swap artifacts (Redis 4.0+)
redis-cli INFO memory | grep -E "allocator_frag_ratio|allocator_rss_ratio"
# Check for active persistence forks that may have triggered COW bloat
redis-cli INFO persistence | grep -E "rdb_bgsave_in_progress|aof_rewrite_in_progress"

All of these are read-only. None change server state.

How to diagnose it

  1. Validate the signal. On instances with used_memory above 100 MB, a mem_fragmentation_ratio below 0.8 indicates swap. On smaller instances the ratio may be ambiguous due to process overhead; corroborate with OS metrics before treating as swap death.

  2. Confirm OS-level swap. Run free -h and check that swap used is non-zero or trending upward. Check VmSwap in /proc/<pid>/status for the Redis process. Non-zero VmSwap is ground truth that the kernel has moved Redis pages out of RAM.

  3. Rule out allocator artifacts. On Redis 4.0+, check allocator_frag_ratio and allocator_rss_ratio. If these sit in normal ranges while mem_fragmentation_ratio is depressed, the low ratio is not an allocator artifact. It is swap.

  4. Correlate with persistence events. Check whether the ratio drop coincides with rdb_bgsave_in_progress=1 or aof_rewrite_in_progress=1. A fork() doubles RSS via copy-on-write; if the host had no headroom, the kernel may have swapped Redis parent pages to make room for COW overhead. This is likely when latest_fork_usec spiked immediately beforehand.

  5. Check system memory pressure history. Inspect dmesg for OOM killer activity or memory-reclaim messages around the time the ratio dropped. Run vmstat 1 and look for sustained si and so columns indicating active swap-in and swap-out.

  6. Determine whether pressure is ongoing or residual. If host free RAM has recovered but mem_fragmentation_ratio remains low, the pages are still swapped. Redis never touched them to trigger a page fault and bring them back. The event loop is running on borrowed time until the next access to a cold page stalls every client.

flowchart TD
    A[mem_fragmentation_ratio < 1.0] --> B{used_memory > 100MB?}
    B -->|No| C[Ambiguous: process overhead dominates]
    B -->|Yes| D[Check OS swap usage]
    D --> E[Swap confirmed]
    E --> F[Identify memory pressure source]
    F --> G[Restart Redis to reload pages in RAM]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
mem_fragmentation_ratioPrimary indicator comparing RSS to allocated memory< 0.8 while used_memory > 100 MB
used_memory_rssWhat the OS sees; the OOM killer and swap subsystem use thisSustained divergence below used_memory
used_memoryAllocator-reported consumption; provides context for the ratioValidates whether ratio thresholds are meaningful
allocator_frag_ratio (Redis 4.0+)True allocator fragmentation, separate from OS swap behaviorNormal value while mem_fragmentation_ratio is low confirms swap
latest_fork_usecCOW fork latency; spikes precede memory pressure events> 500 ms before a ratio drop suggests COW-induced swap
OS swap usedGround truth for whether swapping is activeNon-zero and correlated with ratio below 1.0
rdb_bgsave_in_progress / aof_rewrite_in_progressForks trigger COW that can push a tight host into swapRatio dropping during or after persistence operations

Fixes

Immediate: stop the memory pressure

Identify and terminate or migrate the non-Redis memory consumers that pushed the host over the edge. Stopping the pressure prevents additional pages from being swapped. It does not automatically bring already-swapped Redis pages back into RAM.

Recover the pages

Restart Redis. On startup, the dataset loads from RDB or AOF back into fresh physical memory. All swapped pages are abandoned by the process.

If a restart is not immediately feasible and you have root access, swapoff -a && swapon -a forces the kernel to move swapped pages back to RAM if capacity exists. Warning: if insufficient free memory exists, this command can hang for minutes or hours and may trigger the OOM killer. Do not run it on a memory-constrained host without an escape plan.

Do not rely on MEMORY PURGE. It instructs jemalloc to release dirty pages to the OS, but swapped pages are already on disk, not in allocator arenas. It will not reload them.

Right-size the host

Persistent instances need roughly 50% headroom above used_memory_rss to survive fork copy-on-write without pressuring the OS into swap. Cache-only instances still need headroom for client buffers, replication backlogs, and fragmentation. If the host cannot provide this, shard the dataset or move to larger hardware before the next persistence event triggers the same failure.

Prevention

  • Set vm.swappiness to 0 or 1. This reduces the kernel’s tendency to swap anonymous pages, though it does not eliminate swap under severe memory pressure.
  • Account for COW in capacity planning. For instances with RDB or AOF enabled, maintain used_memory below 50% of physical RAM. For cache-only instances, stay below 75% of available memory.
  • Disable Transparent Huge Pages. THP is the most common cause of excessive fork latency and COW bloat. A fork that takes ten times longer than necessary increases the window during which memory pressure can force swap.
  • Set maxmemory and enforce host-level headroom. A Redis instance with no limit grows until the OOM killer intervenes. Even with maxmemory configured, ensure the host has enough RSS headroom that the OS never needs to reclaim Redis pages.
  • Monitor RSS, not just used_memory. The swap subsystem and the OOM killer operate on RSS. A healthy used_memory means nothing if used_memory_rss is being reclaimed.
  • In containerized environments, ensure the container memory limit includes COW overhead. A limit set equal to maxmemory guarantees fork failures or swap death during the next background save.

How Netdata helps

  • Correlates mem_fragmentation_ratio, used_memory, and used_memory_rss on the same charts to expose the swap gap.
  • Alerts on mem_fragmentation_ratio < 0.8 when used_memory > 100 MB, suppressing noise from small instances where process overhead dominates.
  • Surfaces system-level swap metrics alongside Redis metrics to confirm OS-level swapping.
  • Tracks latest_fork_usec spikes that precede COW-induced memory pressure.
  • Displays allocator_frag_ratio and allocator_rss_ratio on Redis 4.0+ instances to help distinguish true swap from allocator artifacts.
The Netdata solution

Redis monitoring with Netdata

Netdata monitors Redis with per-second metrics and ML anomaly detection. Track memory usage and fragmentation, fork/COW latency, replication backlog, evictions, and connection pressure to spot the failure modes in these runbooks early.