The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / redis / redis-misconf-rdb-snapshots

Operations Guides

MISCONF Redis is configured to save RDB snapshots - what it means and how to fix it

Applications see MISCONF Redis is configured to save RDB snapshots, but it's currently unable to persist to disk (since Redis 7.0; Redis 4.0-6.2 prints MISCONF Redis is configured to save RDB snapshots, but it is currently not able to persist on disk) on every write. Reads still work, but SET, HSET, LPUSH, and all mutating commands are rejected.

Redis makes itself read-only when the last background save failed and stop-writes-on-bgsave-error is yes (the default). The error persists until a subsequent BGSAVE succeeds. Retrying writes will not help. The instance is protecting you from accepting writes that can never be persisted. To recover, fix the underlying persistence failure and clear the error state with a successful save.

What this means

Redis persists data to disk by forking a child process that writes the dataset to an RDB file. This happens automatically based on the save directives in redis.conf, or manually via BGSAVE. If the child fails, rdb_last_bgsave_status flips to err and stays there until a background save succeeds; it does not reset on client reconnections or between save attempts. A Redis process restart resets it to ok, but the restart does not fix the underlying failure - the next failed save flips it back to err.

With stop-writes-on-bgsave-error yes (the default), a failed background save stops all writes. If the server cannot snapshot its data, continuing to accept writes creates a silent durability gap. Every write accepted after the failure would be lost in a crash.

The MISCONF error is a safety mechanism, not a configuration mistake. Do not simply disable the setting unless you are running a pure cache that can repopulate from an authoritative source. Resolve why BGSAVE cannot complete.

flowchart TD
    A[Client sees MISCONF on write] --> B{Check INFO persistence}
    B -->|rdb_last_bgsave_status:err| C{Check disk space}
    C -->|Full or quota| D[Free disk space]
    C -->|Adequate space| E{Check logs and fork}
    E -->|fork failed| F[Fix memory or overcommit]
    E -->|Permission denied| G[Fix dir or file ownership]
    E -->|Child OOM killed| H[Add memory headroom for COW]
    D --> I[Trigger manual BGSAVE]
    F --> I
    G --> I
    H --> I
    I -->|BGSAVE ok| J[Writes resume]
    I -->|BGSAVE fails| K[Re-check root cause]

Common causes

CauseWhat it looks likeFirst thing to check
Disk full or I/O errorsrdb_last_bgsave_status:err, disk utilization at 100%, or filesystem remounted read-only.df -h on the volume holding the Redis dir.
Fork failure from memory pressureLog contains Can't save in background: fork: Cannot allocate memory. Common on large instances or containers with tight memory limits.vm.overcommit_memory and the gap between used_memory_rss and total RAM.
Permission denied on RDB pathThe child process cannot open or rename the RDB file in the configured directory. Often after permission changes or SELinux policy updates.Ownership and mode of the dir and dbfilename path.
Child process killed by OOM killerThe forked child exceeds container or host memory limits during copy-on-write and is killed mid-save.dmesg or kernel logs for OOM killer events targeting the Redis child.
Conflicting RDB paths in containersMultiple Redis instances share a volume and fight over the same dump.rdb file. Common in Docker Compose or misconfigured StatefulSets.Whether CONFIG GET dir and dbfilename point to unique paths per instance.

Quick checks

Run these read-only checks to confirm state and narrow the cause.

# Confirm the error state and see when the last save succeeded
redis-cli INFO persistence | grep -E "rdb_last_bgsave_status|rdb_last_save_time|rdb_bgsave_in_progress"

# Check if write errors are accumulating
redis-cli INFO stats | grep total_error_replies

# Check per-error counters on Redis 6.2+
redis-cli INFO errorstats

# Check free disk space on the persistence volume
df -h "$(redis-cli CONFIG GET dir | tail -1)"

# Check RDB file and directory ownership
ls -l "$(redis-cli CONFIG GET dir | tail -1)/$(redis-cli CONFIG GET dbfilename | tail -1)"

# Check if the kernel allows fork overcommit
cat /proc/sys/vm/overcommit_memory

# Check for recent OOM killer or fork failures in kernel logs
dmesg | grep -iE "oom|fork|redis"

How to diagnose it

  1. Verify the sticky error. Run redis-cli INFO persistence and look for rdb_last_bgsave_status:err. If it is ok but writes are still failing, the issue is likely AOF-related or the error is coming from a replica with a different condition. If it is err, note rdb_last_save_time to see how stale the last successful snapshot is.

  2. Check disk space. Redis needs enough free space to write a temporary RDB file before atomically renaming it to the final file. Run df -h on the persistence volume. A full volume is the root cause. Also check inode usage with df -i on smaller or heavily fragmented filesystems.

  3. Inspect logs for fork failures. If disk space is adequate, check the Redis server log and dmesg for messages like fork: Cannot allocate memory or OOM killer events. The Redis child process duplicates page tables during fork(). On a memory-heavy instance, this can fail if vm.overcommit_memory is 0 or if a container memory limit leaves no room for page table overhead.

  4. Validate permissions. Ensure the user running redis-server has write access to the directory returned by CONFIG GET dir and can create and rename files there. Do not assume the directory is correct after a migration or container restart.

  5. Check for path collisions. If you are running multiple Redis instances on the same host or shared volume, confirm each instance has a unique dir and dbfilename. Two instances writing the same dump.rdb simultaneously will corrupt it or cause one to fail.

  6. Assess whether AOF is active. If aof_enabled is 1 and aof_last_write_status is ok, you still have append-only durability while RDB is failing. This does not fix the MISCONF block, but it changes the urgency: you are not fully without persistence. If AOF is also failing, treat this as a critical data-loss risk.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
rdb_last_bgsave_statusBinary health of the last background save.err persisting beyond one save interval.
Age of rdb_last_save_timeHow long since the last successful snapshot.Older than 2x the configured save interval.
Disk free on persistence volumeRDB and AOF need sufficient free space for temp files.Less than 3x the current dataset size.
latest_fork_usecThe main thread is blocked during fork().Greater than 500ms, or trending upward with dataset growth.
total_error_replies rateDirect evidence that clients are seeing failures.Increasing while rdb_last_bgsave_status is err.
used_memory_rss vs host memoryPredicts fork failure and COW-driven OOM kills.RSS exceeding 80% of available RAM.

Fixes

Disk full or I/O errors

Free space on the persistence volume. Delete old logs, oversized backups, or orphaned temp files left by interrupted saves. After cleanup, trigger a manual save:

redis-cli BGSAVE

Watch rdb_bgsave_in_progress and rdb_last_bgsave_status until it returns ok. Do not simply restart Redis; a restart with a full disk will fail to save during shutdown and may leave you with no recent RDB on startup.

Fork failure or memory pressure

If the fork fails with a memory error and the host is not actually out of RAM, the most common cause is vm.overcommit_memory=0. Set it to 1 to allow the kernel to overcommit memory for the page tables Redis needs during fork():

sysctl -w vm.overcommit_memory=1

Make it permanent in /etc/sysctl.conf. This is a system-wide setting. If Redis runs inside a container, increase the container memory limit to account for copy-on-write overhead. Plan for at least 50% headroom above used_memory_rss on persistent instances; heavy writes during a fork can temporarily double RSS.

If freeing memory or adding capacity is not immediately possible and you need writes to resume, see the temporary workarounds below. Only use them if you understand the durability tradeoff.

Permission or ownership issues

Fix ownership so the Redis process can write to its configured directory:

chown redis:redis /var/lib/redis

If you use SELinux or AppArmor, check for denials in the audit log and adjust the profile. After fixing permissions, run BGSAVE and verify rdb_last_bgsave_status returns ok.

Path collisions in containers

If multiple instances target the same dump.rdb, reconfigure each instance with a unique dir or unique dbfilename, then restart the affected instances. Restarting is acceptable here because the root cause is configuration, not resource exhaustion.

Temporary workarounds and their tradeoffs

If you must unblock writes immediately and you are running a pure cache workload where data can be reconstructed from another source, you can disable the write block:

redis-cli CONFIG SET stop-writes-on-bgsave-error no

This makes the instance accept writes again even though RDB saves are failing. Your data loss window widens to the time since the last successful save. If you do this, also call CONFIG REWRITE to persist the change, or it will revert on restart.

Alternatively, if you rely on AOF for durability and do not need RDB snapshots at all, you can disable automatic RDB saves:

redis-cli CONFIG SET save ""

Again, follow with CONFIG REWRITE if you want the change to survive restart.

Do not use these workarounds on primary databases where RDB is the only persistence mechanism. The correct path is to fix the underlying failure and then run a successful BGSAVE.

Prevention

  • Monitor disk space proactively. Maintain at least 3x the dataset size as free space on the persistence volume to accommodate temp files, AOF rewrites, and RDB snapshots simultaneously.
  • Set vm.overcommit_memory=1 on every Redis host. Fork failures on healthy machines are almost always caused by this setting.
  • Size memory for copy-on-write. If you use RDB or AOF rewrite, keep used_memory_rss below roughly 50% of physical RAM on persistent instances. On cache-only instances, 75% is a safer ceiling.
  • Audit persistence completion. Do not assume that because save directives exist, saves are succeeding. Alert on rdb_last_bgsave_status:err and on stale rdb_last_save_time.
  • Validate container and orchestration configs. Ensure each Redis pod or container has a distinct persistent volume and does not share a dump.rdb path with another instance.

How Netdata helps

  • Netdata collects rdb_last_bgsave_status and rdb_last_save_time from INFO persistence, so you can alert on a sticky err before application write failures spike.
  • Disk space and memory metrics on the same node are correlated with Redis persistence health, making it faster to distinguish a disk-full incident from a fork-failure incident.
  • total_error_replies and throughput charts let you confirm the exact moment writes started failing and whether the failure correlates with a scheduled BGSAVE.
  • RSS and used_memory are plotted together, which helps you forecast when copy-on-write overhead during a fork will exceed available RAM.
  • Fork duration (latest_fork_usec) is tracked over time, revealing when Transparent Huge Pages or memory pressure are slowing the background save process before it fails entirely.
The Netdata solution

Redis monitoring with Netdata

Netdata monitors Redis with per-second metrics and ML anomaly detection. Track memory usage and fragmentation, fork/COW latency, replication backlog, evictions, and connection pressure to spot the failure modes in these runbooks early.