The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / redis / how-redis-works-in-production

Operations Guides

How Redis actually works in production: a mental model for operators

Production incidents do not come from the Redis API. They come from invisible internals: a fork that doubles memory usage, a replication backlog that wraps around, a single slow command that freezes every client. You cannot debug a cascading failure at 3 a.m. without knowing which abstractions compete for which resources.

What it is and why it matters

Redis is a single-threaded event loop around an in-memory dataset. Every incident traces back to resource competition inside one process: memory consumed by the dataset, client buffers, replication backlogs, and allocator fragmentation; CPU consumed by command execution, active expiry, and defragmentation; disk I/O consumed by AOF fsync and RDB snapshots; and network bandwidth consumed by replication and Pub/Sub fan-out.

Without this model, you will misidentify a memory fragmentation spike as a leak, a replication backlog overflow as network instability, or a fork latency spike as generic CPU saturation.

How it works

flowchart TD
    CLIENTS[Client sockets] --> EVENT[Event loop
main thread] EVENT --> COMMANDS[Command execution] COMMANDS --> KEYSPACE[Keyspace hash table] KEYSPACE --> JEMALLOC[jemalloc allocator] EVENT --> BUFFERS[Client buffers] BUFFERS --> JEMALLOC KEYSPACE --> EXPIRY[Lazy + active expiry] JEMALLOC --> FORK[Fork for RDB/AOF
COW duplication] EVENT --> BACKLOG[Replication backlog
PSYNC2] KEYSPACE --> CLUSTER[16384 hash slots
CRC16 modulo] BACKLOG --> REPLICAS[Replica sync]

The single-threaded event loop. Redis uses epoll, kqueue, or IOCP to multiplex client sockets on one main thread. Since Redis 6.0, I/O threads can handle network read and write in parallel, but command execution remains strictly single-threaded. Any command that takes too long blocks every other client. Check SLOWLOG GET for offenders, but remember queue wait time is invisible there.

The keyspace. Keys are stored in a hash table that resizes in powers of two. Rehashing is incremental but adds CPU overhead; activerehashing yes lets the server continue serving reads and writes during resize, but expect latency jitter. Expires are handled lazily on access and actively by sampling random keys on every server cron cycle. The cycle frequency is set by hz (default 10). If keys expire faster than the active cycle deletes them, memory grows and CPU spikes.

Memory allocation. Redis compiles against jemalloc on Linux by default. It does not return memory to the OS eagerly. The gap between used_memory and used_memory_rss reflects allocator free lists, dirty pages, and internal fragmentation. Active defragmentation is available only with jemalloc; switching to libc malloc disables it.

Client buffers. Every connection carries a query buffer governed by client-query-buffer-limit (default 1 GB) and an output buffer controlled by client-output-buffer-limit per class: normal, replica, and pubsub. Normal clients default to unlimited output. During a spike, run CLIENT LIST and inspect the omem field per connection. A forgotten MONITOR client, a slow Pub/Sub subscriber, or a replica that cannot drain its stream allocates from the main heap and can cause sudden RSS growth.

Persistence forks. For RDB snapshots or AOF rewrites, Redis calls fork() to create a child process that serializes the dataset while the parent continues serving commands. Copy-on-write duplicates pages as they are modified. Under a write-heavy workload, RSS can temporarily approach double the dataset size. Disable Transparent Huge Pages before starting Redis; THP multiplies fork latency. During the fork itself, the main thread freezes; latency is proportional to dataset size.

Replication backlog. The primary maintains a fixed-size circular backlog controlled by repl-backlog-size (default 1 MB). Replicas identify themselves by a replication ID and an offset. If the offset falls within the backlog window, a partial resync resumes without a full RDB transfer. Since Redis 4.0, PSYNC2 retains the old master’s replication ID and a second replication offset on promoted replicas, enabling partial resync with downstream replicas after failover. If the backlog wraps before a replica reconnects, a full resync is required, triggering another fork.

Cluster slots. In cluster mode, data is sharded across 16384 hash slots. Keys map via CRC16(key) & 16383. Hash tags using curly braces pin related keys to the same slot, enabling multi-key commands. Gossip uses the client port plus 10000 and must be reachable between all nodes. Partial slot coverage produces a partial outage.

Where it shows up in production

Event loop blocking shows up as uniform latency across all commands. A single KEYS *, a large SMEMBERS, or an unoptimized Lua script freezes the entire server. If SLOWLOG GET shows repeated offenders while overall ops per second drops, the event loop is wedged. Use redis-cli --latency-history to measure queueing delay; if the p99 exceeds the slowlog execution times, the loop is saturated.

Memory pressure shows up in the gap between used_memory and used_memory_rss. The OS OOM killer uses RSS, not Redis’s logical allocator count. In containers, compare RSS to the cgroup limit, not just maxmemory. Fragmentation ratios above 1.5 waste physical RAM and reduce headroom for COW during fork. On instances over 100 MB, a ratio below 1.0 means swapping, which destroys latency. Replicas do not expire keys independently; since Redis 3.2 they mark keys logically expired on read, but the key remains in memory until the primary propagates a DEL. A high-TTL workload can inflate replica memory well beyond the primary’s live dataset size.

Client buffer bloat shows up as sudden RSS spikes even when the keyspace is stable. The default client-output-buffer-limit normal is 0 0 0 (unlimited). A slow Pub/Sub subscriber, a replica that cannot drain its replication stream, or a MONITOR session left running consumes memory from the main heap until either the limit disconnects the client or Redis is OOM killed.

Fork COW pressure shows up during scheduled BGSAVE, AOF rewrite, or an unplanned full resync. In containers with tight memory limits, the RSS spike from dirty page duplication triggers the OOM killer. The fork latency itself, tracked in latest_fork_usec, freezes the main thread. Replicas may time out during long forks, disconnect, and reconnect. If the replication backlog wrapped during the disconnect, another full resync begins.

Replication backlog overflow shows up in INFO stats as sync_full increments and sync_partial_err grows. A brief network blip accumulates more writes than the 1 MB default backlog can hold. The replica reconnects, partial resync fails, and a full resync begins. That resync forks the primary, causing latency. If multiple replicas fall behind simultaneously, the primary enters a fork-latency-resync loop. Increase repl-backlog-size before planned failovers to reduce the chance of full resyncs with downstream replicas.

Cluster slot issues show up as CLUSTERDOWN errors or MOVED and ASK redirects to clients. A node holding a disproportionate share of slots becomes a hot spot. Gossip on port plus 10000 must be reachable between all nodes; missed firewall rules here are the most common cause of cluster_slots_pfail growing.

Tradeoffs and common misuses

Single-threaded command execution trades simplicity and atomicity for a hard throughput ceiling. You cannot add cores to scale command execution. Once main-thread CPU approaches 100% of one core, latency degrades linearly.

jemalloc trades potential fragmentation for allocation performance. Switching to libc malloc eliminates active defragmentation entirely. If your workload has high churn, you need jemalloc and you must monitor fragmentation.

Fork-based persistence trades durability for memory headroom. RDB gives point-in-time snapshots with minimal runtime overhead but large fork cost. AOF gives finer granularity but grows unbounded without rewrite, and rewrite itself requires a fork. Redis 7.0+ uses multi-part AOF with a base file plus incremental deltas to reduce rewrite overhead, but the fork remains.

Replication backlog size trades memory against resync cost. A small backlog saves RAM but guarantees expensive full resyncs after any brief interruption. Production workloads should use at least 100 MB, with write-heavy systems using 512 MB or more.

Hash tags in cluster mode enable multi-key transactions but create slot-level hot spots. A single hash tag receiving heavy traffic pins all load to one node, negating the benefit of sharding.

Common misuses include running KEYS or FLUSHDB without understanding they block the event loop. KEYS scans the entire keyspace. FLUSHDB without the ASYNC flag deletes every key synchronously. Both appear to work in development and destroy latency in production. Using appendfsync always for maximum durability is another misuse in throughput-sensitive environments. Every write waits for fsync, turning disk latency into command latency. The default everysec is the pragmatic compromise, but it requires monitoring aof_delayed_fsync.

Signals to watch in production

SignalWhy it mattersWarning sign
used_memory / maxmemory ratioProximity to memory limit; eviction or write rejection followsRatio > 0.8 trending toward 0.9
mem_fragmentation_ratioAllocator efficiency; high values waste RAM, low values indicate swapSustained > 1.5 or < 1.0 on instances > 100 MB
latest_fork_usecMain thread freeze during RDB/AOF or full resync> 500 ms, or > 20 ms per GB of dataset
evicted_keys rateDataset exceeds memory; cache churn costs CPUSustained rate above baseline, especially with rising misses
sync_full and sync_partial_errBacklog insufficient; full resyncs cost fork latencyAny increase in sync_full or non-zero sync_partial_err
rejected_connectionsHard limit hit; clients are failing immediatelyAny rate > 0
Slowlog growth rateSpecific commands blocking the event loopRepeated entries for KEYS, large SORT, or Lua scripts
aof_delayed_fsyncDisk I/O cannot keep up with appendfsync everysecRate increasing, indicating growing durability window
cluster_state and cluster_slots_failSlot coverage determines availabilitycluster_state:fail or non-zero cluster_slots_fail
connected_clients / maxclientsConnection exhaustion approachingRatio > 0.8

How Netdata helps

  • Charts used_memory and used_memory_rss together to expose fragmentation or COW spikes.
  • Tracks latest_fork_usec alongside rdb_bgsave_in_progress and aof_rewrite_in_progress to isolate fork freeze from slow commands.
  • Surfaces replication offset lag, sync_full, and master_link_status in one view to spot backlog overflow before it forces full resyncs.
  • Correlates slowlog rate and command latency with CPU saturation to distinguish event loop blocking from core exhaustion.
  • Alerts on aof_delayed_fsync and aof_last_write_status to flag disk I/O pressure before write rejection.
The Netdata solution

Redis monitoring with Netdata

Netdata monitors Redis with per-second metrics and ML anomaly detection. Track memory usage and fragmentation, fork/COW latency, replication backlog, evictions, and connection pressure to spot the failure modes in these runbooks early.