The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / redis / redis-monitoring-checklist

Operations Guides

Redis monitoring checklist: the signals every production instance needs

Redis can return PONG while replicating hours behind, during an OOM kill in a background save, or while a KEYS command wedges the event loop. This checklist structures monitoring into four maturity levels. Level 1 is the survival floor. Level 2 adds workload and resource awareness. Level 3 introduces leading indicators that catch degradation before it becomes an incident. Level 4 exposes allocator and encoding internals for granular diagnostics.

Work through the levels in order. Most production incidents are preventable with Level 2 signals that teams never configure. All metrics below are available via standard Redis commands.

flowchart TD
    L1["Level 1 — survival"]
    L2["Level 2 — operational"]
    L3["Level 3 — mature"]
    L4["Level 4 — expert"]
    L1 --> L2
    L2 --> L3
    L3 --> L4

Level 1 — survival

These are binary signals. If any fail, you have an active incident or are one heartbeat away from one. Cover every production instance here before moving on.

PING response. A failed PING means the process is down, unreachable, or the event loop is frozen by a slow command or Lua script. Monitor from the same network path as your application to catch partition issues the node cannot detect itself.

Uptime in seconds. uptime_in_seconds resetting to a low value indicates a crash, OOM kill, or unplanned restart. Use it to gate other alerts during cold-start windows.

Loading state. INFO persistence returns loading:1 while Redis restores an RDB or AOF file. During this phase the instance rejects data commands but still responds to PING. Loading that exceeds your baseline duration indicates disk pressure or an unexpectedly large dataset.

Memory usage versus maxmemory. used_memory approaching maxmemory means eviction or write rejection is imminent. If maxmemory is unset (0), Redis grows until the OS OOM killer terminates it. Set maxmemory explicitly and choose a maxmemory-policy appropriate to your workload; persistent stores should use noeviction, while caches should use allkeys-lru or allkeys-lfu.

Connected clients versus maxclients. connected_clients approaching the effective maxclients limit means new connections will be rejected. The default maxclients is 10000. Replicas and cluster bus connections count toward the limit, and the OS file descriptor ceiling can cap you below the configured value.

Last RDB save status. rdb_last_bgsave_status must be ok. A status of err means your last backup failed. If stop-writes-on-bgsave-error is enabled, writes are already being rejected.

Last AOF write status. aof_last_write_status must be ok when AOF is enabled. A failed AOF write means data is not being persisted and, depending on configuration, writes may soon be blocked.

Master link status (replicas only). master_link_status must be up. A down link means the replica is serving stale data and failover would lose everything written since the disconnect.

Level 2 — operational

These signals move from binary liveness to workload health. They answer whether Redis is doing useful work efficiently and whether resource limits are approaching. Add these only after Level 1 is fully covered.

Operations per second. instantaneous_ops_per_sec establishes your throughput baseline. A sustained drop without a traffic decrease often means the event loop is blocked by a slow command, Lua script, or fork. A spike can precede saturation or signal a traffic pattern shift that invalidates your baseline.

Keyspace hit rate. Compute from keyspace_hits / (keyspace_hits + keyspace_misses). A sustained drop below your baseline after the cold-start window indicates eviction pressure, mass expiry, or a cache stampede. Correlate with evicted_keys and expired_keys rates to distinguish capacity problems from TTL issues.

Evicted keys rate. evicted_keys is a cumulative counter. Any sustained rate of change means the working set exceeds memory. For persistent workloads, any eviction is data loss. For cache workloads, a sudden spike indicates capacity trouble.

Rejected connections rate. rejected_connections incrementing means clients are actively failing to connect. Any increase is an active incident.

Memory fragmentation ratio. mem_fragmentation_ratio above 1.5 indicates wasted memory that brings OOM closer. A ratio below 1.0 indicates part of Redis memory has been swapped to disk, which destroys latency. Confirm with host-level swap metrics.

Latest fork duration. latest_fork_usec above 500 ms means clients experienced a noticeable freeze during the last RDB save or AOF rewrite. Above 1 second, client timeouts are likely. High fork latency on Linux often means Transparent Huge Pages is enabled.

Slowlog growth. SLOWLOG LEN increasing means expensive commands are blocking the event loop. Review SLOWLOG GET to identify specific offenders such as KEYS, SMEMBERS, or unoptimized Lua scripts. Set slowlog-log-slower-than to a threshold that captures queries above your application’s tolerance, typically 10 ms for user-facing workloads.

Expired keys rate. expired_keys rate spikes indicate mass expiry events. Combine this with expired_time_cap_reached_count if available to detect when the active expiry cycle cannot keep up.

Connected replica count. connected_slaves on the primary must match your expected topology. A drop means a replica disconnected, which may trigger a full resync and another fork.

Replication offset lag. Calculate master_repl_offset minus slave_repl_offset. Lag approaching repl-backlog-size means the next reconnection will force a full resync, with associated fork latency and bandwidth.

Network throughput. instantaneous_input_kbps and instantaneous_output_kbps track bandwidth utilization. Asymmetric output spikes can indicate Pub/Sub fan-out, replication pressure, or a forgotten MONITOR session.

Level 3 — mature

These are leading indicators. They detect problems before they trigger Level 1 alarms. They require more instrumentation but separate stable workloads from incidents waiting to happen.

Internal latency events. LATENCY LATEST breaks down delay by event type: command, fork, AOF fsync, and eviction cycle. Enable this by setting latency-monitor-threshold to a low millisecond value in redis.conf or via CONFIG SET. The default of 0 disables it entirely.

Client output buffer memory. CLIENT LIST exposes omem per client. A single client consuming hundreds of megabytes indicates a slow subscriber, a forgotten MONITOR, or an application that cannot read responses fast enough. If omem grows while throughput is flat, the client is likely dead.

Blocked clients count. blocked_clients is expected for queue patterns using BLPOP or XREAD BLOCK. A sustained increase outside your baseline suggests dead consumers or a failed producer.

Keyspace growth trend. INFO keyspace key counts should follow a predictable pattern. Linear or exponential growth without TTL coverage indicates a memory leak or unbounded key creation.

Replication sync quality. sync_full, sync_partial_ok, and sync_partial_err in INFO stats reveal whether replicas are reconnecting cleanly. Any increase in sync_full means the replication backlog is too small. Increase repl-backlog-size if your replicas reconnect frequently enough to exhaust the current buffer.

Copy-on-write memory. rdb_last_cow_size and aof_last_cow_size measure the memory cost of persistence forks. COW exceeding 50 percent of used_memory predicts an OOM kill during the next write spike. If COW grows while write volume is flat, check for large key deletions that dirty pages.

AOF size ratio. aof_current_size / aof_base_size growing above your configured rewrite threshold means AOF rewrites are failing or not triggering. Check aof_rewrite_in_progress and aof_last_rewrite_status to confirm whether a rewrite is stuck, allowing the file to grow unboundedly.

Error statistics. INFO errorstats (Redis 6.2+) breaks down total_error_replies by type. A rising errorstat_OOM count under a noeviction policy confirms writes are being rejected due to memory pressure.

Stream consumer group lag. XINFO GROUPS exposes lag (Redis 7.0+) and pending counts. Growing lag means consumers cannot keep up. Growing pending means entries are delivered but never acknowledged.

Cluster state and slot health. CLUSTER INFO must show cluster_state:ok with all 16384 slots assigned and healthy. Any non-zero cluster_slots_fail is an active outage. Non-zero cluster_slots_pfail is an impending one.

ACL log. ACL LOG (Redis 6.0+) records authentication and authorization failures. Unexplained entries indicate credential rotation gaps or unauthorized access attempts.

Configuration drift. Audit CONFIG GET * against your expected configuration. Changes made via CONFIG SET without CONFIG REWRITE are lost on restart and can mask the root cause of an incident.

Level 4 — expert

These signals expose internal allocator, encoding, and per-client behavior that aggregate metrics hide. Use them when standard dashboards look healthy but sporadic latency or memory growth persists.

Allocator fragmentation ratios. allocator_frag_ratio and allocator_rss_ratio (Redis 5.0+) separate true jemalloc fragmentation from process overhead. Use them when mem_fragmentation_ratio is ambiguous.

Active defrag effectiveness. active_defrag_running, active_defrag_hits, and active_defrag_misses show whether defragmentation is reducing waste or just burning CPU.

Expiry cycle throttling. expired_time_cap_reached_count (Redis 6.0+) increments when the active expiry cycle hits its CPU budget. If this grows, expired keys are accumulating faster than Redis can clean them.

Main thread CPU. used_cpu_user_main_thread and used_cpu_sys_main_thread (Redis 6.2+) isolate command execution CPU from child process and I/O thread usage. Track the rate of change to detect single-core saturation precisely.

I/O thread activity. io_threaded_reads_processed and io_threaded_writes_processed (Redis 6.0+) confirm whether I/O threading is actually engaging under load. Zero values during high throughput indicate a configuration or workload mismatch.

Client-side caching tracking. tracking_total_keys and tracking_total_items (Redis 6.0+) measure the memory overhead of client-side caching invalidation tables. Large values add hidden memory pressure.

jemalloc statistics. MEMORY MALLOC-STATS exposes arena-level fragmentation, dirty pages, and retained memory. Use this when standard ratios do not explain RSS growth. Look for high dirty and retained pages across arenas, which indicate allocator retention rather than true process fragmentation.

Per-client leak analysis. Track CLIENT LIST over time to identify connections with monotonically increasing idle time or address patterns that never close. Match these to application deploys to find leaking pools.

Big key analysis. Run redis-cli --bigkeys or sample MEMORY USAGE across your keyspace. This uses SCAN and adds read load; run it during low traffic or against a replica. A single large sorted set or hash can dominate latency and memory while aggregate metrics look healthy. For precise sizing, call MEMORY USAGE on specific key patterns.

Key encoding analysis. OBJECT ENCODING samples reveal when Redis transitions compact encodings to generic ones, such as listpack (or ziplist before Redis 7.0) to hashtable. These transitions cause step-function memory growth that aggregate counters smooth over.

How Netdata helps

Netdata auto-discovers Redis instances and collects INFO, SLOWLOG, and LATENCY metrics. Use it to:

  • Correlate used_memory with used_memory_rss to spot fragmentation or swap before the OOM killer fires.
  • Visualize replication offset lag per replica alongside master_link_status to distinguish transient blips from growing divergence.
  • Alert on latest_fork_usec spikes tied to rdb_bgsave_in_progress or aof_rewrite_in_progress to separate persistence latency from runaway commands.
  • Track instantaneous_ops_per_sec against main-thread CPU to identify single-core saturation before client timeouts.
  • Surface stream consumer group lag and cluster slot health without custom scripting.
The Netdata solution

Redis monitoring with Netdata

Netdata monitors Redis with per-second metrics and ML anomaly detection. Track memory usage and fragmentation, fork/COW latency, replication backlog, evictions, and connection pressure to spot the failure modes in these runbooks early.