The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-snmp-auth-failure-spikes

Operations Guides

SNMP authentication-failure spikes: misconfiguration vs reconnaissance

SNMP authentication-failure traps are one of the few security signals built into the network monitoring stack. When they spike, the question is never “is something wrong?” - it is “is this a broken poller or someone probing my devices?” The answer changes the response from a quiet config fix to a security incident.

The authenticationFailure trap (OID 1.3.6.1.6.3.1.1.5.5) fires whenever an SNMP agent receives a protocol message that is not properly authenticated. On SNMPv2c, that means a wrong community string. On SNMPv3, it means a wrong username, wrong auth protocol, wrong auth password, or wrong privacy password. The trap is defined in SNMPv2-MIB and every compliant agent can generate it, but many vendors ship with it disabled by default. If you have never explicitly enabled it (for example, snmp-server enable traps snmp authentication on Cisco IOS), you may have no signal at all.

The hard part is interpretation. A burst of auth failures from a single IP can be a monitoring station with a stale credential, or an attacker enumerating community strings. A burst across many devices from one source is almost always scanning. The discriminating signals are source-IP distribution, timing regularity, credential variety, and whether the source is inside or outside the management subnet.

Signal types: traps vs cumulative counters

The authenticationFailure trap is event-based: it fires on each failed authentication attempt. This is distinct from the cumulative counters that track the same failure at the agent level.

For SNMPv2c, the relevant counters are:

  • .1.3.6.1.2.1.11.4 (snmpInBadCommunityNames) - community name not recognized by the agent
  • .1.3.6.1.2.1.11.5 (snmpInBadCommunityUses) - community recognized but the SNMP operation was not permitted for that community

For SNMPv3, the USM (User-based Security Model) exposes two separate counters:

  • .1.3.6.1.6.3.15.1.1.3 (usmStatsUnknownUserNames) - the username was not recognized by the agent
  • .1.3.6.1.6.3.15.1.1.5 (usmStatsWrongDigests) - the username was recognized but the authentication digest did not match

This distinction matters operationally. Rising usmStatsUnknownUserNames with stable usmStatsWrongDigests means someone is probing usernames that do not exist on the device. Rising usmStatsWrongDigests with stable usmStatsUnknownUserNames means the username is valid but the password or auth protocol is wrong, which points to credential rotation or key compromise rather than blind probing.

Common causes

CauseWhat it looks likeFirst thing to check
Misconfigured pollerFailures from one known management IP, periodic at the polling interval (often 300s), same credential each timeVerify the poller’s SNMP credential configuration against the device
Credential rotationFailures from one or a few known IPs, starting at a specific timestamp, correct username but wrong digestCheck whether a credential rotation was applied to one side but not the other
Vulnerability scanner sweepFailures across many devices from one source IP, varied credentials attempted, non-periodic timingIdentify the scanner source and verify it is authorized
External reconnaissanceFailures from IPs outside the management subnet, rapid succession, varied community strings or usernamesCheck perimeter ACL and firewall logs for the source IP
Community string “public” still configuredNo auth-failure traps at all from scanned devices, but snmpInBadCommunityNames is high because scans succeed against “public”Poll snmpInBadCommunityNames and check device config for default community strings

Quick checks

# Check SNMPv2c bad community name counter
snmpget -v2c -c <community> <device> .1.3.6.1.2.1.11.4.0

# Check SNMPv3 USM unknown usernames counter
# Match -a (SHA/SHA-256/MD5) to the agent's configured auth protocol
snmpget -v3 -l authNoPriv -u <user> -a SHA -A <authpass> <device> .1.3.6.1.6.3.15.1.1.3.0

# Check SNMPv3 USM wrong digests counter
snmpget -v3 -l authNoPriv -u <user> -a SHA -A <authpass> <device> .1.3.6.1.6.3.15.1.1.5.0

# Search trap receiver logs for auth failures
grep "authenticationFailure" /var/log/snmptrapd.log | tail -50

# Count auth failure traps from the device syslog (Cisco IOS/IOS XE)
ssh <device> 'show logging | include SNMP-3-AUTHFAIL'

# Verify the auth-failure trap is enabled on the device (Cisco IOS)
ssh <device> 'show run | include snmp-server enable traps snmp authentication'

How to diagnose it

flowchart TD
    A["authFailure spike"] --> B{"Single source IP?"}
    B -- Yes, known poller --> C["Misconfigured poller or credential rotation"]
    B -- Yes, unknown --> D{"Inside mgmt subnet?"}
    D -- Yes --> E["Unauthorized internal tool"]
    D -- No --> F["External scanning: escalate to security"]
    B -- No, many sources --> G{"Varied credentials?"}
    G -- Yes --> H["Reconnaissance or vuln scanner"]
    G -- No --> I["Shared wrong credential across estate"]
  1. Extract source IPs from the trap stream and syslog. The standard authenticationFailure trap carries no varbinds beyond sysUpTime and snmpTrapOID. The source IP of the failed request is more reliably available in device syslog. On Cisco IOS, the SNMP-3-AUTHFAIL syslog message includes the source IP directly. Parse both the trap receiver log and the device syslog to build a source-IP frequency table.

  2. Classify each source IP. For each source, determine: is it a known monitoring station? Is it inside the management subnet? Any nonzero auth-failure rate from a source outside the management subnet is an unauthorized access attempt. Escalate immediately.

  3. Check timing regularity. Misconfigured pollers produce failures at their polling interval, typically every 300 seconds. If failures arrive at a precise, repeating interval, the source is almost certainly a monitoring probe with a stale credential. Random timing or rapid bursts suggest active scanning.

  4. Distinguish USM error types for SNMPv3. Poll usmStatsUnknownUserNames (.1.3.6.1.6.3.15.1.1.3.0) and usmStatsWrongDigests (.1.3.6.1.6.3.15.1.1.5.0) separately. Rising unknown-usernames with stable wrong-digests indicates username probing. The inverse indicates a valid user with a wrong password or auth protocol mismatch, pointing to misconfiguration rather than probing.

  5. Check for multi-vector scanning. Correlate with syslog for SSH and HTTP authentication failures from the same source IP. An attacker probing SNMP is often probing other protocols simultaneously. Check AAA logs (TACACS+/RADIUS) for login failures from the same source.

  6. Verify the absence of traps is not hiding a problem. Many devices still have SNMPv2c community string “public” configured for read-only access. Scans against “public” succeed and therefore do not generate auth-failure traps. Poll snmpInBadCommunityNames even when no traps are seen. If the counter is rising with no corresponding traps, investigate whether “public” or “private” is still configured.

  7. Check for silent v3 failures. Some platforms silently fail SNMPv3 authentication without generating a trap unless auth-failure trapping is explicitly enabled. If you suspect v3 failures but see no traps, verify the trap configuration on the device.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
snmpInBadCommunityNames (.1.3.6.1.2.1.11.4)Cumulative count of community-name mismatches at the agentAny nonzero in production; rising without corresponding traps suggests scans succeeding against “public”
snmpInBadCommunityUses (.1.3.6.1.2.1.11.5)Community recognized but operation not permittedNonzero indicates a tool attempting write operations with read-only credentials
usmStatsUnknownUserNames (.1.3.6.1.6.3.15.1.1.3)SNMPv3 username not recognized by agentRising counter with varied usernames indicates reconnaissance
usmStatsWrongDigests (.1.3.6.1.6.3.15.1.1.5)SNMPv3 username valid but auth digest mismatchRising counter with stable user count indicates credential rotation or key compromise
authenticationFailure trap rate (.1.3.6.1.6.3.1.1.5.5)Event-level signalRate > 10 events/min from a single source is scanning; any event from outside the management subnet is unauthorized
SSH/HTTP auth failure rate (syslog/AAA)Multi-vector scanning probes multiple protocols from one sourceSame source IP failing auth across SNMP, SSH, and HTTP is active reconnaissance
Per-source trap frequency distributionA single sender dominating trap volume is a findingOne source IP accounting for the majority of auth-failure traps

Fixes

Misconfigured poller or stale credential

The most common cause. A monitoring station was recently updated, moved, or reconfigured, and its SNMP credentials no longer match the device.

  1. Identify which poller is generating the failures from the trap source IP or syslog.
  2. Verify the poller’s configured community string or SNMPv3 credentials against the device’s actual configuration.
  3. For SNMPv3, check the auth protocol (MD5 vs SHA), auth password, and privacy password independently. A mismatch on any one produces auth failures. SNMPv3 USM (RFC 3414) requires auth passwords of at least 8 characters; shorter passwords fail silently on compliant agents.
  4. Apply the correct credential to the poller. Do not change the device credential to match the poller unless the poller’s credential is the intended one.

Credential rotation mismatch

Credential rotation applied to one side but not the other.

  1. Verify whether a credential rotation was recently scheduled or executed.
  2. Check both the device and the monitoring system for the current credential.
  3. Synchronize. Prefer rotating on the device first, then updating the poller, to minimize the failure window.

External reconnaissance or scanning

Failures from IPs outside the management subnet require a security response.

  1. Check perimeter firewall and ACL logs for the source IP.
  2. Block the source IP at the perimeter if policy permits.
  3. Verify that SNMP access (UDP 161) is restricted to the management subnet via ACLs on every device. If it is reachable from outside, that is the configuration error that enabled the scanning.
  4. Escalate to the security team. Correlate with SSH, HTTP, or other protocol auth failures from the same source.
  5. If any device still uses community string “public” or “private”, remediate immediately. Scans against “public” succeed silently and never generate auth-failure traps.

SNMPv3 credential special-character issues

On some platforms, shell-special characters ($, backticks, !) in SNMPv3 credentials cause silent authentication failures when passed through CLI or configuration management tools. Test credentials that contain only alphanumerics first to isolate parsing from genuine auth failures.

CVE-2025-20352 exposure

CVE-2025-20352 (CVSS 7.7, disclosed September 2025) is a stack overflow vulnerability in the SNMP subsystem of Cisco IOS and IOS XE. It affects all SNMP versions (v1, v2c, v3) and exploitation requires valid SNMP credentials. An authenticated remote attacker can cause a device reload or execute code as root. If your estate includes Cisco IOS or IOS XE devices and you observe auth-failure spikes from external sources, treat this as a potential precursor to exploitation. Restrict SNMP access to trusted management IPs via ACLs and patch to the fixed release.

Prevention

  • Enable auth-failure traps on every device. Many vendors ship with this disabled. On Cisco IOS, use snmp-server enable traps snmp authentication. Without this, you have no event-level signal.
  • Eliminate default community strings. Any device still using “public” or “private” is a silent finding. Scans against “public” succeed without generating any auth-failure trap, so the absence of traps does not mean the absence of scans.
  • Restrict SNMP to the management subnet. SNMP on UDP 161 should never be reachable from outside the management network. Apply ACLs on every device.
  • Prefer SNMPv3 over v2c. SNMPv2c community strings are transmitted in cleartext and can be captured by passive sniffing on the management VLAN. SNMPv3 with authPriv provides both authentication and encryption.
  • Baseline the auth-failure rate. A healthy estate should have zero auth failures in steady state. Any nonzero value is abnormal. Track the per-source distribution so that a new source is immediately visible.
  • Monitor cumulative counters, not just traps. Traps can be dropped by a flooded receiver. UDP 162 is lossy under burst. The cumulative counters (snmpInBadCommunityNames, usmStatsUnknownUserNames, usmStatsWrongDigests) persist at the agent and survive trap loss.
  • Correlate with SSH and AAA auth failures. Multi-vector scanning probes multiple protocols from the same source. Join SNMP auth-failure events with syslog auth failures by source IP and timestamp.

How Netdata helps

  • Netdata’s SNMP collector can poll snmpInBadCommunityNames, usmStatsUnknownUserNames, and usmStatsWrongDigests at per-second resolution, giving rate-of-change visibility that 5-minute polling misses.
  • Trap receiver metrics expose per-trap-type rates, so an authenticationFailure spike is visible as a distinct signal alongside linkDown, coldStart, and enterprise-specific traps.
  • Correlate SNMP auth-failure spikes with syslog auth failures from the same source IP on the unified timeline, without joining across separate tools.
  • Anomaly detection on the auth-failure counter rate baselines the normal (zero) state and flags any deviation, including slow-rate probing that stays below fixed thresholds.
  • UDP socket buffer drop monitoring (Udp_RcvbufErrors) on the trap receiver surfaces when traps are lost under burst, so an auth-failure spike does not silently disappear at the receiver.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.