The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-trap-syslog-flood

Operations Guides

Trap and syslog flood from link flaps: surviving the storm

A single bad SFP starts flapping. Within seconds, your trap receiver is processing hundreds of linkDown/linkUp pairs per second, your syslog pipeline is drowning in LINK-3-UPDOWN messages, and STP topology change notifications are cascading across the L2 domain. The kernel socket buffer on UDP 162 overflows, and the root-cause hardware alarm is as likely to be dropped as any other datagram in the flood.

The core problem is architectural. Traps and syslog arrive over UDP, a lossy transport with no retransmission. When a receiver socket buffer fills, the kernel silently discards datagrams. There is no per-source accounting. You cannot tell which device’s traps were dropped, only that some were. During a flood, any datagram arriving during the overflow window can vanish, including the one that matters most.

What this means

A link-flap cascade follows a predictable path. A physical-layer fault (bad cable, dirty fiber, failing transceiver, end-station NIC bug, power fluctuation) causes an interface to transition up and down rapidly. Each transition generates a linkDown trap and, moments later, a linkUp trap. On switching platforms, each transition also fires STP topology change notifications as the bridge re-evaluates the forwarding topology. The syslog stream fills with interface state change messages.

A single flapping access port can generate dozens of trap pairs per second. A cascade across multiple ports (STP instability, power event hitting a whole rack) multiplies this by the number of affected interfaces. The trap receiver on UDP 162 and the syslog receiver on UDP 514 share the same kernel UDP infrastructure. When the burst exceeds the socket buffer drain rate, datagrams are silently discarded.

There is no priority queueing for incoming UDP traps. Every datagram that arrives while the socket buffer is full is equally likely to be dropped, regardless of severity. Recovery confirmations, unrelated critical alarms, and the root-cause event itself all face the same odds during the overflow window.

flowchart TD
    A["Bad SFP / dirty fiber / NIC bug"] --> B["Interface flaps rapidly"]
    B --> C["linkDown/linkUp trap pairs"]
    B --> D["LINK-3-UPDOWN syslog messages"]
    B --> E["STP topology change notifications"]
    C --> F["Trap receiver UDP 162 flooded"]
    D --> G["Syslog receiver UDP 514 flooded"]
    F --> H["Kernel socket buffer overflows"]
    G --> H
    H --> I["Root-cause trap silently dropped"]
    I --> J["Operators see noise, miss signal"]

Common causes

CauseWhat it looks likeFirst thing to check
Bad cable, SFP, or dirty fiberSingle interface flapping with incrementing ifInErrorsifInErrors rate and physical-layer syslog
End-station NIC bug or floodingAccess port flapping, no physical errors on the switch sideMAC behavior on the port, end-station driver logs
Power fluctuation to end deviceIntermittent link-loss on one port, often time-correlated with other devices on same PDUPower infrastructure, UPS logs
Speed or duplex mismatchVery high error rates on the interface, carrier transitionsInterface speed/duplex negotiation on both ends
STP instability causing reconvergenceMultiple interfaces transitioning, TCN count rising rapidlydot1dStpTopChanges and root bridge identity
LACP timer mismatchLACP bundle member flapping, especially under CPU loadLACP periodic timer setting (SLOW vs FAST)

Quick checks

# Check trap receive rate. Format depends on your snmptrapd logging config;
# adjust the awk field index to match your delimiter/layout.
awk -F'|' '{print $4}' /var/log/snmptrapd.log | sort | uniq -c | sort -rn | head

# Check syslog rate for LINK-3-UPDOWN flood
grep -c 'LINK-3-UPDOWN' /var/log/network-devices/*.log

# Check kernel UDP socket buffer drops (system-wide, all UDP ports)
cat /proc/net/snmp | grep '^Udp:'
nstat -az Udp_RcvbufErrors

# Check current trap receiver socket fill level
ss -lun '( sport = :162 )' -m

# Capture live trap traffic to see what is arriving
tcpdump -i eth0 -nn 'udp port 162' -c 100

# Check syslog receiver socket fill level
ss -lun '( sport = :514 )' -m

# Identify the flapping interface via SNMP (replace community and host)
snmpwalk -v2c -c <community> <device> .1.3.6.1.2.1.2.2.1.8   # ifOperStatus

# Check input errors on the suspected interface
snmpwalk -v2c -c <community> <device> .1.3.6.1.2.1.2.2.1.14  # ifInErrors

# Check STP topology change count
snmpwalk -v2c -c <community> <device> .1.3.6.1.2.1.17.2.4    # dot1dStpTopChanges

How to diagnose it

  1. Identify the flapping interface. Parse the trap log for interface index values in linkDown/linkUp varbinds. The IF-MIB interface index maps to a physical port via ifDescr or ifName. If traps are being dropped (check Udp_RcvbufErrors), fall back to polling ifOperStatus directly.

  2. Confirm it is a flap, not a single transition. A linkDown/linkUp pair frequency above 1 per second on any interface is a flap. More than 5 transitions per minute warrants investigation. Cross-reference with ifInErrors to distinguish physical-layer faults from logical issues.

  3. Check for STP impact. If dot1dStpTopChanges is incrementing rapidly in correlation with the flap, the L2 domain is reconverging. This widens the blast radius beyond the single port.

  4. Verify receiver health. Check Udp_RcvbufErrors on the trap and syslog collector. Any nonzero increment during the event means datagrams were dropped. There is no per-source attribution. If the counter advanced, assume critical traps may be missing.

  5. Check for correlated events that may have been lost. If the flap coincided with a BGP session drop, a hardware alarm, or a config change, those traps or syslog messages may have been dropped in the flood. Poll the relevant OIDs directly to reconstruct state.

  6. Patch snmptrapd. A buffer overflow vulnerability in net-snmp’s snmptrapd (CVE-2025-68615) could allow a remote attacker to crash the daemon via crafted trap packets. A trap flood is an ideal delivery vector. If your snmptrapd is unpatched, the flood may take down the receiver entirely rather than just dropping packets.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Trap receive rateSudden spike indicates an event cascadeRate 10x or more above rolling baseline
linkDown/linkUp pair frequencyIdentifies the flapping interfaceMore than 1 pair per second on any interface
Udp_RcvbufErrorsSilent datagram loss at the kernel socket bufferAny nonzero increment during a flood
Syslog receive rate and severityParallel flood on UDP 514Rate spike without severity escalation means noise, not a real event
ifOperStatus transitionsConfirms the flap via polling when traps are unreliableMore than 3 transitions per minute
ifInErrorsPhysical-layer root cause indicatorAny sustained increment
STP topology change count (dot1dStpTopChanges)Measures L2 blast radiusRapid increment correlated with the flap
Device control-plane CPUTrap generation and processing consumes device CPUSustained above 70% during the event
Trap source diversityA single sender dominating trap volume is a findingOne device producing more than 50% of all traps

Fixes

Stop the flood at the source

The fastest mitigation is to administratively shut down the flapping interface. This stops trap generation immediately and lets the receiver drain its backlog.

# Identify flapping interfaces (exec mode)
ssh <device> 'show interfaces | include protocol.*down|reset|err-disable'
! Disruptive: shuts down the interface. Required config-mode commands:
device# configure terminal
device(config)# interface <if>
device(config-if)# shutdown
device(config-if)# end

This is disruptive to whatever is connected to that port, but it preserves visibility for the rest of the network. If the port is an access port serving a single end station, the tradeoff is almost always correct.

On Cisco IOS and IOS XE, an interface that flaps more than 5 times within 10 seconds enters errdisable state. Both a syslog message and an SNMP trap are sent upon port shutdown. Enable automatic recovery so the port does not require manual intervention:

device# configure terminal
device(config)# errdisable recovery cause link-flap
device(config)# errdisable recovery interval 300
device(config)# end
# View current flap thresholds and recovery config (exec mode)
ssh <device> 'show errdisable flap-values'
ssh <device> 'show errdisable recovery'

The default 300-second recovery interval gives the physical issue time to stabilize. If the port flaps again after recovery, it re-enters errdisable. This is a self-healing mechanism that prevents sustained floods.

Arista EOS detects continuously flapping interfaces and temporarily holds the link-down state to smooth out flap churn. The feature was introduced in 2024; check the Arista support documentation for your EOS version to confirm availability and whether it applies to your access ports.

Set LACP periodic timer to SLOW (Juniper EX)

On Juniper EX2300 and EX3400 series, LACP bundles can flap during CPU-intensive events such as routing engine switchover or interface flaps elsewhere on the device. Setting the LACP periodic timer to SLOW reduces CPU overhead and prevents spurious member-link flaps that trigger trap floods.

Rate-limit syslog and traps on the device

Cisco IOS provides a global syslog rate-limit that caps messages per second.

device# configure terminal
device(config)# logging rate-limit 10
device(config)# end

For per-message rate limiting, use a logging discriminator to suppress or throttle specific message types:

device(config)# logging discriminator LINKFLAP rate-limit 5
device(config)# logging discriminator LINKFLAP msg-body contains "LINK-3-UPDOWN"
! Apply the discriminator to a logging destination:
device(config)# logging host <host> discriminator LINKFLAP

For trap suppression on specific interfaces, use per-interface link-status trap control. The interface-level command is no snmp trap link-status on most IOS and IOS XE trains. The global form is also available:

! Suppress linkup/linkdown traps globally
device(config)# no snmp-server enable traps snmp linkup linkdown

Suppressing linkup/linkdown traps globally is aggressive. It eliminates the signal entirely. A better approach for access ports is per-interface suppression while keeping traps enabled on uplinks and critical infrastructure ports.

Harden the receiver

Tune the UDP socket buffer on the collector to absorb bursts.

# Increase UDP socket buffer limits (requires root)
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.rmem_default=16777216

# Verify the trap listener is using the larger buffer
ss -lun '( sport = :162 )' -m

Apply these persistently via sysctl configuration (for example, /etc/sysctl.d/). If the collector process sets SO_RCVBUF explicitly, it must request a value at or below rmem_max. If it does not set SO_RCVBUF, it inherits rmem_default, which is why both sysctl keys matter.

Patch snmptrapd against known buffer overflow vulnerabilities. During a link-flap flood, a malicious trap crafted to exploit a vulnerability would be indistinguishable from flood noise.

Prevention

Enable errdisable link-flap detection globally on all Cisco switching platforms. This converts a sustained flap into a single shutdown event with one trap pair instead of hundreds.

Configure syslog rate-limiting on all devices. Prevents the syslog pipeline from being overwhelmed during any event burst.

Suppress linkup/linkdown traps on access ports where the physical state is not operationally critical. Keep traps enabled on uplinks, peer-facing interfaces, and any port in the critical path.

Monitor Udp_RcvbufErrors on trap and syslog collectors continuously. Any nonzero increment means data was lost. This is the single most under-monitored signal in trap and syslog collection.

Watch trap source diversity. If a single device is producing more than 50% of all traps over any sustained window, that device is likely in distress. Alert on this ratio before the flood reaches the receiver.

Be aware of platform-specific blind spots. On Cisco IOS XE, console trap display may be suppressed by interval-based message limiting during a storm, meaning an operator watching the console may miss the flood entirely. After a soft OIR (online insertion/removal) of an interface module, there is approximately a 1-minute delay before storm control detects a storm condition, leaving a window where floods from the flap event itself go unaccompanied by storm control notifications.

Ensure NTP synchronization on all devices. During a flood, postmortem correlation depends on accurate timestamps. Two devices 200ms apart on the same flap will produce trap and syslog records that fail to align in reconstruction.

How Netdata helps

Netdata correlates the signals that matter during a trap flood, providing cross-layer visibility that raw trap logs cannot:

  • Trap receive rate anomaly detection surfaces the spike within seconds, before the receiver buffer overflows.
  • Udp_RcvbufErrors monitoring on the collector catches silent datagram loss that would otherwise go unnoticed until a postmortem fails to reconstruct the event.
  • Syslog ingestion rate and severity distribution distinguish a noise flood (rate spike without severity escalation) from a real event cascade (rate spike with severity escalation).
  • ifOperStatus transition tracking via SNMP polling confirms the flapping interface even when traps are being dropped, providing a ground-truth fallback.
  • Device control-plane CPU monitoring during the event identifies devices under stress from trap generation and STP recalculation.
  • STP topology change correlation with interface state transitions reveals the L2 blast radius of the flap in real time.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.