The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-bgp-flapping

Operations Guides

BGP flapping: why a peer keeps resetting and how to find the cause

A BGP peer cycling between Established and Idle is sending a specific signal. The session tears down because one side sent a NOTIFICATION message, and that message carries an error code and subcode that pinpoints the cause. Most monitoring watches only the FSM state (up or down) and ignores the NOTIFICATION payload, so the operator sees flapping without knowing why.

The second trap is treating the symptom. Clearing the session or increasing the hold timer does not fix the underlying cause. The session re-establishes briefly, then drops again identically. The hold timer is not the problem; the hold timer is detecting the problem.

What this means

BGP flapping is repeated transitions out of Established. Each transition is triggered by one of three events: the local router sends a NOTIFICATION (error detected locally), the remote router sends a NOTIFICATION (error detected remotely), or the TCP session fails (transport loss). The NOTIFICATION message is the single most important diagnostic artifact, and its error code determines the entire investigative path.

The BGP finite state machine progresses: Idle, Active (TCP attempts), OpenSent, OpenConfirm, Established. A peer stuck in Active means TCP connectivity is failing (neighbor unreachable, firewall blocking port 179, or peer down). A peer that reaches Established and then drops back to Idle means the session was established and then torn down by a NOTIFICATION or transport failure.

flowchart TD
    A[BGP peer flapping] --> B{Read NOTIFICATION code}
    B -->|4 Hold Timer Expired| C{CPU high on device?}
    C -->|Yes| D[Control-plane saturation]
    C -->|No| E[MTU black hole or packet loss]
    B -->|6 Cease subcode| F[1 Max Prefixes
2 Admin Shutdown
5 Conn Rejected
10 BFD Down] B -->|2 Open Error| G[Wrong AS or MD5 mismatch] B -->|3 Update Error| H[Malformed route from peer] B -->|No NOTIFICATION| I[TCP or interface failure]

Key BGP NOTIFICATION error codes:

  • Code 2 (Open Message Error): parameter mismatch during session setup. Most commonly wrong remote AS or MD5 authentication failure.
  • Code 3 (Update Message Error): malformed UPDATE from the peer. The peer sent a route the local router could not parse.
  • Code 4 (Hold Timer Expired): the local router did not receive any BGP message (keepalive, UPDATE, or NOTIFICATION) within the negotiated hold time. Hold time defaults vary by vendor; Cisco IOS uses 180/60, Junos uses 90/30. The hold timer expiring means the router heard silence from the peer for the entire hold window.
  • Code 6 (Cease): the session was administratively or policy torn down. The subcode further specifies the reason.

Cease subcodes per RFC 4486 and extensions:

SubcodeMeaningRFC
1Maximum Prefixes Reached4486
2Administrative Shutdown4486
3Peer De-configured4486
4Administrative Reset4486
5Connection Rejected4486
6Other Configuration Change4486
7Connection Collision Resolution4486
8Out of Resources4486
9Hard Reset8538
10BFD Down9384

Common causes

CauseWhat it looks likeFirst thing to check
MTU/PMTUD black holeSession establishes, small keepalives pass, but hold timer expires when UPDATE messages flow. Large BGP UPDATEs are silently dropped.Extended ping with DF bit set toward peer IP at 1400+ bytes
Control-plane CPU saturationHold timer expiry correlates with CPU spikes. Often during massive route table updates or SNMP polling storms.Device CPU counters and process list
Maximum prefix limit (Cease/1)Peer placed in Idle (PfxCt) state after exceeding configured ceiling. Session does not auto-recover.Per-peer prefix count vs configured maximum
MD5 authentication mismatchSession fails to establish or fails right after establishing. Open Message Error or silent TCP failure with %TCP-6-BADAUTH in logs.Router log for auth messages
Interface flap with fast-external-falloverEach BGP reset coincides with an interface transition. Cisco IOS enables fast-external-fallover by default for eBGP.Interface operational status around reset time
BFD-triggered reset (Cease/10)BFD detects path loss faster than BGP hold timer. Reset is immediate after BFD session goes down.BFD session state and underlay path quality
Peer maintenance (Cease/2)Expected reset during a scheduled window. RFC 8203 allows a free-form shutdown message.Change ticket and peer notification

Quick checks

# Check BGP summary for all peers - state and prefix counts
ssh <router> 'show ip bgp summary'

# Get the exact NOTIFICATION reason for a specific peer
ssh <router> 'show ip bgp neighbors <peer> | include notification|Last'

# Check BGP events in syslog
ssh <router> 'show logging | include BGP'

# Poll bgpPeerState via SNMP (RFC 1657 / RFC 4273)
snmpwalk -v2c -c <community> <router> .1.3.6.1.2.1.15.3.1.2

# Check received UPDATE rate per peer via SNMP
snmpwalk -v2c -c <community> <router> .1.3.6.1.2.1.15.3.1.10

# Check control-plane CPU on Cisco
snmpget -v2c -c <community> <device> .1.3.6.1.4.1.9.9.109.1.1.1.1.7
ssh <device> 'show processes cpu sorted | include five sec'

# Check control-plane CPU on Juniper
snmpwalk -v2c -c <community> <device> .1.3.6.1.4.1.2636.3.1.13.1.8

# Check FRR (Linux) BGP state
vtysh -c 'show bgp summary'

How to diagnose it

  1. Decode the NOTIFICATION. Run show ip bgp neighbors <peer> | include notification|Last (Cisco) or show bgp neighbor <peer> (Juniper). The output shows the last reset reason with the error code and subcode. This single line determines the entire investigative path.

  2. For hold timer expiry (code 4), check two things. First, check control-plane CPU on the device. If CPU was above 90% when the session dropped, the BGP process was starved and could not process incoming keepalives in time. Correlate the reset timestamp with CPU history. Second, if CPU was normal, test for an MTU black hole. Run an extended ping with the DF bit set toward the peer IP, stepping up payload sizes from 1400 bytes. If large packets fail while small ones succeed, a middlebox is dropping packets above a certain size and blocking the ICMP “fragmentation needed” response.

  3. For Cease/1 (Maximum Prefixes), check the peer’s prefix count. Compare the current count against the configured maximum-prefix limit. If the peer is sending far more prefixes than usual, investigate a route leak. See BGP route leak and hijack for detection signals.

  4. For Cease/2 (Administrative Shutdown) or Cease/4 (Administrative Reset), verify intent. Check for a change ticket and whether the peer sent an RFC 8203 shutdown message. Modern IOS/XE and Junos releases support this; older releases send the subcode without the human-readable text.

  5. For Cease/10 (BFD Down), check the BFD session state and the underlay path. BFD detects forwarding path failures faster than BGP hold timers. If BFD is triggering, the path is genuinely degraded; the fix is to address the underlay, not to weaken BFD.

  6. For Open Message Error (code 2), verify session parameters. Check the remote AS number and MD5 authentication key on both sides. MD5 mismatches are insidious because the failure happens at the TCP layer before BGP sees the message, so the BGP process may not emit a NOTIFICATION at all. Check the router log for %TCP-6-BADAUTH messages.

  7. For no NOTIFICATION, check the transport layer. Run show tcp brief or equivalent to verify the TCP session. Check the peer-facing interface operational status. If the interface is flapping and the device runs Cisco IOS with bgp fast-external-fallover enabled (the default for eBGP), any interface transition to an eBGP peer immediately resets the BGP session.

  8. Check for a stale Established. Sometimes the FSM reports Established but UPDATE exchange has stopped. Track bgpPeerInUpdates (OID .1.3.6.1.2.1.15.3.1.10) and the timestamp of the last received prefix. If the rate is flat or zero while the session shows Established, a middlebox may be silently dropping UPDATE messages or the peer’s CPU may be too saturated to generate them.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
BGP session state (bgpPeerState)Primary liveness indicator for the peeringRepeated Established-to-Idle transitions; peer in Active for extended time
BGP NOTIFICATION code and subcodePinpoints the exact cause of each resetAny unexpected NOTIFICATION from a critical peer
bgpPeerInUpdates rateDetects stale Established (session up but no data flowing)Rate drops to zero while FSM shows Established
Per-peer prefix countDetects route leaks and maximum-prefix triggersSudden increase (more than 20% in 5 min) or decrease from a peer
Control-plane CPUHigh CPU causes hold timer expiry by starving the BGP processSustained above 90% correlating with session drops
Interface operational statusInterface flap triggers BGP reset if fast-external-fallover is enabledTransitions on the peer-facing interface around reset time
ICMP reachability to peerBasic transport health independent of BGPLoss or latency spike preceding session drop
BFD session stateFaster path-loss detection than BGP hold timersBFD session flapping or down

Fixes

MTU/PMTUD black hole

Configure TCP MSS clamping to prevent BGP UPDATE messages from exceeding the path MTU. On Cisco IOS, use ip tcp adjust-mss <bytes> under the peer-facing interface, or ip tcp mss global <bytes> to clamp all BGP sessions system-wide. On Juniper Junos, use set system tcp-mss. The correct value depends on the smallest MTU along the path; for VPN tunnels, subtract the encapsulation overhead from the physical MTU.

Verify the fix by running the extended ping with DF bit again. Large packets should now succeed.

Control-plane CPU saturation

Identify the process consuming CPU. Common culprits during BGP flapping: massive route table convergence, SNMP polling storms from a collector over-walking the device, or a trap flood from a downstream link-flap cascade. Address the root cause of the CPU spike. If the cause is SNMP polling pressure from your monitoring system, reduce poller concurrency or exclude large MIB table walks from default polling.

Do not increase the BGP hold timer to mask CPU starvation. It delays detection without fixing the underlying problem, and it slows convergence when genuine failures occur.

Maximum prefix limit (Cease/1)

The session does not automatically recover after hitting the maximum-prefix ceiling. An explicit clear is required: clear ip bgp <peer>. This is disruptive; only run it after confirming the peer should be sending that many prefixes.

Investigate why the peer is sending more prefixes than expected. If the count increased suddenly, suspect a route leak from the peer’s upstream. If the limit was set without headroom for organic growth, raise the ceiling. Configure a warning threshold (typically 75-90% of the limit) so you are alerted before the session is torn down.

MD5 authentication mismatch

Verify the MD5 key on both peers. Keys must match exactly. After correcting, the session should re-establish within seconds. Check the router log for %TCP-6-BADAUTH messages to confirm the mismatch was the cause.

Interface flap with fast-external-fallover

If the underlying link to an eBGP peer is unstable and each interface transition causes a BGP reset, consider disabling bgp fast-external-fallover with no bgp fast-external-fallover. This decouples BGP session state from interface state, allowing the BGP hold timer to ride through brief interface flaps. The tradeoff: detection of a genuine link failure is slower (hold timer must expire instead of immediate reset).

BFD-triggered reset (Cease/10)

BFD is doing its job: detecting path degradation faster than BGP would on its own. The fix is to address the underlay path quality, not to remove BFD. If the BFD timers are too aggressive for the path characteristics (lossy wireless, satellite, or VPN underlay), tune the timers rather than disabling BFD entirely.

Prevention

  • Monitor NOTIFICATION subcodes, not just FSM state. A dashboard that shows only “Established” or “Idle” misses the cause. Parse the error code and subcode from syslog, traps, or CLI output and alert on specific subcodes.
  • Track bgpPeerInUpdates rate and last-receive timestamp. This catches stale Established sessions where the FSM reports up but no data is flowing.
  • Proactively clamp TCP MSS on BGP sessions over tunnels. Encapsulation overhead shrinks effective MTU. Clamping prevents the most common silent cause of hold timer expiry.
  • Monitor control-plane CPU trends. CPU above 70% sustained is a leading indicator of hold timer expiry under load.
  • Configure maximum-prefix limits with warning thresholds. Set the hard limit with headroom and the warning at 75-90% so you can act before the session is torn down.
  • Verify MD5 keys as part of change management. Key rotation without coordinated updates on both peers is a common cause of unexpected session drops.
  • Consider BMP (RFC 7854) for deeper visibility. BMP provides Adj-RIB-In (pre-policy and post-policy routes), per-prefix real-time streaming, and withdrawal tracking that BGP4-MIB cannot. BGP4-MIB’s bgp4PathAttrTable only reflects best-path routes.

How Netdata helps

  • Correlates BGP session state transitions with control-plane CPU spikes on the same timeline, making it immediately clear whether hold timer expiry is CPU-induced or path-induced.
  • Parses BGP NOTIFICATION error codes and Cease subcodes from SNMP traps (bgpBackwardTransition, RFC 4273) and syslog, so you see the exact reset reason without manual CLI inspection.
  • Tracks bgpPeerInUpdates rate alongside FSM state to detect stale Established sessions that look healthy but carry no data.
  • Correlates interface operational status with BGP session drops, making fast-external-fallover and link-flap cascades visible in a single view.
  • Monitors per-peer prefix count trends, alerting on sudden increases that may indicate a route leak or an approaching maximum-prefix limit.
  • Integrates ICMP reachability probes to the peer alongside BGP-layer state, so transport-layer loss is visible in context.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.