The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-bgp-notification-cease

Operations Guides

BGP NOTIFICATION and Cease messages: what each subcode is telling you

A BGP NOTIFICATION in your router log is a peer telling you why it tore down the session. The message carries an error code and an error subcode. Those two numbers tell you whether you are looking at a maintenance window, a route leak, a prefix-limit hit, a CPU-starved control plane, or a BFD-triggered teardown.

Cease (code 6) is the most common NOTIFICATION. Its subcodes, defined in RFC 4486 and extended by RFC 8538 and RFC 9384, hold most of the diagnostic value. Codes 2 through 5 appear less often but point to distinct failure classes: parameter mismatch, malformed updates, hold-timer expiry, and FSM errors.

NOTIFICATION message structure

Every BGP NOTIFICATION message has three fields:

  • Error Code (1 byte): the broad category of error.
  • Error Subcode (1 byte): a more specific reason within that category.
  • Data (variable length): diagnostic payload whose format depends on the error code.

Router logs render this as “code/subcode” with a human-readable description. For example, %BGP-3-NOTIFICATION: received from neighbor x.x.x.x active 6/1 (cease: max-prefixes reached) means error code 6 (Cease), subcode 1 (Maximum Number of Prefixes Reached). The subcode is where you look first.

Error codes at a glance

The IANA BGP Parameters registry defines nine error codes (excluding reserved). Most production incidents involve codes 2 through 6.

CodeNameWhat it means operationallyRFC
1Message Header ErrorConnection-level framing problem. Rare in modern implementations.4271
2OPEN Message ErrorBGP parameters mismatched during session setup. Common after config changes.4271
3UPDATE Message ErrorPeer sent a malformed or policy-violating route. Often a bad AS-path or invalid attribute.4271
4Hold Timer ExpiredPeer did not receive keepalives within the negotiated hold time. Frequently indicates control-plane CPU saturation.4271
5Finite State Machine ErrorPeer received an unexpected message for its current FSM state. Usually a software bug or race condition.4271
6CeasePeer intentionally terminated the session. Subcode carries the reason. Most common code in production.4271
7ROUTE-REFRESH Message ErrorMalformed route-refresh request. Rare unless route-refresh is heavily used.7313
8Send Hold Timer ExpiredLocal speaker failed to send within the hold interval. Distinct from code 4 (receive side).9687
9Loss of LSDB SynchronizationBGP-LS deployments only. Not applicable to conventional BGP peering.9815

Codes 8 and 9 are recent additions. Code 8 distinguishes a send-side timeout from the classic receive-side Hold Timer Expired (code 4). Code 9 applies to BGP-LS.

Cease subcodes decoded

Cease (code 6) carries a subcode that tells you why the peer tore down the session. The original eight subcodes come from RFC 4486. Subcode 9 (Hard Reset) was added by RFC 8538, and subcode 10 (BFD Down) was added by RFC 9384.

SubcodeNameRFCWhat triggered it
1Maximum Number of Prefixes Reached4486Peer exceeded the configured prefix limit. Could be a route leak or organic growth.
2Administrative Shutdown4486, 8203Peer intentionally shut down the session, typically for maintenance.
3Peer De-configured4486Peer removed your configuration on their end.
4Administrative Reset4486, 8203Peer reset the session, usually after a policy change.
5Connection Rejected4486Peer refused the TCP connection. Often a policy or peer-group mismatch.
6Other Configuration Change4486Peer changed policy that does not fit subcodes 2 through 5.
7Connection Collision Resolution4486Two simultaneous connection attempts resolved by closing one. Benign.
8Out of Resources4486Peer ran out of memory or other resources.
9Hard Reset8538Peer demands a full session reset, defeating Graceful Restart.
10BFD Down9384Associated BFD session went down, triggering BGP teardown.

Subcode 0 (Reserved) appears when no specific subcode applies. RFC 4271 defines it as “Unspecific.” Some vendors log it as-is; others substitute a generic description.

flowchart TD
    A["NOTIFICATION received"] --> B{"Error code?"}
    B -->|"6 Cease"| C["Decode subcode"]
    B -->|"2 OPEN"| D["AS, MD5, capability mismatch"]
    B -->|"3 UPDATE"| E["Malformed path attribute"]
    B -->|"4 Hold Timer"| F["Control-plane CPU saturation"]
    C --> G{"Cease subcode?"}
    G -->|"1 Max Prefix"| H["Route leak or growth"]
    G -->|"2/4 Admin"| I["Maintenance or policy"]
    G -->|"9/10"| J["Hard reset or BFD down"]

What each Cease subcode means in practice

Subcode 1: Maximum Number of Prefixes Reached. A peer is sending more prefixes than your configured maximum-prefix limit allows. Two scenarios: organic growth that exceeded a stale limit, or a route leak where the peer is advertising prefixes they should not. The session is torn down and all routes from that peer are withdrawn. Check per-peer prefix-count trends to determine whether this was gradual (growth) or sudden (leak). If you see RPKI-invalid routes from the same peer around the same time, treat it as a potential route leak.

Subcode 2: Administrative Shutdown. The peer intentionally brought the session down, typically for planned maintenance. RFC 8203 adds an optional UTF-8 shutdown communication string (up to 128 octets) that the receiving implementation must log via syslog. If your peer supports RFC 8203, the log line will contain a freeform reason such as a ticket number. Check your change management system. If there is no change ticket, this could be an emergency shutdown on the peer side.

Subcode 3: Peer De-configured. The peer removed your BGP configuration entirely. This is not a transient event. Someone on the peer side deleted or commented out your neighbor statement. Contact the peer’s NOC.

Subcode 4: Administrative Reset. Similar to subcode 2 but typically triggered by a policy change rather than a full shutdown. The peer applied a new route-map, changed import or export policy, or reloaded BGP configuration. RFC 8203 shutdown communication also applies to this subcode. If this happens outside a maintenance window, investigate whether the peer changed filtering policy that affects your routes.

Subcode 5: Connection Rejected. The peer refused the incoming TCP connection. On Juniper devices, this often appears with a log line like “no group for IP+port from AS X found”, meaning the connection arrived from an address not belonging to a configured peer group. This is a configuration ordering issue on the peer side, not a protocol error. Verify that the peer’s configuration references the correct source IP for your router.

Subcode 6: Other Configuration Change. A catch-all for policy changes that do not fit subcodes 2 through 5. Less specific, but still indicates a deliberate change on the peer side. Check with the peer’s NOC.

Subcode 7: Connection Collision Resolution. Both sides initiated TCP connections simultaneously, and BGP collision detection resolved the duplicate by keeping one and closing the other. This is normal protocol behavior. No action needed unless it recurs frequently, which may indicate a timer or topology issue causing both sides to reconnect at the same time.

Subcode 8: Out of Resources. The peer ran out of memory, TCAM, or another finite resource. This is a peer-side capacity problem. Correlate with the peer’s control-plane CPU and memory if you have visibility. If this recurs, the peer may need a hardware upgrade or RIB optimization.

Subcode 9: Hard Reset (RFC 8538). The peer demands a full session reset with no Graceful Restart assistance. The triggering NOTIFICATION is encapsulated in the data portion of the Hard Reset message. Upon receipt, the receiving speaker must flush all routes from that peer and perform a complete session reset. This explicitly defeats any Graceful Restart helper behavior. FRRouting had a bug (issue #21822) where the GR helper incorrectly retained stale routes on receipt of Cease(6) or Hard Reset(9). If you run FRR, verify your version includes the fix.

Subcode 10: BFD Down (RFC 9384). The peer tore down the BGP session because the associated BFD session went Down. RFC 9384 makes this a SHOULD, not a MUST. Some implementations send generic Cease without subcode 10 when BFD brings down the session. The underlying BFD failure is the real event. Investigate BFD session state on both endpoints. BFD down typically means path loss or excessive latency and jitter that exceeded BFD thresholds.

Beyond Cease: error codes you will see in production

Code 2: OPEN Message Error. Session setup failed during parameter negotiation. Common causes: wrong remote AS number, MD5 authentication key mismatch, unsupported capabilities, or hold-time disagreement. These almost always indicate a configuration error on one side. If this appears after a key rotation or config change, check the AS number, MD5 key, and address-family or capability negotiation settings. RFC 9234 (2022) deprecated OPEN subcodes 8, 9, and 10 and added subcode 11 (Role Mismatch) for BGP role conflict scenarios.

Code 3: UPDATE Message Error. The peer sent a route with a malformed or invalid attribute. This could be a bad AS-path (including loops detected by the receiver), an invalid ORIGIN, a malformed NLRI, or an optional transitive attribute the receiver could not parse. The data field of the NOTIFICATION identifies the offending attribute. This is a peer-side bug or a route leak with malformed attributes. If it recurs, capture the UPDATE and report it to the peer.

Code 4: Hold Timer Expired. The peer did not receive keepalives or UPDATE messages within the negotiated hold time. In production, the most common root cause is control-plane CPU saturation on the peer. When CPU is pegged, BGP keepalive generation starves. Check the peer’s control-plane CPU. If you have SNMP visibility, poll cpmCPUTotal5sec on Cisco or the vendor equivalent. Correlate with any concurrent SNMP polling that might be contributing to CPU load.

Code 5: Finite State Machine Error. The peer received a BGP message it did not expect in its current FSM state. This is almost always a software bug, a race condition during session establishment, or a duplicate connection. Rare in stable production environments. If it recurs, capture debug output and report it to the vendor.

How to retrieve the last NOTIFICATION

The BGP4-MIB (RFC 4273) exposes the last error per peer via bgpPeerLastError at OID .1.3.6.1.2.1.15.3.1.14. This OCTET STRING encodes the error code and subcode from the last NOTIFICATION received from that peer. The value persists until the next NOTIFICATION or a session reset that clears it.

# Retrieve last error for all peers via SNMP
snmpwalk -v2c -c <community> <router> .1.3.6.1.2.1.15.3.1.14

# Vendor CLI: show the last notification received from a peer
ssh <router> 'show ip bgp neighbors <peer> | include notification|Last'

# Check syslog for recent BGP NOTIFICATION messages
ssh <router> 'show logging | include BGP'

# FRRouting equivalent
vtysh -c 'show bgp neighbors <peer>'

The bgpBackwardTransition notification trap (also defined in RFC 4273) fires on session state transitions. If your trap receiver is configured to accept BGP traps, the trap payload includes the old and new state plus the peer address. Traps are UDP and can be dropped under load. For reliability, pair trap monitoring with periodic SNMP polling of bgpPeerState at .1.3.6.1.2.1.15.3.1.2.

Monitoring signals to correlate with NOTIFICATION messages

A NOTIFICATION tells you what happened. These signals tell you why.

SignalWhy it mattersWarning sign
Per-peer prefix countSudden increase before Cease/1 indicates a route leak or growth past the limitPrefix count rising sharply in minutes before the NOTIFICATION
Control-plane CPUHold Timer Expired (code 4) is frequently CPU-inducedCPU above 90% sustained before the session drop
RPKI/ROA validation stateInvalid routes from the peer confirm a leak or hijackRPKI-invalid count above zero from the same peer
BFD session stateSubcode 10 traces back to a BFD failureBFD session in Down state at the same timestamp
bgpPeerInUpdates rateEstablished with zero updates means stale routingUpdate rate flat for an extended period without a session state change
Syslog severity distributionNOTIFICATION messages often arrive in clusters during incidentsMultiple BGP events from the same peer within minutes

How Netdata helps

  • SNMP-based BGP session monitoring: Netdata polls bgpPeerState per peer and can alert on transitions out of Established. Pair this with bgpPeerLastError to surface the code and subcode without parsing syslog.
  • Syslog ingestion and parsing: Netdata’s syslog collector captures BGP NOTIFICATION messages as they arrive, letting you correlate the exact timestamp with interface state, CPU, and prefix-count changes on the same timeline.
  • Control-plane CPU correlation: When Hold Timer Expired fires, Netdata shows the CPU trend alongside the session drop, making the root cause visible in seconds.
  • Prefix-count trends: Per-peer prefix counts tracked over time reveal whether Cease/1 was a sudden leak or gradual growth that crossed a threshold.
  • Trap collection: The SNMP trap receiver catches bgpBackwardTransition notifications, providing push-based session transition alerts alongside polled state.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.