The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-bgp-rib-fib-growth

Operations Guides

BGP RIB and FIB growth: monitoring route-table size before it bites

The global BGP routing table grows every year. In 2026, the IPv4 default-free zone sits at approximately 940,000 prefixes, with IPv6 adding roughly 190,000 more. These numbers increase steadily, and the hardware that programs forwarding decisions from them has finite capacity. When that capacity runs out, new routes do not get installed in the forwarding plane. Traffic to affected destinations blackholes. The BGP session stays Established the entire time.

This is one of the few failure modes where the control plane reports healthy while the data plane silently breaks. If you monitor BGP session state but not route-table size, you will not know the router has stopped installing routes until users complain.

What it is and why it matters

The Routing Information Base (RIB) is the set of routes a router has learned from all sources: BGP peers, IGP, static routes, connected networks. The Forwarding Information Base (FIB) is the optimized subset of the RIB that the hardware forwarding plane (typically TCAM on modern routers) uses to make packet-forwarding decisions. The RIB lives in control-plane memory. The FIB lives in hardware.

The RIB is always larger than the FIB. A router accepting full Internet feeds from three peers accumulates every received path in the RIB, including rejected, hidden, and dampened routes. On a multi-peer edge router, the RIB can hold 2 to 3 times more entries than the active FIB. Both resources are finite: the RIB consumes control-plane memory, and the FIB consumes TCAM.

When TCAM fills, the router stops installing new routes into the FIB. The RIB keeps growing because BGP sessions are still exchanging UPDATE messages. The control plane has no idea the data plane is rejecting routes. No trap fires. No syslog entry appears. The BGP FSM stays Established. Traffic to the affected destinations starts taking suboptimal paths, loops, or blackholes.

The global table crossed 768,000 IPv4 routes in 2019, triggering TCAM exhaustion on legacy Cisco Catalyst 6500 and 7600 platforms with older supervisor engines. Non-XL supervisors cap at 256,000 IPv4 routes and cannot be extended. The XL variants (SUP720-3BXL and SUP720-3CXL) support extended capacity but still have hard limits. The “768k day” was a preview of what happens when organic growth outpaces hardware capacity without warning.

How it works

The following diagram shows how multiple BGP peers feed the RIB, how best-path selection produces the FIB, and where the silent failure occurs when TCAM capacity is exhausted:

flowchart TD
    PA[Peer A: full feed] --> RIB[RIB: all received paths]
    PB[Peer B: full feed] --> RIB
    PC[Peer C: partial feed] --> RIB
    RIB -->|best path selection| FIB[FIB: active routes only]
    FIB -->|program into hardware| TCAM[TCAM: finite capacity]
    TCAM -->|below limit| OK[Forwarding works]
    TCAM -->|at capacity| DROP[New routes silently rejected]
    DROP --> SESS[BGP session stays Established]
    DROP --> BH[Traffic blackholed for new destinations]

The failure is not between the peers and the RIB. BGP keeps working. The failure is between the FIB and TCAM. Once TCAM is full, the forwarding ASIC cannot accept additional entries. The control plane has no standard mechanism to alert on this condition. Detection requires actively monitoring route counts and comparing them against known hardware limits.

Where it shows up in production

Internet edge routers accepting full feeds. An edge router taking full tables from two or three transit providers is the most common exposure. Each peer contributes roughly 940,000 IPv4 prefixes and 190,000 IPv6 prefixes. The RIB holds all of them. The FIB holds the best-path subset. As the global table grows organically at 5 to 10 percent per year, a router that had comfortable headroom three years ago may be approaching its limit today.

Multi-VRF environments. Standard BGP MIBs (RFC 1657, later updated by RFC 4273) do not expose per-VRF route counts natively. Cisco’s cbgpPeer3Table in CISCO-BGP4-MIB adds VRF context via cbgpPeer3VrfId and cbgpPeer3VrfName. Juniper gained per-VRF RIB monitoring through BMP starting in Junos OS Release 22.4R1. If you run VRFs and rely on standard MIB polling, you may have per-VRF blind spots.

Legacy hardware in critical paths. Older supervisor engines have hard TCAM limits. Cisco Catalyst 6500 and 7600 platforms with non-XL supervisors cannot hold the current full table. The mls cef maximum-routes ip 1000 command can reallocate TCAM from IPv6, multicast, or MPLS toward IPv4 on these platforms. This requires a configuration save and reboot. Check current per-protocol route counts with show mls cef summary before applying.

BGP session teardown from prefix limits. When a peer exceeds a configured maximum-prefix limit, the router sends a BGP Cease notification with subcode 1 (Maximum Prefixes Reached, RFC 4486) and tears down the session. This is a different failure: it is loud, it drops the session, and it is visible. But it only fires if the limit is configured. Routers without maximum-prefix limits on their peers silently absorb whatever the peer sends until TCAM fills.

When this matters

Route-table growth monitoring is most critical when:

  • You accept full or near-full Internet feeds from one or more peers.
  • You operate hardware with known TCAM limits within a few years of the global table size.
  • You run multiple VRFs on shared hardware, where one VRF’s growth can consume another’s capacity.
  • You have recently changed peers, filters, or routing policy, causing sudden shifts in received prefix counts.

It is less critical when:

  • You receive only a default route or partial feeds (your FIB stays small).
  • Your hardware has TCAM capacity orders of magnitude above the global table.
  • You operate in a closed network with controlled prefix growth (enterprise core, campus, OT networks).

A common misuse is monitoring BGP session state as a proxy for route health. Session Established does not mean routes are installed. It only means the TCP session and BGP keepalive exchange are functioning. Track prefix counts and FIB utilization separately.

Signals to watch in production

SignalWhy it mattersWarning sign
BGP RIB prefix count per peerDetects route leaks, upstream changes, mass withdrawals, and organic growthSudden change greater than 20% in 5 minutes
FIB utilization vs TCAM capacityApproaching the cliff edge where routes stop being installedAbove 70% of TCAM capacity
Per-peer prefix count vs maximum-prefix limitPrevents Cease/1 teardown and gives operational headroomWithin 10% of configured limit
RIB-to-FIB size ratioGrowing ratio means more rejected paths consuming memoryRatio exceeding 2-3x on multi-peer full feeds
BGP NOTIFICATION Cease subcode 1Peer exceeded configured prefix limit; possible route leakAny occurrence is actionable
BGP session Established with flat prefix count“Stale Established”: session up, no UPDATE exchangeLast prefix-receive timestamp not advancing
Control-plane CPU during convergenceTable rebuilds and BGP reconvergence are CPU-intensiveSustained above 90% during reconvergence
Device control-plane memory utilizationRIB growth consumes control-plane memory directlyFree memory approaching 5% or declining rapidly
Internet DFZ growth rateIndustry trend that sets your external baselineGrowth exceeding 10% year-over-year

Instrumentation paths

SNMP polling. The standard BGP4-MIB provides bgpPeerInUpdates and bgpPeerOutUpdates at .1.3.6.1.2.1.15.3.1.10 and .1.3.6.1.2.1.15.3.1.11. These track UPDATE message counts, not prefix counts directly. For per-peer accepted prefix counts on Cisco, use cbgpPeer2AcceptedPrefixes in CISCO-BGP4-MIB (1.3.6.1.4.1.9.9.187). The index includes peer address, AFI, and SAFI. Verify the exact suffix on your target IOS XE version.

The older bgp4PathAttrTable indexes by prefix first, requiring a full table walk to isolate a single peer’s routes. This is operationally inefficient at scale. The cbgpRouteTable indexes by peer address first, then prefix, enabling efficient per-peer polling without walking the entire BGP RIB. Prefer the newer table where available.

CLI verification. The most reliable on-box commands:

# Per-peer prefix counts
ssh <router> 'show ip bgp summary | begin Neighbor'

# Route table size by protocol
ssh <router> 'show ip route summary'

# FIB summary (route counts; TCAM utilization requires platform-specific commands)
ssh <router> 'show ip cef summary'

# Per-peer received routes (requires soft-reconfiguration inbound on the peer;
# memory-intensive on full-feed peers)
ssh <router> 'show ip bgp neighbors <peer> received-routes | include Total'

For TCAM utilization specifically, the command varies by platform. On Catalyst 6500/7600, use show mls cef summary. On Nexus 9000, use show hardware access-list resource profile. On Juniper MX, use show chassis forwarding.

BMP for deep visibility. BGP4-MIB’s bgp4PathAttrTable only reflects best-path routes. Adj-RIB-In entries (pre-policy and post-policy routes) are not exposed via standard SNMP. For full visibility into what each peer is actually sending before your local policy filters it, deploy BMP (RFC 7854). BMP streams per-prefix, real-time route data to a monitoring station, including withdrawals and AS-path changes that BGP4-MIB cannot surface. Juniper supports per-VRF BMP monitoring from Junos 22.4R1 onward.

Note that show ip bgp summary resets counters on clear ip bgp *, so baselines are not preserved across a reset. Track prefix counts via SNMP or BMP for persistent trending.

How Netdata helps

  • Per-peer prefix count trending at minute granularity. Netdata’s SNMP collector can poll BGP4-MIB and vendor-specific BGP MIBs at per-minute intervals. A sudden 20% spike triggers alerting before it becomes a route leak or TCAM overflow. Correlate with the detection signals in BGP route leak and hijack.
  • FIB-to-TCAM ratio monitoring. If your platform exposes FIB utilization via SNMP or CLI scrape, Netdata can alert when utilization crosses 70% of capacity.
  • Control-plane CPU and memory correlation. When BGP reconvergence occurs (new peer, mass withdrawal, route reflector change), CPU and memory spikes accompany the RIB rebuild. Netdata correlates these alongside prefix-count changes so you can see whether a table rebuild is progressing or stuck.
  • Stale session detection. A BGP session that stays Established while prefix counts flatten is a silent failure. Per-minute polling of bgpPeerInUpdates catches the moment UPDATE traffic stops, even when the FSM reports healthy. See BGP session Established but stale.
  • Cease subcode alerting. A Cease/1 (Maximum Prefixes) notification is always actionable. Netdata can parse BGP NOTIFICATION events from syslog or traps and alert immediately, distinguishing routine maintenance shutdowns from prefix-limit violations.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.