The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-monitoring-checklist

Operations Guides

Network monitoring checklist: the signals every production network needs

This checklist covers the signals production networks need, organized by detection priority and mapped to maturity levels from survival to expert.

An NPM stack is a federation of collectors, parsers, enrichment services, storage tiers, and an analytics core. Most production incidents are not “the network broke” but “a collector’s UDP buffer dropped packets,” “the NetFlow v9 template cache went stale after a device reboot,” or “the polling worker pool fell behind and now a healthy device looks down.” The checklist is organized to surface those failure modes, not just the top-level symptoms.

The federation at a glance

When a signal is missing, stale, or wrong, the fault is usually one or two subsystems upstream of the dashboard. The typical NPM stack includes:

  • Time synchronization substrate (NTP/PTP). Every cross-collector correlation depends on accurate, monotonic time across collectors, polled devices, and API endpoints.
  • Polling transport. ICMP, UDP/161 (SNMP), TCP/22 (SSH/CLI scrape), and HTTPS (vendor APIs) reaching each managed endpoint.
  • SNMP polling engine. A scheduler fanning OID requests across devices with timeouts, counter tables, and a device state machine (UP / STALE / UNKNOWN / DOWN).
  • Flow collection subsystem. NetFlow v5/v9 and IPFIX collectors with template caches, sampling-rate awareness, and flow record storage. sFlow is sample-datagram oriented, not template-flow, and has a different failure profile.
  • Topology inference engine. Fuses CDP/LLDP neighbor tables, FDB entries, ARP tables, STP state, and routing tables to derive Layer-2 and Layer-3 topology.
  • BGP monitoring subsystem. Active or passive sessions tracking FSM state, prefix announcements, AS-path changes, and RPKI validity.
  • Syslog and trap ingestion. UDP/TCP/TLS listeners with parser backpressure, facility/severity handling, and deduplication.
  • Vendor API integration layer. Pull-mode clients for SD-WAN controllers, cloud platforms, and modern firewalls, each with their own auth, rate limit, and pagination semantics.
  • Storage tiers. Counter TSDB (downsampled for long retention), full-resolution flow store, topology graph DB, raw syslog store, and event/alert log.

Signal domains by detection priority

The domains below are ordered by detection priority: the earliest surfacing of real issues with the best signal-to-noise comes first. Within each domain, the most operationally critical signals are listed first.

Availability

SignalSourceWhy it matters
SNMP agent reachability (sysUpTime)SNMP GET .1.3.6.1.2.1.1.3.0No response means agent down, partition, ACL block, or credential issue. Value decrease means reboot. SNMP down with healthy ICMP means agent problem, not device outage.
ICMP reachabilityping, fpingLiveness independent of SNMP. ICMP down plus SNMP down equals network problem. ICMP down plus SNMP up equals ICMP rate-limited or blocked (common on firewalls, CoPP).
Vendor API reachability and validityHTTPS to vendor endpointFor SD-WAN/cloud, the API may be the only telemetry source. HTTP 200 with empty or error payload (PAN-OS <response status="error"> inside HTTP 200) is a silent failure.
Flow UDP packet receipt rate/proc/net/udp, collector stats, nstatDrop to 0 from one exporter means exporter stopped or partitioned. Drop from all exporters means collector-side failure.
Syslog receipt rate and severityUDP/TCP/TLS port 514 listenerRate spike with severity escalation means device event. Spike without escalation means noise storm. Silence from a normally-chatty device means isolation or logging failure.
SNMP trap rate and typeUDP port 162 listenerlinkDown/linkUp pairs mean flap. coldStart means reboot. Silence from a noisy device means trap path broken.
BGP session state (FSM)BGP4-MIB .1.3.6.1.2.1.15.3.1.2, CLI, BMPEstablished means exchanging routes. Established with no UPDATE traffic (stale session) is a worse failure than Idle.
Interface operational statusIF-MIB .1.3.6.1.2.1.2.2.1.8 (ifOperStatus)Admin up plus oper down means physical or link-layer failure. Flapping means link instability.

Errors

SignalSourceWhy it matters
Interface errors (ifInErrors, ifOutErrors)IF-MIB .1.3.6.1.2.1.2.2.1.14, .20Incrementing counters mean cable/fiber degradation, SFP failure, duplex mismatch, or EMI. Rate of change matters more than absolute value.
Interface discards (ifInDiscards, ifOutDiscards)IF-MIB .1.3.6.1.2.1.2.2.1.13, .19Queue or buffer overflow, or ACL drops. Often the leading indicator of congestion before utilization shows 100%.
UDP socket buffer drops/proc/net/snmp, nstat -az Udp_RcvbufErrorsThe number one silent killer for flow, trap, and syslog collectors. Datagrams arrive at the kernel but the application was too slow to drain. Any nonzero value means lost telemetry.
SNMP timeout and retry rateCollector stats, time snmpgetRising across many devices means collector-side issue. Rising on one device means device-side agent or CPU issue.
BGP NOTIFICATION and Cease messagesbgpBackwardTransition trap, CLI, syslogCease/1 is maximum prefixes reached. Cease/2 is administrative shutdown. Hold Time Expired (NOTIFICATION code 4) indicates CPU saturation.
License and feature validityVendor MIBs, PAN-OS API, Meraki API, Cato GraphQLFeature silently disabled at midnight. Users complain at 09:00. The most common root cause of “the firewall stopped doing what we paid for.”

Saturation

SignalSourceWhy it matters
Interface utilization (% of ifHighSpeed)IF-MIB ifHCInOctets .1.3.6.1.2.1.31.1.1.1.6, ifHCOutOctets .10, ifHighSpeed .1595% sustained for over 5 min on critical interface means congestion with drops and latency. Use 64-bit HC counters. 32-bit ifInOctets wraps in approximately 3.4 seconds at 10G line rate.
NIC RX/TX drops on collector/proc/net/dev, ethtool -SRing buffer overflow before packets reach the socket layer. rx_missed_errors is the most actionable counter.
Collector CPU (per-core, %soft)mpstat -P ALL, /proc/softirqsHigh %soft on one core means RSS funneling all packet processing to one CPU. Total CPU may look fine while one core is pinned.
Collector disk and TSDB write queuedf, iostat, collector metricsCardinality inflation (new subnet, NAT pool, scanner traffic) can fill disk in hours. Write queue growing means TSDB cannot keep up with ingestion.
Device control-plane CPUCisco .1.3.6.1.4.1.9.9.109.1.1.1.1.7, Juniper .1.3.6.1.4.1.2636.3.1.13.1.8Sustained over 90% means SNMP starvation, BGP hold-time expiry, and session drops.
Device memory utilizationCisco .1.3.6.1.4.1.9.9.48.1.1.1.5, HOST-RESOURCES-MIBFree memory approaching 0 means OOM imminent. Rate of increase over 1%/min means memory leak.
BGP RIB and FIB sizeBGP4-MIB prefix counts, CLISudden change over 20% in 5 min means route leak or mass withdrawal. Full IPv4 DFZ in 2026 is approximately 940k prefixes.
NAT and session table utilizationPAN-OS API, vendor CLIApproaching limit means new connections denied. Sustained growth means traffic outpacing NAT capacity.
API rate-limit remainingHTTP headers (Retry-After, X-RateLimit-Remaining)Meraki: 10 req/sec/org with burst of 30 in 2 sec. Cato: 120/min general, accountSnapshot 1/sec, accountMetrics 15/min, eventsFeed 100/min.

Internal state, replication, and correctness

SignalSourceWhy it matters
Device uptime (sysUpTime)SNMP .1.3.6.1.2.1.1.3.0Decrease means reboot. 32-bit wrap at approximately 497 days looks like reboot; track wraps separately.
Temperature, fan, power supplyENTITY-SENSOR-MIB .1.3.6.1.2.1.99.1.1.1.4Thermal failure, cooling failure, or redundancy lost. Use vendor-defined thresholds, not arbitrary absolute numbers.
Interface counter discontinuityifCounterDiscontinuityTime .1.3.6.1.2.1.31.1.1.1.3Counter reset without sysUpTime reset means SNMP agent inconsistency or counter-source bug.
Cross-collector time skewntpq -p, chronyc trackingOver 100ms drift breaks cross-site flow correlation. Over 1s breaks it entirely.
NTP offset on monitored deviceshrSystemDate .1.3.6.1.2.1.25.1.2.0Device clock drift causes postmortem correlation failure. Consistently the most under-monitored NTP signal.
Topology view consistencyCDP/LLDP vs FDB vs ARP cross-validationInconsistency means stale data, topology change in progress, or device bug. Three sources agreeing is high confidence; one source alone is low.
Flow sampling rate consistencysFlow MIB, NetFlow v9 template fieldsMismatch means analytics wrong by orders of magnitude. Without sampling-rate correction, sFlow at 1:1000 reports 1/1000 of true traffic.
STP root bridge and TCNBRIDGE-MIB .1.3.6.1.2.1.17.2Root bridge change means reconvergence. TCN rate over 5/min means instability.

Latency, throughput, and security

SignalSourceWhy it matters
SNMP poll response latencytime snmpget, collector statsOver 1s on a normally-fast device means agent or management-network degradation.
ICMP round-trip timeping, fpingp99 over 2x rolling baseline means congestion or path change. High jitter means unstable path.
Active path probesCisco IPSLA RTTMON MIB, TWAMP, HTTP GETRTT and loss per path, independent of application. Loss over 1% sustained is degraded.
Flow bytes per conversationNetFlow/sFlow/IPFIX recordsTop talkers, DDoS patterns, data exfiltration signals. sFlow requires sampling-rate multiplication for accurate byte counts.
Poller poll cycle durationCollector internal statsCycle exceeding configured interval means data is drifting stale. The most under-monitored meta-signal in NPM.
Flow exporter drop rate (device-side)Cisco CISCO-NETFLOW-MIB (walk .1.3.6.1.4.1.9.9.387 to find the drop counter for your platform)Device dropped flows that never reached collector. Invisible to collector alone. Compare device-exported rate against collector inbound rate for end-to-end loss detection.
Unauthorized SNMP accesssnmpInBadCommunityNames .1.3.6.1.2.1.11.4, USM statsBurst from single source means scanning. Persistent events from many sources means community string “public” still configured.
BGP RPKI/ROA invalid acceptanceVendor CLI show bgp rpki, validatorsAny RPKI-invalid route accepted in production is a security event. Verify with public validators before alerting; stale cache produces false invalids.
Config changes without ticketSyslog CONFIG-I, AAA logs, config diffChange outside maintenance window without change ticket means unauthorized or emergency. Change followed within 30 min by incident is a high-correlation root-cause candidate.

Monitoring maturity levels

These levels are sequential and cumulative. Each level includes everything below it.

flowchart TD
    L4["L4 Expert
BMP, RPKI integrity, per-VRF,
sampling-rate forensics"] --> L3["L3 Mature
UDP drops, RSS, flow end-to-end loss,
NTP on devices, topology confidence"] L3 --> L2["L2 Operational
CPU and memory, traps, licenses,
topology discovery, API status"] L2 --> L1["L1 Survival
sysUpTime, ifOperStatus, BGP state,
utilization, errors, syslog, trap port"]

L1: survival

The absolute minimum to know if the network is alive and not on fire:

  • SNMP reachability (sysUpTime GET) for every critical-path device
  • Interface operational status (ifOperStatus) for critical interfaces
  • BGP FSM state for critical eBGP and iBGP peers
  • Interface utilization (ifHCInOctets / ifHCOutOctets vs ifHighSpeed) for top-10 interfaces
  • Interface error counters (ifInErrors, ifOutErrors) for critical interfaces
  • Syslog severity 0-3 (EMERG through ERR) forwarded from critical devices
  • Flow collector port listening (UDP 2055 for NetFlow, 6343 for sFlow, 4739 for IPFIX)
  • Trap receiver bound on UDP 162

A team at L1 catches hard outages. Nothing else.

L2: operational

Everything in L1, plus:

  • All interfaces for status, utilization, errors
  • All BGP peers for FSM state and prefix count
  • SNMP poll latency and timeout rate per device
  • Device control-plane CPU and memory
  • Flow records received per second; syslog source count
  • License days-to-expiry for all licensed features
  • Temperature, fan, power supply state
  • Topology discovery (CDP/LLDP)
  • STP root bridge identity and topology change count
  • Vendor API HTTP status for SD-WAN and cloud
  • SNMP authentication failure rate
  • ColdStart/warmStart detection with alerting

A team at L2 has visibility into most failures. They still miss silent failures, license cliffs, and topology staleness.

L3: mature

Everything in L2, plus:

  • UDP socket buffer drops (Udp_RcvbufErrors) on flow, trap, and syslog collectors
  • NIC RX/TX drops and RSS IRQ distribution
  • Collector CPU (per-core, %soft) and disk space
  • TSDB write queue depth and series cardinality
  • Flow export-to-ingest latency and sampling rate consistency
  • Flow exporter drop rate (device-side) and inbound-vs-exported comparison
  • Interface counter discontinuity detection
  • NAT/session table utilization
  • Vendor API request latency, error rate, and rate-limit remaining
  • Active path probes (IPSLA/TWAMP/HTTP) on critical paths
  • Cross-collector time skew and NTP offset on monitored devices
  • Topology view consistency and inference confidence score
  • BGP route advertisement vs reception symmetry
  • Poller poll cycle duration vs configured interval
  • ARP cache entry count and staleness
  • RPKI/ROA validation state for all BGP sessions
  • BGP NOTIFICATION Cease subcode parsing (RFC 4486, RFC 8538, RFC 9384)
  • Configuration drift detection
  • Endpoint positioning orphan rate

A team at L3 catches most incidents in their early stages.

L4: expert

Everything in L3, plus the signals operators add after multiple major incidents:

  • BMP (RFC 7854) for Adj-RIB-In visibility (pre-policy and post-policy routes). BGP4-MIB bgp4PathAttrTable only reflects best-path routes; Adj-RIB-In entries return “NA” over SNMP.
  • BGP AS-path baseline deviation detection for own prefixes and upstreams
  • RPKI validator health monitoring; alert on “Unknown” rate changes (signals validator outage)
  • Sub-prefix hijack detection: alert when a more-specific appears without a less-specific in the RIB
  • Smart License and vendor license server reachability monitored continuously
  • Per-VRF and per-tenant isolation: BGP RIB size, flow volume, license utilization tracked per VRF
  • Per-priority-queue discard counters (vendor QoS MIBs) revealing QoS queue saturation behind moderate utilization
  • CoPP (control-plane policer) drop counters
  • NIC per-queue drop counters via ethtool -S (rx_missed_errors, rx_no_dma_resources)
  • /proc/net/softnet_stat for kernel packet processing backpressure
  • FDB/ARP entry freshness (time since last refresh, computed from polling deltas)
  • Flow template cache hit/miss ratio for NetFlow v9/IPFIX
  • License grace-period state with feature-specific counter validation (IPS drops at 0 when traffic flows after license expiry)
  • Asymmetric routing detection (forward vs reverse probe comparison)

The signals most teams miss

These are the systematic blind spots that keep causing incidents:

  1. UDP socket buffer drops are not monitored. Udp_RcvbufErrors is the number one missed signal in flow collection. Charts show declining traffic during incidents that are actually traffic spikes. Production flow collectors need net.core.rmem_max raised to 16 MB or higher, tuned to actual ingress volume.

  2. License expiry is monitored only when too late. A licensed feature (IPS, VPN, threat prevention) silently disables at midnight. The device stays up. The syslog message is low severity and buried. Users notice at 09:00.

  3. NTP drift on monitored devices is not watched. Two devices 200ms apart on the same flap produce records that do not correlate. Postmortems fail to reconstruct events because timestamps are seconds apart.

  4. BGP “Established but stale” is not detected. The FSM reports Established but UPDATE exchange stopped. Graceful Restart keeps the FSM green while the session is gone. Track bgpPeerInUpdates rate and the timestamp of last received prefix.

  5. Trap receiver drops are invisible. During a trap flood, the highest-priority trap (root cause) is statistically the most likely to be dropped. There is no per-source drop counter.

  6. Vendor API silent failures are not detected. HTTP 200 with empty payload is treated as “no data” rather than “API is broken.” PAN-OS returns <response status="error"> inside HTTP 200.

  7. NetFlow v9/IPFIX template desync is invisible. After a device reboot or upgrade, templates arrive on a 5-30 minute interval. Until then, all data records are silently discarded.

  8. Sampling rate normalization is skipped. sFlow analytics report raw counts without scaling. Bandwidth charts are wrong by the sampling factor (often 1:1000 or worse).

  9. 32-bit counter rollover is treated as a real spike. ifInOctets wraps in approximately 3.4 seconds at 10G line rate. Naive differencing produces terabit spikes or negative utilization.

  10. Poller fall-behind is not detected. The scheduler oversubscribes devices, retries compound, control-plane CPU spikes, and healthy devices appear “down.” The platform is the problem; the network is not.

How Netdata helps

Netdata collects many of the collector-side signals in this checklist that most NPM platforms miss:

  • SNMP data collection polls sysUpTime, ifOperStatus, ifHCInOctets/ifHCOutOctets, ifInErrors/ifOutErrors, ifInDiscards/ifOutDiscards, and device CPU/memory with configurable intervals down to 1 second.
  • Linux system plugins expose Udp_RcvbufErrors, NIC RX/TX drops from /proc/net/dev, per-core softirq from /proc/softirqs, and /proc/net/softnet_stat for kernel packet processing backpressure.
  • Network interface metrics include per-NIC ring buffer drops and ethtool -S counters like rx_missed_errors, so you can distinguish NIC-level drops from socket-buffer drops.
  • Cross-layer correlation lets you join rising Udp_RcvbufErrors with rising flow receive rate and rising collector CPU in a single view, which is the diagnostic chain for silent UDP flow loss.
  • NTP metrics surface offset and drift on collectors and, where SNMP exposes hrSystemDate, on monitored devices.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.