The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-flow-export-ingest-latency

Operations Guides

Flow export-to-ingest latency: why your NetFlow data is minutes behind

Flow export-to-ingest latency accumulates across a pipeline: the exporter’s active timeout, the UDP transport path, the kernel socket buffer, the collector’s parser, and the storage write queue. Each stage can add seconds or minutes, and each has a different fix.

The most common cause is the active timeout default on most network devices: 30 minutes. Long-lived flows (VPN tunnels, database connections, bulk transfers) are not exported until the timer expires. The collector is not slow and the network is not congested. The device is behaving as configured. But if you need near-real-time visibility, a 30-minute export delay is indistinguishable from broken telemetry.

What this means

Flow export-to-ingest latency is the time between when a flow record leaves the exporter and when the collector ingests it. On NetFlow v9, IPFIX, and sFlow, you can compute this directly from timestamps embedded in the records. On NetFlow v5, there is no exporter-side timestamp, so you must infer latency by comparing flow end times against wall clock time or SNMP-derived traffic patterns.

Sustained latency above 30 seconds means something in the pipeline is bottlenecked. Above 5 minutes, the data is too stale for real-time detection use cases such as DDoS visibility, anomaly detection, or operational troubleshooting. The causes fall into two categories: configuration-driven latency (active timeout too long, template refresh misaligned) and resource-driven latency (collector backlog, UDP buffer drops, parser saturation).

flowchart LR
    A["Device flow cache
active timeout: 30 min default"] --> B["Export buffer"] B -->|"UDP datagram"| C["Kernel socket buffer
rmem_max varies by distro"] C -->|"drain"| D["Parser / aggregator"] D --> E["TSDB write queue"] E --> F["Stored and queryable"] B -.->|"overflow = silent drops"| G["cnfESPktsDropped"] C -.->|"overflow = silent drops"| H["Udp_RcvbufErrors"] D -.->|"backlog = growing latency"| I["Queue depth rising"]

Common causes

CauseWhat it looks likeFirst thing to check
Active timeout too long (30 min default)Long-lived flows arrive in batches every 30 minutes; short flows arrive fineCheck the exporter’s active timeout configuration
Collector UDP buffer overflowTraffic charts show declining volume while SNMP interface counters show traffic is normal or risingcat /proc/net/snmp | grep '^Udp:' for RcvbufErrors
Collector parser or write backlogFlow latency grows steadily; queue depth metric risingCheck collector intake and write queue depth
Device-side export buffer dropsDevice exports fewer records than expected; gap between device and collector countsCisco: snmpget .1.3.6.1.4.1.9.9.387.1.4.6 (cnfESPktsDropped)
Template desync (NetFlow v9 / IPFIX)Datagrams arriving but decoded records are zero or anomalously lowCheck collector logs for “template not found” or cache miss
Clock skew on exporterLatency values are impossible (negative or wildly inconsistent)Check NTP offset on the exporter

Quick checks

# Kernel UDP socket buffer drops on the collector
cat /proc/net/snmp | grep '^Udp:'
# RcvbufErrors column: any nonzero value is silent data loss

# Current socket buffer fill for the flow listener
ss -lun '( sport = :2055 )' -m
# NetFlow v5/v9 uses port 2055; IPFIX uses 4739; sFlow uses 6343

# Current rmem_max setting
sysctl net.core.rmem_max

# Inspect flow record timestamps for latency (nfdump)
nfdump -R /var/nfdump/ -o fmt:'%ts %te %fl' | head -50
# Compare flow end time against current wall clock

# Device-side export drops (Cisco IOS-XE)
snmpget -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4.6.0
# cnfESPktsDropped: should be near zero relative to exported count
<!-- TODO: verify these Cisco OIDs. They vary by platform, IOS version, and MIB support. Test against your specific device. -->

# Device-side total exported (Cisco IOS-XE)
snmpget -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4.4.0
# cnfESPktsExported: compare rate against collector inbound rate

# Collector CPU per core (look for single-core saturation from RSS)
mpstat -P ALL 1 5
# Watch %soft (softirq) and %sys columns

# Collector write queue depth (vendor-specific stats endpoint)
curl -s http://localhost:<stats-port>/metrics | grep -E 'write|queue'

# NTP offset on the exporter (if SNMP accessible)
snmpget -v2c -c <community> <device> .1.3.6.1.2.1.25.1.2.0
# Returns device-local time (hrSystemDate); compare against collector time

How to diagnose it

  1. Determine whether the latency is configuration-driven or resource-driven. If all flows from a single exporter are uniformly delayed (long-lived flows arriving at 30-minute intervals), the active timeout is the cause. If latency varies and correlates with traffic volume, the bottleneck is in the collector pipeline.

  2. Check the active timeout on the exporter. On Cisco IOS-XE, inspect the flow monitor configuration. The default is typically 30 minutes. If you need near-real-time visibility, set it to 60 seconds. The tradeoff is more export packets and higher collector load. Juniper and other vendors have equivalent settings under flow sampling or monitoring configuration.

  3. Check for UDP socket buffer drops. Run cat /proc/net/snmp | grep '^Udp:' and look at the RcvbufErrors column. Any nonzero increment means the kernel received datagrams but could not deliver them to the application because the socket buffer was full. This is silent data loss, not just latency. The dropped records never arrive.

  4. Compare device-side export counts against collector inbound counts. Poll cnfESPktsExported on the device and compare the rate against what the collector reports receiving. A gap means records are being lost in transit through the UDP buffer, NIC ring, or network path.

  5. Check collector parser and write queue depth. If the collector exposes a queue depth metric, watch it over time. A growing queue means the parser or storage layer cannot keep up with the incoming rate. The backlog adds latency to every record in the queue.

  6. Verify template state if using NetFlow v9 or IPFIX. After a device reboot or configuration change, templates must be re-received before data records can be decoded. Check collector logs for template-related messages. The blind window can be 5 to 30 minutes depending on the template refresh interval.

  7. Check for clock skew. If the exporter’s clock is off, the computed latency will be wrong. Verify NTP synchronization on the exporter. Even 200ms of drift can break cross-device correlation in postmortem analysis.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Flow export-to-ingest latencyDirect measure of pipeline delaySustained above 30 sec; above 5 min means data is stale
UDP socket buffer drops (Udp_RcvbufErrors)Silent data loss at the kernel layerAny nonzero increment
Flow packets received rateConfirms exporter is sending and collector is receivingSudden drop from one exporter means exporter or path issue
Flow exporter drops (cnfESPktsDropped)Device-side overflow before records leave the deviceRising rate during traffic bursts
Collector inbound vs device exported rateEnd-to-end loss detectionGap above 0.1% of exported rate
Collector CPU per coreSingle-core saturation from RSS funnelingOne core at 100% while others are idle
TSDB write queue depthStorage-layer backlogGrowing queue without bound
NTP offset on exporterCorrupts latency computation and cross-device correlationOffset above 100ms
Template cache state (v9 / IPFIX)Records discarded if template is missingCollector logs showing template cache misses

Fixes

Active timeout tuning

The active timeout controls how often long-lived flows are exported. The default of 30 minutes means a persistent flow is invisible until either it ends or the timer expires. For real-time use cases, set it to 60 seconds.

The tradeoff is increased export volume. Shorter timeouts generate more flow records and more UDP datagrams. A 60-second active timeout on a busy exporter can increase export packet rate by an order of magnitude compared to 30 minutes. Ensure your collector can handle the additional load before making this change globally.

Cisco ASA note: The flow-export active refresh-interval command was removed in ASA versions 8.5(1) through 9.1(1), causing persistent tunnels to export only at teardown. If you are running an affected version, upgrading to 9.1(2) or later restores the capability.

UDP buffer sizing

If Udp_RcvbufErrors is incrementing, the kernel socket buffer is overflowing. The fix has two parts:

  1. Raise the system maximum. Run sysctl -w net.core.rmem_max=16777216 to set it to 16 MB. For very high-volume collectors, use 33 MB. This is a runtime change; persist it in /etc/sysctl.d/ for reboot survival.

  2. Ensure the collector application explicitly sets SO_RCVBUF on its listener socket. Without this, the socket uses rmem_default (typically around 212 KB on most Linux distributions), which is far too small for flow collection regardless of rmem_max.

Also check RSS (Receive Side Scaling) configuration. If all flow traffic is funneled to a single CPU core, that core becomes the bottleneck even though the system has spare capacity. Verify IRQ distribution with cat /proc/interrupts | grep eth0.

Collector parser backlog

If the write queue is growing, the parser or storage layer is the bottleneck. Options:

  • Increase parser thread count if the collector supports it.
  • Check whether the parser is doing expensive work per record (regex matching, enrichment lookups) that can be deferred or batched.
  • Verify disk IOPS are not saturated. Flow record storage is disk-intensive.
  • Separate flow storage from syslog storage on different volumes.

Template refresh alignment

For NetFlow v9 and IPFIX, templates must be received before data records can be decoded. After a device reboot or config change, the collector is blind until the next template arrives. Set the template refresh interval shorter than the active timeout. A misaligned configuration where template refresh is longer than active timeout causes template expiry before data records arrive, leaving the collector unable to decode flows.

RFC 7011 Section 8.4 requires that exporting processes using UDP periodically retransmit active templates. Ensure your exporter’s template refresh rate is configured and that the collector’s template cache lifetime is at least three times the retransmission interval.

Prevention

  • Set active timeout to 60 seconds on exporters where real-time visibility matters. Document the increased export load and size collectors accordingly.
  • Monitor Udp_RcvbufErrors continuously. This is the single most missed signal in flow collection. Any nonzero value represents silent data loss. Set rmem_max to 16 MB or higher and confirm the collector sets SO_RCVBUF.
  • Compare device-side export counts against collector inbound counts. This is the only reliable end-to-end loss detection method. A sustained gap means records are lost somewhere in the path.
  • Verify NTP on exporters. Clock skew corrupts latency computation and breaks cross-device correlation.
  • Test template cache persistence across collector restarts. Some collectors lose in-memory template state on restart, creating a 5 to 30 minute blind window.

How Netdata helps

  • UDP buffer drop detection. Netdata monitors Udp_RcvbufErrors from /proc/net/snmp with per-second resolution, catching socket buffer overflow before it becomes a data gap.
  • Per-core CPU breakdown. Netdata’s CPU collector shows softirq and system time per core, making RSS funneling visible without manual mpstat sessions.
  • Collector process metrics. Netdata monitors the collector’s own CPU, memory, and disk I/O, correlating collector-side bottlenecks with flow data latency.
  • NIC-level drop counters. Netdata tracks /proc/net/dev RX drops and ethtool counters such as rx_missed_errors, distinguishing NIC-level drops from socket-buffer drops.
  • Cross-signal correlation. When flow data latency spikes, Netdata’s unified timeline lets you correlate it with CPU saturation, disk I/O wait, and UDP buffer errors in a single view.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.