The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-netflow-vs-sflow-vs-ipfix

Operations Guides

NetFlow vs sFlow vs IPFIX: what they measure and how each one fails

Flow telemetry protocols are often lumped together as “flow data,” but they measure fundamentally different things. NetFlow and IPFIX build stateful flow records by tracking conversations in device memory. sFlow captures random packet samples without maintaining any flow state. This architectural split determines not only what you can see but how the data breaks when something goes wrong.

Operators who conflate the three protocols tend to misdiagnose the dominant failure they all share (silent UDP loss) and miss the protocol-specific failures that produce the most misleading data: template desync in NetFlow v9 and IPFIX, sampling blindness in sFlow, and exporter-side drops that neither NetFlow nor IPFIX reports in-band.

What each protocol measures

The three protocols divide into two architectural families: stateful flow aggregation (NetFlow v5, NetFlow v9, IPFIX) and stateless packet sampling (sFlow).

NetFlow v5 is effectively legacy. It exports a fixed 7-field record layout with no templates, no IPv6 support, no MPLS or VXLAN awareness, and ingress-only tracking. It still appears on older Cisco IOS deployments but should be treated as deprecated for any new build. NetFlow v5 also lacks an exporter-side timestamp, which means you cannot compute export-to-ingest latency from its records alone.

NetFlow v9 is the de facto current NetFlow standard. It exports stateful flow records aggregated in device memory (RAM or TCAM). Each record tracks L3/L4 attributes: source and destination IP, ports, protocol, ToS, byte and packet counts, and start and end timestamps. NetFlow v9 is template-based: the exporter sends a template record describing the field layout, followed by data records that conform to that template. Collectors must receive and cache the template before they can decode any data records.

IPFIX (RFC 7011/7012) is a direct derivative of NetFlow v9. Its template-based architecture is identical in concept, but it uses an IETF-maintained, vendor-neutral Information Element registry instead of Cisco-specific field definitions. IPFIX supports enterprise-specific custom fields and variable-length field encoding for arbitrary data such as URL fragments or usernames. It is the forward-looking standard, with active IETF governance and formal interoperability testing.

sFlow (RFC 3176, maintained by sflow.org) takes a fundamentally different approach. It performs stateless random packet sampling: on average, 1 in N packets is captured. The sampling is pseudorandom to avoid synchronizing with periodic traffic patterns. sFlow exports the full packet header (Ethernet through L4) plus the first 64 or 128 bytes of payload, and it also exports time-based interface counter samples. There is no unsampled mode. There are no flow timestamps in the NetFlow sense. sFlow is sample-datagram oriented, not template-flow oriented.

flowchart LR
    subgraph Stateful["NetFlow v9 / IPFIX (stateful)"]
        A1[Packet stream] --> A2[Device flow cache]
        A2 --> A3[Aggregated flow records]
        A3 --> A4[Template + data records]
        A4 --> A5[UDP export]
    end
    subgraph Stateless["sFlow (stateless)"]
        B1[Packet stream] --> B2[1-in-N random sampling]
        B2 --> B3[Sample datagrams]
        B2 --> B4[Counter samples]
        B3 --> B5[UDP export]
        B4 --> B5
    end

How they fail differently

All three protocols share one dominant failure mode: UDP transport with no retransmission, no acknowledgment, and no sender-side notification of loss. When a UDP datagram carrying flow data is dropped, every record inside it is permanently lost. The kernel counter Udp_RcvbufErrors (in /proc/net/snmp) is often the only signal that this happened.

Beyond the shared UDP weakness, each protocol fails in architecturally distinct ways.

NetFlow v9 and IPFIX: template desync

NetFlow v9 and IPFIX require collectors to receive and cache template records before data records can be decoded. Templates are sent over UDP on a configurable interval, typically 5 to 30 minutes. After a device reboot, software upgrade, or flow-exporter configuration change, the template may change. Until the collector receives the new template and rebuilds its cache, all data records from that exporter are silently discarded.

This window can be 5 to 30 minutes per exporter. If the restart coincides with a security event, the forensic traffic evidence is gone. Neither protocol provides an in-band signal that data is being discarded for this reason. The collector may log “template not found” or “cache miss” internally, but these messages rarely surface on operational dashboards.

NetFlow v9 and IPFIX: exporter-side silent drops

Most operators monitor collector-side packet loss (kernel buffer drops, NIC drops) but the exporter itself can silently discard flow records. Full internal buffers, FIB lookup failures, and VRF mismatches all cause the device to drop flow records before they are exported. Neither NetFlow nor IPFIX provides an in-band signal for these drops.

On Cisco devices, the CISCO-NETFLOW-MIB exposes counters for this purpose:

<!-- TODO: verify these OIDs map to the claimed objects in CISCO-NETFLOW-MIB;
     these may require table index suffixes; use snmpwalk first to discover the correct instance OIDs -->
# Verify the OIDs exist on your device first:
snmpwalk -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4

# Then query specific counters:
snmpget -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4.6   # cnfESPktsDropped
snmpget -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4.4   # cnfESPktsExported

If the device-exported rate exceeds the collector inbound rate, loss is happening in transit or at the collector. If the device drop counter is rising, loss is happening on the device itself. Comparing these two numbers gives you end-to-end loss visibility that Udp_RcvbufErrors alone cannot provide.

sFlow: sampling accuracy degrades at low volumes

sFlow accuracy depends on absolute sample count, not total packet volume. At 0.25% sampling (1-in-400), hourly estimates for rare flows can swing plus or minus 65%. Monthly aggregation at the same rate yields plus or minus 2.4%. Cloudflare recommends not exceeding a 1-in-5000 sampling rate before noticeable accuracy loss on typical network volumes.

This means sFlow systematically underrepresents bursty or low-volume traffic. A flow generating 100 packets per second may yield zero samples in a given interval, or three, purely due to Poisson variance. This makes sFlow unsuitable for forensic investigation of short-duration events or detection of low-volume exfiltration. Operators relying on sFlow for security alerting should understand this architectural blind spot.

sFlow: sampling rate not normalized

sFlow reports raw sampled counts. If the analytics layer does not multiply by the sampling rate to recover true byte and packet counts, bandwidth charts are wrong by the sampling factor. At a 1-in-1000 sampling rate, charts report one-thousandth of actual traffic. The error is consistent across all conversations, so relative comparisons look right while absolute numbers are off by orders of magnitude.

On high-speed links with low aggregate traffic, static sampling rates produce extremely sparse samples. A 10 Gbps interface carrying less than 500 Mbps with a static 1-in-10000 rate may yield so few samples that the data is unusable for meaningful analysis. Misconfigured sFlow on 40G inter-switch links routinely produces data that operators mistake for a traffic drop when the real problem is sampling sparsity.

sFlow vs NetFlow: different collector load profiles

sFlow and NetFlow have very different packet-rate characteristics at equal bandwidth. sFlow generates samples proportional to packet rate, not bandwidth. NetFlow v9 exports are bursty: they concentrate on flow creation and teardown events. Buffer sizing and collector capacity planning must be protocol-specific. A collector sized for NetFlow v9 may be overwhelmed by sFlow at the same link speed, because sFlow export volume scales with packet count rather than flow count.

Where these failures show up in production

The table below maps each protocol to its dominant failure mode, what it looks like on a dashboard, and the first signal to check.

ProtocolDominant failureWhat it looks likeFirst thing to check
NetFlow v9 / IPFIXTemplate desync after rebootCollector receives datagrams but decodes zero recordsCollector logs for “template” errors; exporter uptime
NetFlow v9 / IPFIXExporter-side dropsDevice exports fewer records than traffic suggestscnfESPktsDropped via SNMP at .1.3.6.1.4.1.9.9.387.1.4.6
sFlowSampling rate not normalizedBandwidth charts off by sampling factor (e.g., 1000x low)Sampling rate config vs analytics scaling factor
sFlowLow-volume flow blindnessKnown traffic absent from flow dataCompare sample count to expected rate for that flow
All threeUDP socket buffer overflowCharts show declining traffic during an actual traffic spikeUdp_RcvbufErrors in /proc/net/snmp

The shared UDP failure is the most common and the most damaging. The Linux default net.core.rmem_max of 4,194,304 bytes (4 MB) is inadequate for high-pps sFlow collectors. Production deployments should target 16 MB or higher, with 33 MB for very high-volume collectors. The socket buffer must also be explicitly set on the listener socket via SO_RCVBUF.

A collector that looks healthy (process running, port listening, disk not full) can still be losing a significant fraction of incoming flow datagrams if the socket buffer is undersized. The “packets received” counter at the application layer increments normally for packets that make it through, while dropped packets increment a separate kernel counter that most teams never monitor.

Signals to watch in production

SignalWhy it mattersWarning sign
Udp_RcvbufErrors in /proc/net/snmpDatagrams arriving at kernel but dropped before application reads themAny nonzero increment on a flow collector
Collector inbound rate vs device exported rateEnd-to-end loss detection across the UDP pathCollector receiving fewer records than device exported
Flow template cache state (NetFlow v9/IPFIX)Templates must be cached to decode data recordsDecoded records = 0 with received packets > 0
sFlow sampling rate consistencyRaw counts must be scaled by sampling rate for accurate analyticsBandwidth values inconsistent with SNMP interface counters
Exporter drop counters (Cisco cnfESPktsDropped)Device-side loss invisible to collector-side monitoringCounter incrementing during traffic spikes
NIC RX drops on collector (/proc/net/dev)Packets lost at hardware level before reaching socket layerRising RX drops on the flow-ingress NIC
Flow export-to-ingest latency (NetFlow v9/IPFIX)Indicates device backlog, network delay, or collector ingestion lagSustained latency greater than 30 seconds

For deeper coverage of specific failure modes, see the related guides on UDP flow loss, template desync, sFlow sampling rate, and export-to-ingest latency.

How Netdata helps

Netdata correlates flow telemetry signals across the collection stack, which is where most silent failures live.

  • UDP socket buffer drops are monitored at the kernel level (Udp_RcvbufErrors), with per-collector visibility so you can see which collector is losing data.
  • NIC RX drops are tracked per interface, catching hardware-level loss before it reaches the socket layer.
  • Collector CPU utilization is monitored per core, which catches RSS misconfiguration where one core saturates at 100% while others sit idle. This is a common cause of flow data loss that aggregate CPU metrics hide.
  • TSDB write queue depth and disk space are tracked, catching the slow-consumer bottleneck that backs up the socket buffer and causes silent drops.
  • Cross-signal correlation lets you compare SNMP interface counters (which reflect real traffic) against flow-derived analytics (which may be missing data). When SNMP shows rising traffic but flow charts show decline, the gap points directly to collector-side loss.
  • Exporter-side counters for Cisco devices (cnfESPktsDropped, cnfESPktsExported) can be polled alongside collector-side metrics, giving you end-to-end loss visibility without relying on a single signal.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.