The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-sflow-sampling-rate

Operations Guides

sFlow sampling rate: why your traffic totals are off by 1000x

Your sFlow-derived bandwidth charts show a 10G link carrying 12 Mbps. SNMP counters on the same interface show 8.4 Gbps. The switch is not broken and the collector is not dropping packets. The analytics pipeline is summing raw sampled bytes without multiplying by the sampling rate.

sFlow is not NetFlow. It does not maintain a flow cache on the device, aggregate bytes per conversation, and export summary totals. sFlow exports individual packet samples, one per datagram, each carrying the packet’s header data and metadata about the sampling process. The collector is responsible for turning those samples into traffic estimates through multiplication. When that multiplication is missing, every chart, alert, capacity plan, and billing report built on the data is wrong by the sampling factor.

On a typical 1:1000 sampling configuration, the numbers are off by 1000x. A link carrying 8 Gbps appears to carry 8 Mbps. A DDoS attack at 5 Gbps looks like ordinary background traffic. Capacity decisions get made on data that is three orders of magnitude below reality.

What this means

sFlow is a sample-datagram protocol, not a template-flow protocol. The distinction determines where the math happens.

In NetFlow v5, v9, or IPFIX, the device maintains a flow cache. It aggregates bytes and packets per 5-tuple conversation. When a flow expires or the cache fills, the device exports an accumulated record with the total byte and packet counts. The collector receives numbers that already represent observed traffic volume.

sFlow works differently. The device samples packets at a configured rate (1 in N), wraps each sampled packet’s header and metadata into a UDP datagram, and sends it to the collector. There is no flow cache on the agent. Each datagram contains:

  • frame_length: the original packet length on the wire. Whether this includes the FCS depends on the implementation and platform.
  • sampling_rate: the integer N used for this sample (a value of 1000 means 1 in 1000 packets was sampled)
  • Additional metadata: interface index, source and destination information, header protocol data

The collector receives raw samples. To estimate actual traffic volume, it must multiply: estimated_bytes = SUM(frame_length) * sampling_rate.

Without this multiplication, you are reporting the volume of sampled packets as if it were the total. On a 1:1000 link, you report 0.1% of reality.

flowchart TD
    A["Device samples
1 in N packets"] --> B["sFlow datagram
frame_length + sampling_rate"] B --> C["UDP to collector
port 6343"] C --> D["Collector receives
raw samples"] D --> E{"Multiply by
sampling_rate?"} E -->|Yes| F["Correct traffic
estimate"] E -->|No| G["Off by factor N
often 1000x"]

A critical subtlety: in sFlow v5, sampling_rate is embedded per-sample. It can change between samples from the same device. Some platforms adjust the rate dynamically based on interface speed or traffic load. Collectors that hardcode a single rate will produce wrong totals whenever the device adjusts.

Common causes

CauseWhat it looks likeFirst thing to check
Normalization skipped entirelyAll traffic uniformly low by the same factor across all exportersWhether the collector or analytics layer applies sampling_rate multiplication
Wrong base metric usedConsistent undercount, roughly 14 bytes per packet below real trafficVerify the pipeline uses frame_length, not IP payload length
Variable rate not handled per-sampleSome exporters correct, others off by varying factors over timeInspect whether sampling_rate changes between samples from the same source
Hardware sample dropsOff by a variable factor despite correct nominal rateCheck the drops field in flow_sample structures and compare sample_pool against received count
UDP transport lossIntermittent undercounting, worse during traffic burstsCheck Udp_RcvbufErrors and compare collector inbound rate against device export counters

Quick checks

These commands are read-only. None modify device or collector state.

# Verify sFlow datagrams are arriving at the collector
tcpdump -i eth0 -nn 'udp port 6343' -c 1000

# Check for UDP socket buffer drops (silent sample loss)
cat /proc/net/snmp | grep '^Udp:'

# Check current UDP receive buffer sizing on the sFlow listener
ss -lun '( sport = :6343 )' -m

# Compare SNMP-derived interface utilization against sFlow-derived values
# ifHCInOctets gives ground truth for bytes on the wire
snmpwalk -v2c -c <community> <device> .1.3.6.1.2.1.31.1.1.1.6

# Check device-side flow export statistics (Cisco)
# TODO: verify the correct OID subtree for Cisco sFlow export stats, as this may require a table index
snmpwalk -v2c -c <community> <device> .1.3.6.1.4.1.9.9.387.1.4

How to diagnose it

  1. Cross-validate against SNMP counters. Poll ifHCInOctets and ifHCOutOctets on the same interface the sFlow data covers. Compute utilization from SNMP. If SNMP shows 8 Gbps and sFlow analytics show 8 Mbps, the ratio tells you the sampling factor immediately.

  2. Calculate the expected factor. Divide SNMP-derived bytes by sFlow-derived bytes for the same time window and interface. If the ratio is close to the configured sampling rate (1000, 4096, 8192, etc.), normalization is missing. If the ratio is close but not exact, investigate hardware sample drops or UDP loss.

  3. Inspect raw sFlow samples. Decode a stream of sFlow datagrams and check two fields per sample: frame_length and sampling_rate. Confirm that sampling_rate matches what you expect. Confirm that frame_length reflects full wire-layer length, not just IP payload.

  4. Check for per-sample rate variation. If the device uses adaptive sampling, different samples may carry different sampling_rate values. A collector that applies a single hardcoded rate will miscalculate for any sample whose actual rate differs.

  5. Verify the normalization formula with a worked example. 50 sFlow samples of 1500-byte packets at a 1:64 sampling rate over 60 seconds should yield 50 * 1500 * 64 * 8 / 60 = 640,000 bps. If your analytics report 50 * 1500 * 8 / 60 = 10,000 bps, the sampling rate multiplication is missing.

  6. Check for hardware sample loss. Compare the sample_pool field (total packets eligible for sampling) against the number of samples actually received. A large gap means the device is generating fewer samples than its nominal rate implies, often due to hardware rate limiting. The drops field in flow_sample tracks packets the sampling mechanism could not process.

  7. Check for UDP transport loss. Compare the device’s export counter against the collector’s receive counter. If the device exported 50,000 datagrams and the collector received 35,000, the 15,000 gap is silent loss that compounds the sampling error. Check Udp_RcvbufErrors in /proc/net/snmp for kernel-level confirmation.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Flow sampling rate consistencyDetects when the rate reported by the exporter differs from what the collector appliesAny mismatch between reported and applied rate
sFlow-derived vs SNMP-derived utilizationGround truth comparison; SNMP counters reflect actual wire trafficsFlow total below 50% of SNMP total for the same interface and window
UDP socket buffer drops (Udp_RcvbufErrors)Silent sample loss at the kernel level; compounds sampling errorAny nonzero increment on a flow collector
Device-side flow exporter dropsPackets generated but not exported by the deviceDrop rate exceeding 0.1% of total exported
Per-exporter sampling rate over timeDetects adaptive rate changes that collectors must handleRate changing without a corresponding change ticket

Fixes

Apply normalization in the ingestion pipeline

The fix depends on where the gap is. If the collector library does not apply sampling rate multiplication, add it as a post-ingestion step: for each sample, multiply frame_length by sampling_rate before aggregating. For a constant rate, the formula is estimated_bytes = SUM(frame_length) * sampling_rate. For per-sample rates, use estimated_bytes = SUM(frame_length * sampling_rate).

Use frame_length, not IP payload length

Some pipelines use the Layer 3 total length field from the decoded packet header as the base for multiplication. This omits Layer 2 overhead: Ethernet header (14 bytes), VLAN tags (4 bytes each), and FCS (4 bytes). The sFlow v5 spec defines frame_length as the original packet length on the wire. Use that field as the base for all byte calculations.

Handle variable sampling rates per-sample

Do not assume a single sampling rate per exporter. Read the sampling_rate field from each flow_sample structure and apply it individually. If your collector or analytics platform only supports a global rate, it will produce wrong totals whenever a device adjusts its sampling rate. This is particularly important on platforms that use adaptive sampling keyed to interface speed.

Tune UDP buffers if loss is compounding

If the normalization is correct but totals are still low, UDP loss may be dropping samples before they reach the analytics layer. The Linux default net.core.rmem_max is often under 1 MB on many distributions and is inadequate for high-pps sFlow collectors.

Production deployments should target 16 MB or higher, with explicit SO_RCVBUF set on the listener socket.

# WARNING: These commands change system state and affect ALL UDP sockets on the host.
# Verify no other UDP consumers will be impacted before applying.
# These changes are non-persistent and will be lost on reboot without sysctl.conf or systemd-sysctl entries.
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.rmem_default=16777216

Prevention

Always cross-validate sFlow against SNMP. Build a periodic check that compares sFlow-derived throughput against ifHCInOctets for critical interfaces. If the ratio drifts from the expected sampling factor, something has changed: the sampling rate, the collector normalization, or the UDP transport.

Track sampling rate changes over time. Log the sampling_rate from each exporter and alert on changes. If a device silently switches from 1:1000 to 1:4096, every downstream chart will shift by a factor of 4 without any error message.

Monitor hardware sample drops. The drops field in flow_sample and the ratio of sample_pool to received samples reveal when the device cannot keep up with its configured rate. This produces a systematic undercount that multiplication alone cannot fix.

Validate vendor defaults. Default sampling rates vary by platform and vendor, with common values ranging from 1:1000 to 1:16384 or higher. A device inherited from another team may have a sampling rate that differs from what your analytics assume. Always verify the configured rate on the device, not just in the collector.

How Netdata helps

Netdata collects both SNMP interface metrics (ifHCInOctets, ifHCOutOctets) and system-level UDP statistics (Udp_RcvbufErrors) alongside flow data, enabling direct cross-validation between sFlow-derived and SNMP-derived throughput on the same dashboard. Per-core CPU metrics and softirq percentages help identify RSS misconfiguration that funnels sFlow processing to a single core, causing buffer drops that silently reduce sample counts. UDP socket buffer drop trends correlate with flow data quality degradation, providing an early warning before analytics drift becomes visible in charts.

The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.