The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-cloud-onprem-flow-correlation

Operations Guides

Correlating cloud VPC flow logs with on-prem NetFlow

Cloud flow logs and on-prem flow records share the 5-tuple concept but diverge in nearly every dimension that matters for correlation: transport, latency, sampling, timestamps, topology, and NAT visibility. Cloud providers emit VPC flow logs via push to object storage with implicit sampling and aggregation intervals measured in minutes. On-premises devices export NetFlow v5/v9, IPFIX, or sFlow over UDP with configurable sampling and near-real-time delivery.

This gap is a recurring contributor to operational incidents. An attacker pivoting from a compromised cloud workload to on-prem via VPN is invisible across the boundary if no join exists between the two telemetry sources. The same gap hides legitimate operational issues: cross-boundary packet loss, asymmetric routing through cloud transit gateways, and NAT translation mismatches between cloud NAT and on-prem firewalls.

This reference covers the format incompatibilities, sampling semantics, timestamp behavior, and NAT opacity that make cross-environment flow correlation operationally difficult. It is aimed at operators building or maintaining a normalization layer across these two domains.

Why cloud and on-prem flow data resist correlation

Cloud flow logs are not flow records in the NetFlow/IPFIX sense. They are aggregated, sampled, and delivered asynchronously. There is no SNMP, no CDP, no LLDP. The flow record is the only native signal. Topology is the cloud provider’s graph. Delivery is push-based to object storage. Lag is minutes, not seconds. Sampling is implicit.

DimensionOn-prem NetFlow/IPFIXCloud flow logs
TransportUDP push to collectorPush to object storage, polled by consumer
LatencyNear real-time (seconds)Minutes (aggregation plus delivery)
SamplingExplicit, configurable per exporterImplicit, provider-controlled
TimestampsExporter clock (NTP-dependent)Provider-aggregated start and end epochs
TopologyCDP/LLDP/FDB/ARP availableCloud graph only
NAT visibilityPost-NAT at perimeterPre-NAT inside VPC (if pkt-* fields enabled)
Template or cacheNetFlow v9/IPFIX template cacheNo template concept
flowchart LR
    CL["Cloud flow logs
AWS / GCP / Azure
push to object storage
minutes of lag"] -->|"timestamp skew
up to 60s (AWS)"| NORM["Normalization layer
windowed 5-tuple join
NAT translation logs
direction inference"] OP["On-prem NetFlow / IPFIX
UDP export
near real-time"] -->|"NTP-dependent
sampling-aware"| NORM NORM --> CORR["Cross-boundary
correlation
same conversation
different vantage points"]

Provider reference: cloud flow log semantics

Each cloud provider’s flow log format has distinct semantics that affect how records can be joined with on-prem data.

AWS VPC Flow Logs

AWS VPC Flow Logs aggregate captured packets into intervals. The default aggregation interval is 10 minutes, reducible to 1 minute. Format versions 2 through 11 exist, each adding fields without removing prior ones.

Core 5-tuple fields: srcaddr, dstaddr, srcport, dstport, protocol (IANA number). Additional fields include packets, bytes, start and end (Unix epoch seconds), action (ACCEPT or REJECT), and log-status.

For correlation through NAT gateways or EKS pods, pkt-srcaddr and pkt-dstaddr are essential. Without them, srcaddr and dstaddr reflect the translated IP, not the original. EKS pods have separate pod IPs from node ENI IPs. pkt-srcaddr exposes the pod IP while srcaddr shows the node ENI IP.

The log-status field distinguishes data gaps:

  • SKIPDATA: records were dropped internally by AWS due to capacity constraints. One SKIPDATA record can represent multiple uncaptured flows.
  • NODATA: no traffic on that ENI during the interval. Not a gap, but operators frequently confuse it with SKIPDATA.

The tcp-flags field is a bitmask aggregated across the entire aggregation interval: FIN=1, SYN=2, RST=4, SYN-ACK=18. ACK and PSH flags are not included in the bitmask; their values stay 0 unless paired with a supported flag. A short-lived connection that opens and closes within a single aggregation interval may appear as a single record with combined flags (SYN+FIN = 3, SYN-ACK+FIN = 19). This makes TCP state machine reconstruction unreliable compared to on-prem NetFlow v5, which records a single OR’d flags value per flow representing the union of all flags seen during that flow’s lifetime. For packet-level TCP handshake analysis, neither cloud nor NetFlow records are a substitute for a packet capture.

The flow-direction field (added in version 5) resolves initiator ambiguity for AWS-side flows.

GCP VPC Flow Logs

GCP uses dual-stage sampling. The primary sampling rate is opaque and dynamic, varying with host load. The secondary rate is configurable from 0.0 to 1.0. Primary sampling is uncontrollable by the operator.

Aggregation intervals are configurable: 5 seconds (default), 30 seconds, 1 minute, 5 minutes, 10 minutes, or 15 minutes.

GCP flow logs do not indicate which endpoint initiated a flow. They identify packet direction relative to the interface only. This complicates correlation with on-prem NetFlow, which also lacks initiator context natively.

A critical asymmetry exists in firewall interaction:

  • Egress packets are sampled before egress firewall rules are evaluated. Denied packets can still appear in logs.
  • Ingress packets are sampled after ingress firewall rules are evaluated. Dropped packets are not logged.

This means GCP flow logs overestimate allowed egress and underestimate blocked ingress compared to on-prem firewalls that log both directions.

Azure NSG Flow Logs and VNet Flow Logs

Azure NSG Flow Logs version 2 introduces flow state tracking with B (Begin), C (Continue), and E (End) states, plus bidirectional byte and packet counters.

Byte and packet counts are not recorded for flows affected by non-default inbound rules. Reported totals will be lower than actual traffic for those flows. This affects chargeback correlation and volumetric validation.

Azure NSG Flow Logs are officially retiring on 30 September 2027. The successor is VNet Flow Logs, which operates at the virtual network level rather than per-NSG and captures platform-rule traffic that NSG flow logs miss. Teams staying on NSG flow logs should use version 2 only. Version 1 lacks byte and packet counters entirely.

Timestamp skew: the biggest correlation killer

Timestamp alignment is the single hardest problem in cross-environment flow correlation. Each source introduces skew differently.

AWS VPC Flow Logs: the start and end timestamps can be up to 60 seconds off from actual packet receipt or transmission. AWS documentation states these values might be either up to 60 seconds before the packet was received on the network interface or up to 60 seconds after.

GCP VPC Flow Logs: the primary sampling stage adds interpolation layers. Missed packets are compensated by interpolation from captured packets, which introduces additional timestamp uncertainty.

On-prem NetFlow/IPFIX: timestamps depend on exporter clock accuracy. Juniper has documented IPFIX timestamp inaccuracy on certain platforms where timestamps diverge from system time despite apparent NTP sync.

NetFlow v5: flow records do not carry absolute timestamps. Flow timing is derived from system uptime offsets (the FIRST and LAST fields relative to export uptime), which means accuracy depends on the exporter’s clock and uptime counter. NetFlow v9, IPFIX, and sFlow include timestamps that can be used directly to compute export-to-ingest latency.

The operational rule: always use a windowed join, never an exact timestamp match. A 2-minute tolerance is conservative for most paths. Five minutes is safer for high-latency cross-VPN paths where aggregation and delivery lag compound. Note that NTP drift between collectors is a separate problem from flow-log aggregation lag. Even sub-second NTP offset between a cloud flow log epoch and an on-prem NetFlow exporter clock will not matter if the join window is set to minutes, but it will compound with aggregation lag on paths where every second counts.

Sampling and completeness gaps

Cloud flow logs introduce data gaps that have no direct equivalent in on-prem flow collection.

AWS SKIPDATA: one SKIPDATA record can represent multiple uncaptured flows. The on-prem side shows traffic. The cloud side shows nothing. Treat any SKIPDATA record during a known traffic window as a data-quality finding.

GCP secondary sampling: when set below 1.0, flow entries are discarded randomly. The default secondary sampling rate is 0.5 (50%) for Compute Engine API-created configs; configs created via the Network Management API default to 1.0 (100%). Correlation against on-prem NetFlow will show missing entries proportional to the secondary sample rate. At 0.5 secondary sampling, expect roughly half the cloud-side flows to be absent.

Cloud delivery lag: cloud flow logs arrive via object storage poll, not real-time UDP push. The cloud side of a correlated view is always delayed relative to the on-prem side. Real-time alerting on cross-boundary patterns is not feasible with cloud flow logs alone. Retrospective correlation is the realistic use case.

NAT and identity correlation

NAT boundaries break 5-tuple joins. The endpoint IP inside the flow record is the NAT’d address. Security teams investigate the wrong host. Topology inference places the endpoint at the NAT device port, not the actual endpoint.

Inside the VPC: cloud flow logs may show pre-NAT IPs. AWS pkt-srcaddr and pkt-dstaddr expose the original pod or instance IP before NAT gateway translation. Without these fields, only the translated IP is visible.

At the perimeter: on-prem NetFlow typically reflects post-NAT IPs at the perimeter device. The cloud side may show pre-NAT IPs inside the VPC. The on-prem side shows the translated address as traffic exits the cloud.

To correlate across the NAT boundary, operators must normalize to whichever vantage point they are correlating from, and integrate NAT translation logs (cloud NAT logs, firewall session logs) into the enrichment pipeline. Without translation logs, identity recovery is impossible for flows that crossed the NAT boundary.

See locating endpoints behind NAT and wireless for the related problem of placing endpoints whose IP appears only behind a NAT device.

Direction ambiguity

Direction is ambiguous in both cloud and on-prem flow data, but in different ways.

GCP: no initiator direction at all. Flow logs identify packet direction relative to the interface, not which endpoint started the conversation.

AWS: the flow-direction field (version 5 and later) resolves this for AWS-side flows. Earlier versions lack it.

On-prem NetFlow: lacks initiator context natively. Flow direction is typically inferred from port assignment (low port = server) or from template metadata, not from packet inspection.

For cross-environment correlation, a flow initiated from cloud to on-prem appears as ingress on the on-prem side and egress on the cloud side. Without explicit direction fields, the join must rely on 5-tuple symmetry (same source and destination pair, ports swapped) rather than directional matching.

The normalization layer

Cross-domain correlation requires a normalization layer that translates cloud flow log formats and on-prem flow records into a common schema. Commercial platforms exist for this. For teams building their own normalization, the minimum requirements are:

  • Common timestamp field: normalize all timestamps to UTC epoch seconds. Apply a windowed join with tolerance appropriate to the path. Two to five minutes is the practical range for cross-cloud-to-on-prem paths.
  • Sampling-rate awareness: multiply sampled counts by the sampling rate for both cloud and on-prem sources. GCP’s opaque primary sampling rate means cloud-side byte and packet counts may be inherently unreliable for volumetric comparison against on-prem data.
  • NAT translation integration: join cloud NAT logs and on-prem firewall session logs into the enrichment pipeline to recover pre-NAT and post-NAT identity.
  • Direction normalization: infer conversation direction from 5-tuple symmetry rather than relying on provider-specific direction fields.
  • Completeness tracking: track SKIPDATA (AWS), sampling gaps (GCP), and template-cache misses (NetFlow v9/IPFIX) as data-quality signals, not just as absent records. A gap on one side with traffic on the other is itself a finding.

Signals to watch across the boundary

SignalWhy it mattersWarning sign
Timestamp offset between sourcesWindowed joins fail silently when skew exceeds toleranceFlows present on one side, absent on the other despite known traffic
Cloud log delivery lagPrevents real-time cross-boundary alertingCloud records arrive 5 to 10 minutes after on-prem records for the same conversation
SKIPDATA or NODATA frequencyIndicates cloud-side data loss that creates false gaps in correlationSudden increase in SKIPDATA records during a traffic spike
GCP secondary sampling rateBelow 1.0, random flow entries are discardedCorrelation shows missing cloud entries proportional to the configured rate
NAT translation log retentionWithout translation logs, pre-NAT identity is unrecoverableInvestigation window exceeds NAT log retention period
NTP offset on on-prem exportersDrift compounds with aggregation lag to shift flow records outside the join windowFlow records from different devices do not align for the same event
Azure non-terminating flow countsByte and packet totals are silently absent for affected flowsCloud-side volumetric totals consistently lower than on-prem for same path

How Netdata helps

Netdata can serve as the on-prem half of the correlation equation:

  • On-prem flow collection: Netdata collects NetFlow v5/v9, IPFIX, and sFlow data from network devices, providing the on-prem telemetry that cloud flow logs must be joined against.
  • Per-second metric resolution: Netdata’s collection frequency allows tight temporal correlation between on-prem flow data and contextual signals such as interface counters, BGP state, and syslog events.
  • Gap corroboration: when cloud flow logs show a gap, Netdata’s on-prem signals (interface utilization, error counters, discard counters) can confirm whether traffic actually flowed during the gap or whether the cloud-side absence reflects a real outage.
  • NTP monitoring: Netdata tracks NTP offset on collectors and can alert when clock skew exceeds thresholds that would compound with cloud-side aggregation lag.
  • UDP buffer health: for on-prem flow collectors, Netdata monitors Udp_RcvbufErrors and NIC RX drops, ensuring the on-prem half of the correlation is not silently losing data before it reaches storage.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.