The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-vendor-api-latency-pagination

Operations Guides

Vendor API latency and pagination: monitoring pull-mode collection

For SD-WAN controllers, cloud-managed networking, and modern firewall platforms, vendor API pull-mode collection is now the primary telemetry path. Operators depend on HTTPS calls to Meraki, Cato, PAN-OS, RESTCONF, and gRPC endpoints to learn tunnel state, license validity, session counts, and topology. Each API has its own authentication model, rate-limit budget, and pagination semantics. A collector that ignores these constraints will silently lose data, get throttled, or report healthy when the payload is empty.

The dominant failure mode is not the API going down. It is the collector treating HTTP 200 as success without validating the payload, exhausting a shared rate-limit budget, or stalling on paginated responses that never complete. These are silent failures: dashboards flat-line, no error is logged at INFO, and the first signal is a user complaint hours later.

What it is and why it matters

Pull-mode collection means the monitoring platform initiates a request to the vendor API and waits for a response. This is distinct from push-mode telemetry such as streaming gNMI, YANG Push (RFC 8641), or syslog traps, where the device or controller sends data unsolicited.

For API-first and controller-managed platforms, pull-mode is often the only option. Meraki, Aruba Central, Juniper Mist, and Cisco DNA/Catalyst Center expose most operational state through their APIs, not through SNMP. SD-WAN overlays like Cato, Viptela, Versa, and Fortinet make tunnel state, SLA metrics, and application-aware routing decisions available via the orchestrator API. For these platforms, the API is the telemetry source.

The shift from SNMP to vendor APIs changes the failure surface. SNMP is UDP-based with no session overhead; a request and response are a single round-trip. Vendor APIs run over HTTPS, incurring TCP connection setup and TLS handshake cost on every request unless the collector pools connections. Authentication adds another layer: API keys expire, tokens refresh on schedules, and SAML/SSO administrators on some platforms (Meraki) cannot generate API keys at all.

How it works

A pull-mode collection cycle has four stages, each with its own failure mode.

flowchart LR
  S[Scheduler] --> A[Auth / token refresh]
  A --> R[API request]
  R --> RL{Rate limit remaining?}
  RL -->|Exhausted| T[HTTP 429 + Retry-After]
  RL -->|OK| P{Paginated response?}
  P -->|Yes| PG[Fetch next page via cursor or offset]
  P -->|No| V[Validate payload]
  PG --> V
  V --> E{Schema valid?}
  E -->|No| G[Silent gap: HTTP 200 + empty or error body]
  E -->|Yes| W[Write to TSDB]

Stage 1: Authentication. The collector presents credentials to the vendor API. PAN-OS uses an API key requested via /api/?type=keygen&user=<user>&password=<password>, then passed as an X-PAN-KEY header or key= query parameter on subsequent calls. Meraki uses a bearer token in the Authorization header. Cato uses an x-api-key header against a GraphQL endpoint at https://api.catonetworks.com/api/v1/graphql2. Key expiration and rotation are the leading cause of silent collection failure: a key that worked yesterday returns 401 today (or, for Meraki, 404 by design to avoid leaking resource existence).

# Check PAN-OS API key validity and response payload
# WARNING: the key appears in the URL and will be visible in shell history
# and process listings. Use this only for debugging, not in production scripts.
curl -sk "https://<fw>/api/?type=op&cmd=<show><system><info></info></system></show>&key=<apikey>"

# Check Meraki API reachability and HTTP status separately
curl -s -o /dev/null -w "%{http_code}\n" -H "Authorization: Bearer $KEY" \
  https://api.meraki.com/api/v1/organizations

# Isolate DNS resolution time from total API latency
curl -s -o /dev/null -w "dns:%{time_namelookup} total:%{time_total}\n" \
  https://api.meraki.com/api/v1/organizations

Stage 2: Rate-limit budget. Every vendor imposes per-token, per-minute, or per-query limits. Meraki allows 10 requests per second per organization, with a burst of 30 in 2 seconds, and 100 per second per source IP. Cato enforces per-query, per-account limits: 120/min general, but accountSnapshot at 1/sec, accountMetrics at 15/min, and eventsFeed at 100/min. Multiple collectors sharing the same API key share the same budget. When the budget is exhausted, the API returns HTTP 429, and collection stops until the window resets. Meraki returns a Retry-After header. Cato does not formally publish response header names for rate limiting, so verify empirically.

Stage 3: Pagination. Large responses (FDB tables, event feeds, license inventories) are delivered across multiple pages. The IETF is standardizing pagination for RESTCONF and NETCONF via two active drafts: draft-ietf-netconf-list-pagination and draft-ietf-netconf-list-pagination-rc. These define query parameters including:

ParameterPurpose
limitMaximum entries returned per response
offsetNumber of entries to skip
cursorOpaque position marker for stateless resume
directionforwards (default) or backwards
sort-byNode identifier for ascending sort
whereXPath 1.0 filter expression

The defined processing order is: where, then sort-by, then direction, then (offset or cursor), then limit. Offset and cursor are mutually exclusive per the draft specification. Cursor values are opaque and ephemeral; collectors must not cache or reuse them across requests. Before these drafts existed, vendors implemented proprietary parameters (such as ?page=1&size=50), and operators had to discover behavior per vendor empirically.

The drafts explicitly warn that retrieving all entries without pagination “can lead to inefficiencies (e.g., long loading time, memory overconsuming, or crash) in the server, the client, and the network in between.” A paginated response that stalls partway through (due to a timeout, cursor expiry, or rate-limit hit mid-walk) produces a partial dataset, and the collector may not detect the incompleteness.

Stage 4: Response validation. The collector must verify not just the HTTP status code but the payload contents. PAN-OS returns <response status="error"> inside an HTTP 200. A collector that checks only the HTTP status will treat this as success and record the error payload as valid telemetry. Schema validation catches vendor-side API changes that break the adapter without changing the HTTP status.

Where it shows up in production

Five deployment variants depend heavily on vendor API pull-mode collection:

  • API-first / controller-managed (Meraki, Aruba Central, Juniper Mist, Cisco DNA/Catalyst Center). Many operational facts live in the controller API, not on the device. Topology is often controller-asserted rather than inferred from CDP/LLDP. SNMP may be entirely absent.
  • SD-WAN overlay (Cato, Viptela, Versa, Fortinet). The tunnel is the unit of interest, not the physical interface. Tunnel up/down state, SLA probe latency, jitter, loss, and application-aware routing decisions all come from the orchestrator API. Underlay versus overlay confusion is the most common misdiagnosis.
  • Cloud-native (AWS VPC Flow Logs, Azure NSG Flow Logs, GCP VPC Flow Logs). No SNMP, no CDP, no LLDP. Flow records arrive via object storage poll, not UDP. Delivery lag is measured in minutes, not seconds. Sampling is implicit in the cloud provider’s collection mechanism.
  • Modern firewalls (PAN-OS, FortiGate). Session counts, NAT translation tables, license state, and threat logs are available via XML or JSON-RPC APIs. These typically supplement, not replace, SNMP for interface counters.
  • Hybrid / multi-vendor. Most enterprises. Requires a normalization layer to translate vendor-specific API schemas into a common data model, which is the highest operational complexity for topology inference and flow normalization.

Tradeoffs and when to use it

Pull-mode vendor APIs solve real problems that SNMP cannot: structured, model-driven payloads, controller-asserted topology, and access to state that exists only in the management cloud. But they introduce failure modes that SNMP does not have.

Connection overhead. Each RESTCONF or vendor HTTPS request incurs TCP and TLS session cost absent from SNMP’s UDP model. Connection pooling is essential for high-volume polling. Without it, per-request handshake overhead dominates latency measurements and reduces effective throughput.

Rate-limit contention. SNMP has no equivalent of a shared per-token quota. When multiple tools (NMS, security scanner, automation script, ad-hoc dashboard) share one Meraki or Cato API key, a single aggressive consumer can exhaust the budget and blind every other consumer. Track consumption against the published limits and watch for Retry-After headers before throttling occurs.

The 200-with-empty-payload trap. This is the most operationally damaging failure mode. The API is up, the network path is fine, ICMP to the vendor cloud is healthy, but the payload is empty or schema-mismatched. This happens during vendor maintenance windows, after an API schema change without a version bump, or after a vendor-side incident. Collectors that validate only HTTP status record silence as success. The signal that catches it is tracking data freshness (time since last valid payload) independently of HTTP status.

Streaming as an alternative. Where available, streaming telemetry (gNMI Subscribe RPC in STREAM mode, YANG Push per RFC 8641) removes the polling bottleneck entirely. gNMI also supports a POLL mode for collector-initiated snapshots. But streaming coverage is uneven across vendors and device generations; older devices require polling regardless. The choice between pull and push is often dictated by what the platform supports, not by operator preference.

Retry and backoff discipline. When a vendor API returns 429 or 5xx, the collector must back off. The standard pattern is exponential backoff with jitter: start at 1 to 2 seconds, double each attempt, add random jitter, and cap retries at 3 to 5. Without jitter, synchronized retry storms from multiple collectors can amplify load on the vendor endpoint.

Signals to watch in production

SignalWhy it mattersWarning sign
API request latency (p99)Rising latency indicates vendor-side backpressure or network path degradation. Latency approaching the configured timeout is a data-loss risk.p99 greater than 5x rolling baseline sustained
HTTP 429 rateAny sustained 429 means the collector is over-consuming or sharing a key with another aggressive consumer.Greater than 0 sustained, or greater than 1% of calls over 15 min
Rate-limit remainingProactive headroom tracking prevents cliff-edge throttling. Track consumption against published per-vendor limits.Less than 20% of quota per window
HTTP 401/403 (or Meraki 404)Authentication failure. API key rotated, revoked, or expired. Any nonzero value is abnormal and security-relevant.Any occurrence
Response payload validityHTTP 200 with empty or error payload is a silent gap. PAN-OS <response status="error"> inside HTTP 200 is the canonical case.Payload schema mismatch, empty data, or error body inside 200
API-sourced data freshnessTime since last successful valid payload. Flat-lines during silent gaps even when HTTP status is 200.Stale beyond 2x configured poll interval
ICMP to vendor cloud endpointDistinguishes vendor-side outage from network path issue. API down with ICMP healthy points to vendor cloud problem.Packet loss or elevated RTT to the API endpoint
Collector-side DNS resolution timeDNS resolution is included in total API latency. Isolate with time_namelookup to separate DNS from API processing.Resolution time trending upward

How Netdata helps

Netdata instruments per-API-call latency, HTTP status distribution, and response validation as first-class metrics, and correlates them alongside SNMP, flow, and syslog signals on a unified timeline:

  • Correlate vendor API latency spikes with SNMP timeout rates and device control-plane CPU on the same dashboard to distinguish collector-side issues from device-side problems.
  • Track HTTP 429 counts and rate-limit consumption trends alongside API-sourced data freshness to catch throttling before it causes a silent data gap.
  • Alert on payload validation failures (HTTP 200 with error or empty body) as a signal distinct from transport-level errors, so the 200-with-empty-payload trap does not go unnoticed.
  • Monitor collector CPU, DNS resolution time, and connection pool behavior to detect collector-side bottlenecks in the pull pipeline.
  • Combine vendor API health with ICMP reachability to the vendor cloud endpoint to separate local network path issues from vendor-side outages.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.