The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-vendor-api-silent-gap

Operations Guides

Vendor API silent data gap: HTTP 200 with an empty payload

Your SD-WAN controller dashboard shows flat lines. The Meraki organization API has not updated in twenty minutes. The PAN-OS firewall telemetry stopped at 03:00. Your collector logs show zero errors, every request returned HTTP 200, and no 5xx or timeout appears anywhere. But the data is gone.

The API endpoint is reachable, the TCP connection succeeds, the HTTP status code says OK, and the response body is empty, null, or contains an error wrapped inside a success envelope. Your collector accepted the response as valid because it checked the status code and nothing else. Many API adapters treat a 200 with an empty payload as “no data to report” rather than “the API is broken.” Charts go flat, but no error fires. If the API is your only telemetry source for an SD-WAN overlay or a cloud-managed firewall estate, you are blind without knowing it.

What this means

HTTP 200 signals a successful HTTP transaction. It places no obligation on the server to include a meaningful body. A vendor API that returns an empty JSON object, a null data field, or an error flag buried inside a 200 envelope is conformant to the HTTP specification. The problem is that collectors and monitoring adapters that rely solely on the HTTP status code cannot distinguish between “success with data” and “success without data.”

This failure mode affects vendor northbound APIs: RESTCONF, NETCONF, gNMI, gRPC-based telemetry streams, controller REST APIs (Cisco Catalyst Center, Meraki Dashboard, Cato GraphQL), and firewall XML APIs (PAN-OS). SNMP and ICMP do not have this problem because they either return data or timeout. The gap exists only where an application-layer protocol carries its own success and failure semantics inside an HTTP envelope.

PAN-OS is the canonical example. The XML API returns HTTP 200 for virtually every request, including those that fail. A response body containing <response status="error"> with an error message is delivered inside a 200 envelope. An adapter that checks only the HTTP status code will see 200 and treat the response as successful. Meraki and Cato APIs return structured JSON where the error is a field inside the body, not a separate status code. For gRPC-based telemetry (Juniper JTI, gNMI), the HTTP/2 frame always carries :status=200; the actual call outcome lives in the grpc-status trailer, which many load balancers and access logs do not inspect.

flowchart TD
    A["API returns 200"] --> B{"Body present?"}
    B -- "empty/null" --> C["Silent empty gap"]
    B -- "has content" --> D{"Error marker in body?"}
    D -- "yes" --> E["Semantic error"]
    D -- "no" --> F{"Schema matches?"}
    F -- "no" --> G["Schema drift"]
    F -- "yes" --> H["Healthy"]
    C --> I{"Auth valid?"}
    I -- "401/403/404" --> J["Key expired or revoked"]
    I -- "valid" --> K["Vendor backend fault"]
    E --> K
    G --> L["API version changed"]

Common causes

CauseWhat it looks likeFirst thing to check
API key expired or revoked200 with empty body, or intermittent 401/403 that retry logic swallowedManually curl with the current key and inspect the full response
Vendor API schema change200 with data, but parser fails or produces null fieldsCompare response structure to the vendor API changelog
Vendor maintenance or backend fault200 with empty payload across multiple endpoints simultaneouslyCheck the vendor status page
Rate limiting (HTTP 429)Intermittent empty responses during high-poll periodsInspect response headers for Retry-After
Collector adapter bug200 with valid data, but adapter mishandles the payloadEnable debug logging on the adapter and inspect raw response

Quick checks

# PAN-OS: inspect the full response body, not just the status code.
# Note: key in URL is visible in shell history and process list; use X-PAN-KEY header in production.
curl -sk "https://<fw>/api/?type=op&cmd=<show><system><info></info></system></show>&key=<apikey>" | head -20

# Meraki: check organization listing and HTTP status separately
curl -s -H "X-Cisco-Meraki-API-Key: $KEY" https://api.meraki.com/api/v1/organizations | python3 -m json.tool
curl -s -o /dev/null -w "%{http_code}\n" -H "X-Cisco-Meraki-API-Key: $KEY" https://api.meraki.com/api/v1/organizations

# Cato: GraphQL query with full response inspection
curl -s -H "x-api-key: $KEY" -H "Content-Type: application/json" \
  -d '{"query":"{ accountSnapshot(accountID: \"<id>\") { sites { name connectivityStatus } } }"}' \
  https://api.catonetworks.com/api/v1/graphql2 | python3 -m json.tool

# Check rate-limit headers on Meraki
curl -sI -H "X-Cisco-Meraki-API-Key: $KEY" https://api.meraki.com/api/v1/organizations | grep -i 'ratelimit\|retry'

# Check rate-limit headers on Cato (Cato does not formally publish these; verify empirically)
curl -sI -H "x-api-key: $KEY" https://api.catonetworks.com/api/v1/graphql2 | grep -i 'ratelimit\|retry'

# Verify ICMP path to vendor cloud is healthy (note: some cloud providers deprioritize ICMP)
ping -c 5 api.meraki.com

How to diagnose

  1. Manually reproduce the API call. Use curl against the same endpoint the collector polls. Inspect the full response body, not just the HTTP status code. Look for error markers inside the payload: PAN-OS <response status="error">, JSON fields like "isError": true, or empty result arrays where data should exist.

  2. Distinguish auth failure from data failure. A 401 or 403 response means the API key is expired, revoked, or the SAML/SSO admin context changed. Meraki v1 returns 404 (not 403) on a bad API key by design, to avoid leaking resource existence. If the response is 200 but empty, the auth may still be valid but the backend is not returning data.

  3. Check for rate limiting. Look for HTTP 429 responses in collector logs. Inspect response headers for Retry-After (Meraki returns this; Cato does not formally publish rate-limit headers, so verify empirically). Cato rate limits are per-query, per-account: general 120/min, accountSnapshot 1/sec, accountMetrics 15/min, eventsFeed 100/min. Multiple collectors sharing the same API key share the same counter.

  4. Check the vendor status page. If the API is returning empty payloads across multiple endpoints, check whether the vendor has an active incident. If the status page shows green but your API calls are failing, you may be early to a multi-customer incident.

  5. Compare SNMP and API data. If SNMP is available for the same device, check whether SNMP-polled data is also stale. API down with SNMP up points to a vendor-cloud-side issue. Both down points to a device or network issue.

  6. Check the API changelog. A schema change without a version bump can cause your parser to silently fail. The response may contain data, but the fields your adapter expects may have moved or been renamed.

  7. Check token expiry timing. API key rotation is the most common cause of silent failures. Tokens often expire on schedules that do not align with operator memory. Check when the current key was issued and when it expires.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
API response payload validityHTTP 200 with empty body is the primary failure modeNon-zero count of 200-with-empty-payload responses
Data freshness for API-sourced metricsStaleness is the downstream symptom of the gapTime since last non-empty response exceeding 2x poll interval
API request latencyRising latency indicates vendor-side backpressurep99 latency exceeding 5x rolling baseline
HTTP 429 rateThrottling produces intermittent gapsAny sustained 429 rate above 0
HTTP 401/403/404 rateAuth failures are security-relevantAny nonzero value is abnormal
API rate-limit remainingApproaching zero means imminent throttleBelow 20% of quota
ICMP to vendor cloudConfirms network path healthLoss or latency spike to the vendor API hostname
Vendor API schema versionSchema drift breaks parsers silentlyVersion mismatch between expected and actual

Fixes

API key expired or revoked

Rotate the key immediately. For PAN-OS, generate a new key via /api/?type=keygen&user=<u>&password=<p> and pass it as the X-PAN-KEY header or key= query parameter. For Meraki, generate a new key from the dashboard (SAML/SSO admins cannot generate API keys; use a local dashboard admin). For Cato, generate a new API key from the admin portal. Update the collector configuration and verify the new key returns valid data.

Track key expiration dates as an operational metric. Alert when any API key is within 7 days of expiry. This is the single most effective prevention measure for silent API gaps.

Vendor API schema change

Update the collector adapter to handle the new schema. Check the vendor API changelog for breaking changes. If the vendor changed the schema without a version bump, add schema validation logic that checks for expected fields in the response and flags their absence.

Vendor maintenance or backend fault

Wait for the vendor to resolve the issue. Switch to supplementary telemetry if available. If SNMP is configured on the same devices, SNMP data may continue flowing while the API is down. If the API is the only telemetry source, which is common for SD-WAN overlays, there is no fallback. Document the gap in the incident record and set expectations for when data will resume.

Rate limiting

Reduce polling frequency. Shard collectors to use separate API keys if the vendor supports multiple keys per organization. Implement exponential backoff in the collector and respect the Retry-After header where present. For Meraki, the documented limit is 10 req/sec per organization, with a burst allowance of 30 requests in 2 seconds. For Cato, different query types have different limits: accountSnapshot at 1/sec is the tightest constraint and is easily exceeded if multiple collectors share a key.

Collector adapter bug

Enable debug logging on the adapter to capture raw API responses. Compare the raw response to what the adapter is parsing. If the adapter is silently dropping valid data due to a parsing bug, file a bug report with the raw response and the expected parsed output.

Prevention

Validate response payloads, not just status codes. Every API adapter should check for expected fields in the response body. For PAN-OS, check for <response status="success"> in the XML. For JSON APIs, check that expected data arrays or objects are non-null and non-empty when they should contain data.

Monitor data freshness explicitly. Track the timestamp of the last successful, non-empty API response per endpoint. Alert when this exceeds 2x the poll interval. This catches silent gaps that status-code-only checks miss.

Track API key lifecycle. Maintain an inventory of all API keys, their creation dates, and their expiration or rotation schedules. Alert when any key approaches expiry. Some vendor keys do not have formal expiration timestamps, but they can be revoked or rotated by dashboard admins at any time.

Implement schema validation. At minimum, check that the response contains the top-level fields your adapter expects. Full JSON Schema validation is better but may be excessive for most use cases.

Share API keys deliberately. If multiple collectors share a single API key, they share a single rate-limit counter. Document which collectors use which keys, and monitor aggregate usage against the vendor’s published limits.

How Netdata helps

  • Netdata can inspect HTTP response codes and body content from vendor APIs, alerting on 200-with-empty-payload patterns that status-code-only checks miss.
  • Correlate API response validity with downstream metric freshness: if the Cato API starts returning empty payloads, Netdata shows the downstream effect on SD-WAN tunnel metrics in the same view.
  • Track API request latency alongside response validity to distinguish vendor-side backpressure from silent failures.
  • Monitor SNMP data in parallel with API data for the same devices, making it immediately visible when the API is down but SNMP is still flowing.
  • Alert on data freshness thresholds: time since last non-empty API response exceeding 2x poll interval triggers an alert regardless of HTTP status code.
  • Collect per-endpoint API error rates, including 401/403/404 responses that indicate auth issues before they become silent gaps.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.