The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / network / network-syslog-severity-normalization

Operations Guides

Normalizing syslog severity across vendors: why 'critical' isn't critical

RFC 5424 defines eight syslog severity levels, numbered 0 through 7: Emergency, Alert, Critical, Error, Warning, Notice, Informational, and Debug. Every major network vendor implements the same numeric scale. The integer that means “Critical” on a Cisco router means “Critical” on a Juniper switch.

But the severity a device assigns to a given event is not standardized. A BGP session reset might arrive as severity 5 (Notice) from one vendor and severity 3 (Error) from another. A hardware alarm that one platform logs as Critical (2), another logs as Alert (1) or Warning (4). Same operational condition, different severity label, same RFC.

Teams that forward raw vendor severity into a SIEM without a normalization layer build alerting rules on an unstable foundation. A rule that triggers on “severity at or above Critical” fires for some vendors and silently misses the same event class from others.

The production failure pattern

A firewall sends a “Warning” severity message about a threat detection. The SIEM alerting rule triggers only on “Critical” and “Alert.” It never fires. The threat goes undetected for hours.

Any multi-vendor network that feeds syslog into a SIEM, log aggregator, or alerting pipeline inherits this problem. The question is not whether severity mismatch exists. It is whether your normalization layer accounts for it before the data reaches your alert rules.

How it works

The RFC 5424 severity scale

The eight levels, from most to least severe:

NumericLabelTypical operational meaning
0EmergencySystem is unusable
1AlertAction must be taken immediately
2CriticalCritical conditions
3ErrorError conditions
4WarningWarning conditions
5NoticeNormal but significant condition
6InformationalInformational messages
7DebugDebug-level messages

All vendors agree on this mapping. No vendor assigns a different label to numeric 2. The disagreement is upstream: which events get which numbers. That decision is left to each vendor’s firmware developers.

Threshold filtering is inclusive upward

When you configure a severity threshold on a network device, the device sends that level and all numerically lower (more severe) levels. Configuring “error” (3) on a Juniper device captures Emergency (0), Alert (1), Critical (2), and Error (3). The same inclusive behavior holds on Cisco, Arista, and FortiOS.

Set the threshold too permissively and you flood the collector with debug and informational noise. Set it too restrictively and you silently suppress events the vendor assigned to Notice or Informational but that you consider operationally important. Device-side threshold filtering controls volume. It does not normalize severity across vendors.

Facility codes differ across vendors

Severity is only half of the PRI field. The other half is the facility code, which identifies the subsystem that generated the message. Facility codes also differ across vendors, and this breaks rules that filter on facility.

Juniper uses various facility codes for different subsystems, including kernel (0), user (1), daemon (3), authorization (4), and local facilities for firewall and forwarding-plane messages.

Cisco routing-plane messages commonly use local6 or local7 depending on platform and configuration. A SIEM rule that filters on a specific facility will silently drop identical-severity messages from a different vendor using a different facility code. If your normalization layer remaps severity but not facility, you have solved only half the problem.

Two message formats in the wild

RFC 5424 obsoleted RFC 3164 (the BSD syslog format). RFC 5424 added structured-data elements, a version field, and a strictly defined message structure. Most legacy network gear still emits RFC 3164-style messages. Both formats remain in active use, and collectors must parse both.

The PRI field (and therefore the severity encoding) is calculated the same way in both formats: PRI = facility * 8 + severity. What differs is the surrounding message structure. A parser that expects RFC 5424 and receives RFC 3164 will fail to extract fields correctly, not because the severity is encoded differently, but because the message layout differs. Test your collector against actual device output for both formats.

flowchart TD
    A[Multi-vendor syslog stream] --> B[Collector / forwarder]
    B --> C{Per-vendor severity normalization?}
    C -->|No: raw passthrough| D[SIEM rules match on raw severity]
    C -->|Yes: remap table| E[SIEM rules match on normalized severity]
    D --> F[Missed alerts on some vendors]
    E --> G[Consistent alerting across vendors]

Where it shows up in production

The SIEM alert gap

A SIEM rule pages on “critical” events from network devices. It matches on the severity keyword or numeric code in the syslog message. It works for the vendor the team tested against. It silently fails for other vendors that assign different severity to the same event class.

This is most dangerous for security events. A threat detection message logged as Warning (4) by one firewall vendor may represent the same operational urgency as Critical (2) from another. If the alert threshold is set at Critical, the Warning event passes through unflagged.

Compound hostnames breaking inventory matching

Junos OS Evolved appends the node name to the hostname by default, producing hostnames like “ptxhost-re0” or “ptxhost-fpc0” in syslog messages. Some monitoring systems fail to match these compound hostnames against their inventory database. The syslog message arrives with the correct severity but is orphaned because the SIEM cannot map it to a known device.

The workaround is set system syslog alternate-format, which prepends the node name to the process identifier instead, keeping the bare hostname intact. This is a Juniper-specific configuration.

Management instance changes after upgrade

Starting in Junos OS Release 24.2R1, syslog traffic no longer defaults to the dedicated management instance when management-instance is configured. You must explicitly configure mgmt_junos for system log traffic to use the management VRF. Operators upgrading from older Junos OS Evolved releases may find that syslog stops arriving after the upgrade. The cause is not a severity change but a routing-instance change that breaks the transport path. Verify syslog routing-instance configuration post-upgrade.

Windows Event Log does not map to syslog severity

If your estate includes Windows servers alongside network devices, the severity mismatch extends beyond network vendors. Windows Event Log uses a different taxonomy: Critical, Error, Warning, Information, Verbose. There is no Emergency or Alert level. A Windows “Critical” event maps approximately to syslog Error (3) or Critical (2) depending on the application, not to Emergency (0). Correlating Windows and network-device severity in the same SIEM dashboard requires explicit mapping logic.

Common misuses and normalization failures

Alert rules that match on keyword, not numeric code. A rule that matches the string “critical” will miss numeric severity 2 if the syslog message encodes severity numerically. Conversely, a rule that matches numeric severity at or below 2 will match Critical, Alert, and Emergency but will miss a vendor that logs the same event as Warning (4). Match on numeric codes where possible, and normalize the codes per vendor before the alert layer sees them.

Assuming facility consistency. Facility codes differ across vendors. A normalization layer that remaps severity but passes facility through unchanged will produce rules that work for some vendors and silently fail for others.

Relying on device-side threshold filtering as the only control. Device-side filtering reduces volume but cannot normalize across vendors. Two devices configured with the same threshold (“error”) will send different event sets because they assign different severities to the same events. Normalization must happen at the collector or SIEM layer.

Mixing timezones. Timestamps in syslog are device-local-time unless explicitly UTC (RFC 3339 / RFC 5424). Timestamps from devices in different timezones (or with NTP drift) will break event correlation even with perfect severity normalization. Normalize time alongside severity. Force UTC everywhere and monitor device NTP offset.

Ignoring parser format differences. If your parser handles only RFC 3164 or only RFC 5424, messages in the other format will be mis-parsed or dropped. Verify that the collector handles both and test with actual device output.

Building a normalization layer

Audit current severity assignments. Before building remapping rules, query your SIEM or log aggregator for the same event class (BGP peer down, interface down, fan failure) across vendors and compare raw severity values. This reveals the mismatch surface area in your specific estate.

Per-vendor severity remapping. Build a lookup table that maps (vendor, event-class, raw-severity) to a normalized severity. This requires per-vendor calibration: review the syslog output from each vendor in your estate, identify the event classes that matter operationally, and assign a normalized severity based on operational impact, not the vendor’s label.

Facility normalization. Alongside severity, remap facility codes to a normalized scheme. Group vendor-specific facilities into operational categories: routing-plane, firewall, authentication, hardware, config-change. Filter and alert on the normalized category, not the raw facility code.

Format-aware parsing. Ensure the collector handles both RFC 3164 and RFC 5424. Validate that severity extraction works for both formats. Test with actual device output from each vendor, not synthetic messages.

Signals to watch in production

SignalWhy it mattersWarning sign
Syslog severity distribution per vendorReveals vendor-specific severity patterns; baseline what “normal” looks like per vendorSudden shift in distribution for one vendor indicates config change or new event type
SIEM alert rate by source vendorIf one vendor generates most alerts, others may be under-alerting due to severity mismatchOne vendor disproportionately quiet compared to peers
UDP receive buffer errors (RcvbufErrors in /proc/net/snmp)Drops during syslog storms lose the highest-priority messages first, including root-cause indicatorsCounter incrementing during incidents means root-cause syslog likely lost
Syslog source count (devices actively sending)A device that stops sending syslog is invisible to severity-based alertingDrop to zero from one device while others still report indicates isolation or logging subsystem failure
Syslog parser error rateFormat mismatch (RFC 3164 vs 5424) produces parse errors, not silent dropsRising parse errors after a firmware upgrade indicates format change
Timestamp skew between syslog and corroborating signalsPerfect severity normalization fails if timestamps do not align across sourcesEvents that should correlate appearing minutes or seconds apart indicates clock drift

How Netdata helps

  • Netdata collects syslog metrics alongside SNMP, flow, and system-level signals, letting you correlate severity spikes with interface state changes or BGP session drops.
  • Per-device syslog collection lets you baseline severity distribution per vendor and detect shifts that indicate config changes before they break alert rules.
  • Collector health metrics, including receive rate and buffer drops, surface silent syslog data loss at the collector.
The Netdata solution

Network monitoring with Netdata

Netdata monitors network infrastructure with per-second interface metrics, SNMP, NetFlow/sFlow/IPFIX, and ML anomaly detection. Correlate interface flapping, packet drops, routing changes, and traffic spikes with the systems that depend on them.