The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / docker / docker-dns-not-working

Operations Guides

Docker DNS not working inside containers

Your application logs show connection timeouts. Health checks against dependency names are failing. curl from inside a container returns “Could not resolve host” while the same name resolves fine on the host. DNS inside Docker is not a simple passthrough to the host resolver. It is a stack of namespace-specific forwarders, embedded resolvers, and inherited configuration that breaks in specific, repeatable ways.

This guide will help you determine whether the failure is a missing embedded DNS, a poisoned resolv.conf, an upstream forwarding issue, or a version regression. You will be able to distinguish between inter-container name resolution failures and external lookup failures, identify the root cause with safe read-only checks, and apply the correct fix without guessing.

What this means

On user-defined bridge networks, including the networks Docker Compose creates automatically, Docker injects an embedded DNS forwarder at 127.0.0.11:53 into each container’s /etc/resolv.conf. This resolver handles container-to-container name lookups and forwards external queries to upstream servers that Docker cached from the host at container creation time.

On the default bridge network, there is no embedded DNS. Containers receive a snapshot of the host’s /etc/resolv.conf at startup. They cannot resolve each other by container name unless you use the deprecated --link flag.

Because the DNS configuration is snapshotted when the container starts, dynamic changes on the host, such as a VPN connecting or systemd-resolved switching upstream servers, are invisible to running containers. The embedded DNS is also a forwarder, not a recursive resolver. If the upstream it cached becomes unreachable, all external resolution fails, even though the container’s network stack is otherwise healthy.

Common causes

CauseWhat it looks likeFirst thing to check
systemd-resolved stub copied into containerresolv.conf shows nameserver 127.0.0.53; DNS returns SERVFAIL or times outHost’s /etc/resolv.conf symlink target and daemon-level DNS config
Default bridge network limitationsContainers cannot resolve each other by name, only by IPContainer network mode via docker inspect
Container attached to wrong networkOne service cannot reach another by name while other pairs workdocker network inspect membership for both containers
ndots and search domain stormsSlow startup, high query volume, apparent hangs on unqualified namesoptions line inside container /etc/resolv.conf
Upstream DNS unreachableExternal names fail; container-to-container names may still workReachability of upstream IPs from the container namespace
Docker 29.1.0 regressionAfter restarting containers created by Docker 29.0.x, external DNS resolution failed on custom bridge networksdocker version output; recreation fixes it
Firewall or VPN intercepting port 53IP connectivity works but all name resolution failsping -c 3 1.1.1.1 works from inside the container
HTTPS/TYPE65 lookup for internal hostnameEmbedded DNS forwards it upstream; internal-name lookup failsQuery type and whether the name should resolve internally

Quick checks

Run these read-only checks before making changes.

# Check the container's current DNS configuration
docker exec <container> cat /etc/resolv.conf

# Check which network mode the container uses
docker inspect --format '{{.HostConfig.NetworkMode}}' <container>

# Test name resolution from inside the container
docker exec <container> nslookup <hostname>

# Check if the host uses systemd-resolved
systemctl is-active systemd-resolved

# Check the actual file Docker copied resolv.conf from
readlink -f /etc/resolv.conf

# Check Docker version for known regressions
docker version --format '{{.Server.Version}}'

# Test embedded DNS directly in the container's network namespace
CONTAINER_PID=$(docker inspect --format '{{.State.Pid}}' <container>)
nsenter -t ${CONTAINER_PID} -n dig @127.0.0.11 <hostname>

# Check container logs for DNS-related errors
docker logs <container> 2>&1 | grep -iE "dns|resolve|lookup|nxdomain"

# Verify layer-3 connectivity independent of DNS
docker exec <container> ping -c 3 1.1.1.1

How to diagnose it

Follow this flow from symptom to root cause.

  1. Determine the container’s network mode. Run docker inspect --format '{{.HostConfig.NetworkMode}}' <container>. If the result is default, the container is on the default bridge. Name resolution between containers by name is not supported here. If you need service discovery, move the containers to a user-defined network or use Docker Compose, which creates one automatically.

  2. Inspect /etc/resolv.conf inside the container. If the nameserver is 127.0.0.11, the embedded DNS is in use. If it is 127.0.0.53, Docker copied the systemd-resolved stub address into the container namespace. Inside the container, 127.0.0.53 refers to the container’s own loopback, not the host’s resolver. This produces SERVFAIL or timeouts.

  3. Test internal versus external names. Run nslookup for another container name on the same network, then for an external domain like cloudflare.com. If internal names fail but external names work, the embedded DNS may not know about the target container. Verify both containers are attached to the same network. If both fail, the embedded DNS cannot reach its upstream forwarders.

  4. Verify IP connectivity. Run docker exec <container> ping -c 3 1.1.1.1. If this fails, the problem is not DNS. Check network interfaces, firewall rules, and bridge status. If IP works but DNS fails, the problem is strictly in name resolution.

  5. Check the host’s live resolver state. On hosts with systemd-resolved, /etc/resolv.conf often symlinks to /run/systemd/resolve/stub-resolv.conf. Docker copies this file at container creation. The real upstream servers live in /run/systemd/resolve/resolv.conf. If the host’s stub is in the container, configure Docker to use real upstream IPs or the real servers file.

  6. Look for ndots misconfiguration. If options ndots:5 appears in the container’s resolv.conf, unqualified names like postgres trigger multiple suffix-appended lookups before falling back. This causes slow starts and CPU-bound blocking in some runtimes. Docker’s default on user-defined networks is ndots:0, but host inheritance or Kubernetes overrides can push ndots:5 into containers.

  7. Check Docker version against known regressions. Docker 29.1.0 introduced a regression affecting restarted containers created by Docker 29.0.x: after upgrading, Docker misread their generated /etc/resolv.conf as user-modified and skipped setting up external DNS resolvers. On custom bridge networks, resolv.conf still showed 127.0.0.11, but external queries failed. This was fixed in 29.1.1. If you are on 29.1.0, recreate the affected containers or upgrade the engine.

  8. Review daemon logs for embedded DNS or libnetwork errors. Run journalctl -u docker.service | grep -iE "dns|error|network". Errors here can reveal a hung embedded DNS, corrupted iptables rules, or plugin failures that do not show up inside the container.

flowchart TD
    A[DNS fails inside container] --> B{On user-defined network?}
    B -->|No| C[Migrate to custom network or use IPs]
    B -->|Yes| D{resolv.conf shows 127.0.0.53?}
    D -->|Yes| E[Configure daemon.json with real upstream DNS]
    D -->|No| F{Internal or external failure?}
    F -->|Internal| G[Check network attachment and container names]
    F -->|External| H{IP reachable?}
    H -->|No| I[Check host network, VPN, or firewall]
    H -->|Yes| J[Check upstream DNS, ndots, and Docker version]

Metrics and signals to monitor

DNS failures rarely happen in isolation. Correlate these signals to distinguish a DNS incident from a general network or daemon health issue.

SignalWhy it mattersWarning sign
Container network errorsveth pair issues or bridge drops manifest as packet loss that breaks DNS queriesrx_errors or tx_errors increasing
Docker DNS resolution latencyEmbedded DNS at 127.0.0.11 is part of dockerd; slowness here degrades application startupResolution time >100 ms for external names
Container restart countApplications that depend on DNS for service discovery may enter crash loops when resolution failsRestart count increasing with exit code 1
Docker daemon response latencyA stressed or deadlocked daemon slows the embedded DNS forwarder/_ping or docker ps latency >500 ms sustained
Docker bridge connection countconntrack exhaustion silently drops packets, including UDP port 53nf_conntrack_count / nf_conntrack_max >70%
Container OOM killed statusDNS lookup storms from ndots misconfiguration can spike memory or CPUOOMKilled: true in container state

Fixes

If the cause is systemd-resolved

Configure explicit upstream DNS servers in /etc/docker/daemon.json:

{
  "dns": ["1.1.1.1", "8.8.8.8"],
  "dns-search": ["corp.example"]
}

Restart Docker with systemctl restart docker. Existing containers are not affected. Recreate them to pick up the new resolver. Alternatively, configure Docker to read from /run/systemd/resolve/resolv.conf, which contains the real upstream addresses, instead of the stub file.

If the cause is default bridge network limitations

Move containers to a user-defined bridge network. Docker Compose does this automatically for every project. Do not use --link; it is deprecated and absent from modern Compose.

If the cause is ndots or search domains

Add "dns-opts": ["ndots:0"] or ["ndots:1"] to /etc/docker/daemon.json. This prevents suffix expansion on unqualified names. Recreate containers after changing daemon-level DNS options.

If the cause is upstream unreachability

Verify the upstream servers configured in daemon.json are reachable from the host. If the host relies on a VPN that dynamically rewrites resolvers, be aware that running containers continue using the old set. You must recreate containers after the VPN changes host DNS, or use fixed daemon-level DNS IPs that are valid in all network states.

If the cause is the Docker 29.1.0 regression

Upgrade to Docker 29.1.1 or later. As a workaround without upgrading, recreate all affected containers. For Compose workloads, run docker compose down && docker compose up -d.

If the cause is firewall or VPN interception

Third-party firewalls that intercept outbound DNS on port 53 can block Docker’s embedded DNS from reaching upstream resolvers. If ping -c 3 1.1.1.1 works from inside the container but name resolution fails, whitelist the container network’s path to port 53 or switch to an upstream that is not intercepted.

If the cause is IPv6 AAAA queries

Docker’s embedded DNS handles A and PTR records for container names. It does not synthesize AAAA answers from IPv4-only container records. HTTPS and TYPE65 queries are forwarded to upstream DNS rather than answered from internal container records; for internal names with no upstream answer, the query fails. Use IPv4 A lookups for internal container names, or handle the absence of AAAA at the application level.

Prevention

  • Use user-defined networks for all multi-container workloads. This enables the embedded DNS and service discovery by default.
  • Set explicit DNS in daemon.json. Never rely on the host’s dynamic resolv.conf if the host runs systemd-resolved. Point Docker to real, stable upstream IPs.
  • Keep ndots low. Set "dns-opts": ["ndots:0"] in daemon.json to avoid lookup storms in containerized applications.
  • Recreate containers after host network changes. Docker’s DNS snapshot does not update dynamically. Treat VPN connections and resolver changes as events that require container recreation.
  • Test Docker engine upgrades in staging. Version regressions like 29.1.0 can break DNS for existing containers. Validate upgrades with live workloads before production rollout.
  • Monitor the signals. Alert on container restart counts, daemon latency, and DNS resolution latency from within representative containers.

How Netdata helps

Netdata collects the signals that surround a DNS incident, letting you correlate rather than guess.

  • Container network error charts show rx_errors and tx_dropped per container. A jump here alongside DNS timeouts points to a veth or bridge issue, not a resolver bug.
  • Docker daemon latency monitoring tracks how fast the API responds. Since the embedded DNS lives inside dockerd, rising /_ping latency is an early warning that DNS forwarding will degrade.
  • Container restart count and exit code tracking reveal when an application is crash-looping because it cannot resolve a dependency name.
  • conntrack utilization on the host shows whether the connection tracker is approaching exhaustion, which silently drops UDP DNS packets before they ever reach a resolver.
  • CPU throttling metrics catch the secondary effect of ndots-induced lookup storms, where a container spends excessive time in resolver syscalls and hits its CFS quota.
The Netdata solution

Docker monitoring with Netdata

Netdata auto-discovers every container and monitors Docker with per-second CPU, memory, disk, and network metrics plus ML anomaly detection. Catch disk exhaustion, OOM cascades, daemon hangs, and log explosions before they take the host down.