The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / mongodb / mongodb-scatter-gather-queries

Operations Guides

MongoDB scatter-gather queries: when mongos fans out to every shard

In a sharded MongoDB cluster, mongos reads chunk-to-shard mappings from the config servers and routes each query to the smallest set of shards that can satisfy it. When the predicate includes the shard key, mongos targets one shard or a bounded subset. When it does not, mongos broadcasts the query to every shard that owns chunks for the collection and merges the results. This is a scatter-gather query.

On a cluster with fifty shards, a scatter-gather query can consume fifty times the CPU, memory, and I/O of a targeted query. Latency is gated by the slowest responding shard. Aggregations, reporting queries, and some multi-document operations legitimately need to touch all shards. The operational problem is when scatter-gather queries dominate OLTP traffic, turning horizontal scale into a tightly coupled liability.

What it is and why it matters

A scatter-gather query, or broadcast operation, is any read or write where the query filter lacks the shard key and mongos cannot narrow execution to a specific shard. Mongos opens a cursor on every shard that owns chunks for the collection, collects results, and merges them.

For aggregations, blocking stages such as $group, $bucket, $bucketAuto, $count, and $sort split the pipeline. Shards execute the first half in parallel; results return to mongos for merge except when the pipeline uses $out, $lookup into an unsharded collection, or a sorting/grouping stage with allowDiskUse enabled.

This removes the isolation that sharding is meant to provide. Every shard pays the cost of parsing the query, acquiring WiredTiger tickets, scanning indexes, and returning results. Each shard uses its own ticket pool, so a broadcast consumes tickets on every node simultaneously and can block targeted queries on otherwise cold shards. Mongos holds a client cursor for each shard cursor; if one shard stalls, the whole operation blocks until that shard responds or the cursor times out.

How it works

Mongos targets or broadcasts based on shard key metadata in the config servers. For a compound shard key such as { a: 1, b: 1, c: 1 }, a query containing { a: 1 } or { a: 1, b: 1 } routes to a subset of shards. A query on { b: 1 } alone broadcasts because it misses the prefix.

flowchart TD
    Client([Client query])
    Mongos[mongos router]
    Config[config servers]
    Decision{Shard key
in predicate?} Target[Targeted shard subset] Broadcast[Every shard] Merge[Merge on mongos
or single shard] Client -->|sends query| Mongos Mongos -->|reads metadata| Config Mongos -->|routes| Decision Decision -->|yes| Target Decision -->|no| Broadcast Target -->|returns| Client Broadcast -->|returns cursors| Merge Merge -->|returns| Client

Multi-document writes such as updateMany() and deleteMany() broadcast to all shards unless the query contains an equality match on the full shard key. Single-document operations such as updateOne(), replaceOne(), and deleteOne() must include the shard key or _id to target; otherwise MongoDB returns an error rather than broadcasting.

Warning: Untargeted updateMany() and deleteMany() can cause sudden load spikes and replication lag on all secondaries. Test the predicate on a staging cluster before running against production.

Untargeted count() and distinct() operations also broadcast to all shards.

Use explain("executionStats") on a suspect aggregation to see the routing decision. The output includes mergeType ("anyShard", "specificShard", or "router" in current releases; older 6.0/7.0 branches use "primaryShard", "anyShard", or "mongos"), splitPipeline, and per-shard execution stats. A shards array containing every shard in the cluster confirms a broadcast. When mergeType is "specificShard", the mergeShard field names the node that performs the final merge.

Write concern and read concern are enforced independently on each shard. A scatter-gather write with writeConcern: "majority" must be majority-committed on every shard before mongos acknowledges success. A read with readConcern: "snapshot" or "majority" waits for the slowest shard to satisfy the guarantee, adding latency tails.

Skip and limit introduce merge work that cannot be pushed to shards. Mongos applies skip() locally after retrieving results. When a query combines skip(n) with limit(m), mongos passes (n+m) to each shard to reduce the merge volume, but the skip still happens on the router. Paginated APIs that rely on deep skip values perform poorly on sharded collections.

Pipelines that target a single shard can still merge on a shard. Official behavior pins the merge to the output collection for $out and to the unsharded collection for $lookup. With allowDiskUse: true and a sorting or grouping stage, MongoDB runs the merge on a randomly selected targeted shard to use that shard’s local disk headroom. The final phase of a split $sort is a streaming merge. The final phase of $group is blocking, which means the chosen merger shard can become a bottleneck.

When mongos fans out, it applies the configured read preference independently on each shard. The selected member is governed by both read preference and replication.localPingThresholdMs, re-evaluated per operation. Secondary reads in a scatter-gather query add replication lag to the overall response time on every shard.

Where it shows up in production

Check the mongos slow query log for nShards equal to the total shard count. On the shell, explain("executionStats") shows a shards array that includes every shard, or a top-level SHARD_MERGE stage. A mergeType field in aggregation explain output also signals a broadcast.

Uniform load across all shards for a query that should be isolated is another tell. If opcounters, ticket utilization, or opLatencies spike in lockstep across every replica set, a scatter-gather query is likely running. This is common with OLTP queries that filter on a non-shard-key field, such as looking up a user by email when the collection is sharded on customer_id. It also appears when collections are sharded by _id, where every query that does not know the exact document ID fans out. Background analytics jobs that query by created_at range without the shard key prefix will broadcast even for small time windows.

updateMany and deleteMany operations in application code are frequent culprits when they omit the shard key. Aggregation pipelines with $lookup to a sharded foreign collection force scatter-gather on the joined side. Queries using skip() for pagination without shard-key prefixes are another source.

Low-cardinality shard keys and monotonically increasing shard keys create related problems. A key on continent yields at most seven chunks regardless of cluster size, so adding shards provides no benefit. An ascending shard key such as a timestamp concentrates writes on the chunk bounded by MaxKey, creating a hot shard. Both issues force a re-examination of shard key design.

Tradeoffs and when to use it

Scatter-gather is correct behavior when the system cannot determine where data lives. Reporting queries, cross-shard aggregations, and administrative operations legitimately scan all shards. The concern is frequency.

When evaluating a shard key, remember that prefix matching matters. A compound key serves queries on its leading fields. If the application most often queries by email but the shard key is { tenant_id: 1, user_id: 1 }, email lookups broadcast unless they also include tenant_id.

The tradeoff is between write distribution and read targeting. A hashed shard key on _id spreads writes evenly but forces almost every read to broadcast unless the _id is known. A monotonically increasing key like a timestamp targets time-range queries but creates a hot insert shard. The ideal shard key has high cardinality, appears in the majority of query predicates, and distributes writes evenly.

If you must run scatter-gather queries in production, set maxTimeMS to bound the impact of a slow shard.

Signals to watch in production

SignalWhy it mattersWarning sign
Slow query log nShardsConfirms fan-out scopenShards equals total shard count on mongos
explain("executionStats") shards arrayReveals routing decisionArray contains all shards for OLTP patterns
db.serverStatus().shardingStatisticsmongos-level scatter-gather rateIncreasing untargeted query count
Per-shard opcounters and latencyBroadcast loads all shards uniformlyAll shards spike together for one query pattern
metrics.document.returned vs opcounters.queryHigh ratio means each query returns many documentsCombined with scatter-gather, indicates broad filters
Chunk distributionImbalance compounds routing inefficiencyGreater than 20% skew between most and least loaded shard
db.currentOp() on mongosShows active scatter-gather operationsLong-running find or aggregate without shard-key filters
replSetGetStatus.members[].optimeDateReplication lag adds to broadcast latencyAll secondaries lagging together after scatter-gather read phase
explain("executionStats") executionTimeMillisPer-shard vs total execution timeLarge gap indicates merge or network overhead on mongos

How Netdata helps

Netdata monitors opcounters and opLatencies per shard, making uniform load spikes from broadcast queries visible in one view.

Correlating mongos query latency with per-shard CPU, disk I/O, and WiredTiger ticket utilization reveals whether a latency spike comes from one hot shard or from all shards responding to a scatter-gather query.

Tracking globalLock.currentQueue across all shards exposes when a broadcast query is queuing on multiple replica sets simultaneously.

Process-level metrics on mongos highlight merge latency, helping determine whether slowness is in shard execution or in the router’s merge phase.

The Netdata solution

MongoDB monitoring with Netdata

Netdata monitors MongoDB with per-second metrics and automatic dashboards. Watch WiredTiger cache pressure, oplog window, connection counts, checkpoint stalls, and replication health in one place, correlated with the underlying host.