The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / mongodb / mongodb-connection-storm-spiral

Operations Guides

MongoDB connection storm spiral: reconnection floods after an election or deploy

Connection count on a primary jumps from 200 to 4,000 in under a minute. Resident memory climbs, query latencies double, and application logs fill with timeout errors. The slow query log shows nothing unusual. Individual queries are not the problem. The database is drowning in threads.

This is a connection storm spiral. A trigger event, usually a replica set election, application deploy, or network blip, invalidates existing connections across your application fleet. Every driver reconnects at once. Each new connection costs MongoDB a dedicated thread and roughly 1 MB of stack memory. The resulting RSS spike and ticket contention slow down operations already in flight, causing more timeouts, which drives even more reconnections. The feedback loop ends in OOM kill or unresponsiveness.

The differentiator from normal pool growth is connection churn. A steady current count with a rapidly climbing totalCreated means threads are being created and destroyed faster than the server can sustain.

flowchart TD
    A[Election, deploy, or network blip] --> B[Drivers invalidate pooled connections]
    B --> C[Mass simultaneous reconnection]
    C --> D[New thread per connection]
    D --> E[Memory RSS spikes]
    E --> F[Ticket contention and scheduling overhead]
    F --> G[Operations slow and timeout]
    G --> H[More reconnections]
    H --> C

What this means

MongoDB uses a one-thread-per-connection model. When a trigger event causes mass reconnection, the server must create thousands of threads almost instantly. This consumes resident memory, increases kernel scheduling overhead, and floods the WiredTiger storage engine with concurrent operations competing for read and write tickets.

As tickets exhaust, new operations queue in globalLock.currentQueue. Latencies rise. Application drivers time out and retry, opening yet more connections. The cycle feeds itself until the node runs out of memory or file descriptors, or until the underlying trigger is resolved and reconnections stop.

A spike in current alone can be normal pool warmup. A sustained delta in totalCreated means the server is burning resources on thread lifecycle overhead. Treat totalCreated churn rate as the primary diagnostic signal.

Common causes

CauseWhat it looks likeFirst thing to check
Primary election or failoverConnection spike seconds after a new primary is elected; rs.status() shows a recent electionDaters.status() member states and MongoDB logs for "Starting an election"
Rolling application restart or deployConnections spike from many new application processes simultaneously; client IPs are distributed uniformly across the fleetApplication deployment timestamps and process start events
Network blip or DNS failureBrief drop in network throughput followed by a flood; driver logs may show pool cleared eventsNetwork latency and DNS resolution health between apps and database
Load balancer health check failureRegular cadence of connection spikes matching the LB probe interval; many connections from the LB IP rangeLB health check configuration and probe logs

Quick checks

Run these read-only commands to confirm the storm and assess severity.

// Check connection count, available slots, and total churn
var c = db.serverStatus().connections;
print("Current: " + c.current + "  Available: " + c.available + "  Total created: " + c.totalCreated);
// Check WiredTiger ticket availability (MongoDB <=7.x)
var t = db.serverStatus().wiredTiger.concurrentTransactions;
print("Read tickets available: " + t.read.available + " / " + t.read.totalTickets);
print("Write tickets available: " + t.write.available + " / " + t.write.totalTickets);
// Check queue depth
var q = db.serverStatus().globalLock.currentQueue;
print("Total queued: " + q.total + "  Readers: " + q.readers + "  Writers: " + q.writers);
// Identify heaviest client sources among active operations
db.currentOp({ active: true }).inprog.forEach(function(op) {
  print((op.client || "internal") + " | " + op.op + " | " + (op.secs_running || 0) + "s");
});
# Check OS-level RSS in MB (adjust if multiple mongod processes exist)
ps -o rss= -p $(pgrep -x mongod) | awk '{print "RSS: " $1/1024 " MB"}'
# Check for recent elections or stepdowns in the log
grep -iE "Starting an election|stepping down" /var/log/mongodb/mongod.log | tail -10
// Check server-side average latency
var lat = db.serverStatus().opLatencies;
var rOps = lat.reads.ops;
var wOps = lat.writes.ops;
print("Read avg ms: " + (rOps ? (lat.reads.latency / rOps / 1000).toFixed(2) : "N/A"));
print("Write avg ms: " + (wOps ? (lat.writes.latency / wOps / 1000).toFixed(2) : "N/A"));
// Check if throughput has collapsed
db.serverStatus().opcounters

How to diagnose it

  1. Confirm the trigger. Check rs.status() for a recent election, application deployment logs for a restart, or network metrics for a blip. The spiral almost always has an identifiable trigger within the last 1-2 minutes.
  2. Measure churn, not just count. Sample totalCreated twice, 30 seconds apart. A large delta with a stable or slowly changing current confirms a storm rather than legitimate growth.
  3. Correlate memory with connections. Compare db.serverStatus().mem.resident against the expected baseline of WiredTiger cache + (current connections x ~1MB) + internal overhead. If RSS is significantly higher, thread stacks are the likely cause.
  4. Identify the top client sources. Use db.currentOp() to see which hosts are driving the most load. In a storm, you will see a broad distribution across your application fleet rather than a single misbehaving host.
  5. Check ticket exhaustion. If available read or write tickets have dropped below 25% of total, the storage engine is saturated and operations are queuing. This is what turns a reconnect burst into a latency spiral.
  6. Assess memory exhaustion risk. If RSS is within 1 GB of system memory or the available connection count is approaching zero, the node is at risk of OOM or file descriptor exhaustion.

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Connection totalCreated churn rateThread creation and destruction consume CPU and memory; high churn with stable current is the defining storm signalDelta > 100 per minute while current is flat
connections.currentEach connection allocates a thread and stack memoryRapid 10x increase within seconds
Memory RSSThread stacks drive resident memory upward during a stormRSS growth correlates with connection spike and exceeds expected baseline
WiredTiger ticket availabilityMore connections means more operations competing for storage engine admissionAvailable read or write tickets drops below 25% of total
globalLock.currentQueue depthOperations queue when tickets or locks are scarceSustained total queue > 20
opLatencies reads and writesServer-side latency degradation from contentionAverage latency doubles from baseline for > 5 minutes
opcounters throughputTicket exhaustion and thread overhead cause throughput collapseSustained drop > 50% from baseline

Fixes

Immediate containment

The fastest way to stop the node from dying is to throttle new connections without a restart. Note that maxIncomingConnections is startup-only; it cannot be changed with setParameter. On MongoDB 8.0.12, 8.1.1, and 8.2 or later, tune the connection-establishment rate limiter at runtime instead:

db.adminCommand({setParameter: 1, ingressAdmissionControllerTicketPoolSize: <value>})

Lower values throttle connection establishment. On older versions, throttle reconnections at the network (firewall) or driver-pool layer instead. Rejected connections appear in application logs as connection failures, but this breaks the feedback loop before OOM.

Persist the change in mongod.conf after the incident. Do not restart mongod to apply this during a storm. Avoid reactive rs.stepDown() calls; an election invalidates connections and can worsen the flood.

Identify the heaviest client sources with db.currentOp() and coordinate with application owners to slow or pause restarts until the topology is stable.

Stabilize the topology

If the trigger was an election, determine why the primary stepped down. Check for heartbeat timeouts, disk stalls, or memory pressure that caused the original instability. Fixing the trigger without stabilizing the root cause will simply restart the spiral on the next event.

Throttle client reconnect behavior

After the storm subsides, review application driver configuration. Large connection pool sizes multiply the impact of any trigger. Ensure that total possible connections across all application instances leaves substantial headroom below the server limit. A common target is to operate at less than 50% of maxIncomingConnections during normal load, leaving capacity for reconnection bursts.

Prevention

  • Monitor totalCreated delta, not just current. Churn is the leading indicator.
  • Keep normal connection utilization below 50% of the configured maximum. This leaves headroom for reconnection floods.
  • Track WiredTiger ticket availability continuously. Ticket exhaustion is what turns a reconnect burst into a spiral.
  • Ensure replica set elections are rare outside of maintenance. Frequent elections indicate network instability, storage latency, or misconfigured electionTimeoutMillis.
  • Review application logs for connectionPoolCleared events after any deploy or failover. These indicate driver-level reconnection behavior that may need tuning.

How Netdata helps

  • Netdata samples serverStatus every second. A totalCreated delta that dwarfs baseline churn is visible immediately.
  • A single dashboard correlates mongodb.connections_current, system mem.rss, and mongodb.global_lock_current_queue_total, confirming the spiral in seconds.
  • Ticket availability is trended automatically, so you can see when a reconnect burst crosses into storage engine saturation.
  • Alerts on connection spikes coupled with rising queue depth or falling ticket availability reduce false positives from benign pool resizing.
The Netdata solution

MongoDB monitoring with Netdata

Netdata monitors MongoDB with per-second metrics and automatic dashboards. Watch WiredTiger cache pressure, oplog window, connection counts, checkpoint stalls, and replication health in one place, correlated with the underlying host.