The only agent that thinks for itself

Autonomous Monitoring with self-learning AI built-in, operating independently across your entire stack.

Unlimited Metrics & Logs
Machine learning & MCP
5% CPU, 150MB RAM
3GB disk, >1 year retention
800+ integrations, zero config
Dashboards, alerts out of the box
> Discover Netdata Agents

Centralized metrics streaming and storage

Aggregate metrics from multiple agents into centralized Parent nodes for unified monitoring across your infrastructure.

Stream from unlimited agents
Long-term data retention
High availability clustering
Data replication & backup
Scalable architecture
Enterprise-grade security
> Learn about Parents

Fully managed cloud platform

Access your monitoring data from anywhere with our SaaS platform. No infrastructure to manage, automatic updates, and global availability.

Zero infrastructure management
99.9% uptime SLA
Global data centers
Automatic updates & patches
Enterprise SSO & RBAC
SOC2 & ISO certified
> Explore Netdata Cloud

Deploy Netdata Cloud in your infrastructure

Run the full Netdata Cloud platform on-premises for complete data sovereignty and compliance with your security policies.

Complete data sovereignty
Air-gapped deployment
Custom compliance controls
Private network integration
Dedicated support team
Kubernetes & Docker support
> Learn about Cloud On-Premises

Powerful, intuitive monitoring interface

Modern, responsive UI built for real-time troubleshooting with customizable dashboards and advanced visualization capabilities.

Real-time chart updates
Customizable dashboards
Dark & light themes
Advanced filtering & search
Responsive on all devices
Collaboration features
> Explore Netdata UI

Monitor on the go

Native iOS and Android apps bring full monitoring capabilities to your mobile device with real-time alerts and notifications.

iOS & Android apps
Push notifications
Touch-optimized interface
Offline data access
Biometric authentication
Widget support
> Download apps

The future of infrastructure observability

See our strategic direction across AI-native observability, full-stack signals, operational intelligence, and enterprise platform maturity.

AI-native observability
Full-stack signal coverage
Operational intelligence
Enterprise platform maturity
Agent releases every 6 weeks
Cloud continuous delivery
> Explore Product Roadmap

Best energy efficiency

True real-time per-second

100% automated zero config

Centralized observability

Multi-year retention

High availability built-in

Zero maintenance

Always up-to-date

Enterprise security

Complete data control

Air-gap ready

Compliance certified

Millisecond responsiveness

Infinite zoom & pan

Works on any device

Native performance

Instant alerts

Monitor anywhere

AI-native observability

Continuous delivery

Open source foundation

80% Faster Incident Resolution

AI-powered troubleshooting from detection, to root cause and blast radius identification, to reporting.

True Real-Time and Simple, even at Scale

Linearly and infinitely scalable full-stack observability, that can be deployed even mid-crisis.

90% Cost Reduction, Full Fidelity

Instead of centralizing the data, Netdata distributes the code, eliminating pipelines and complexity.

See and Map Your Entire Network

Live topology, flow analytics, and SNMP device and trap monitoring — unified with your full-stack observability.

Control Without Surrender

SOC 2 Type 2 certified with every metric kept on your infrastructure.

Integrations

800+ collectors and notification channels, auto-discovered and ready out of the box.

800+ data collectors
Auto-discovery & zero config
Cloud, infra, app protocols
Notifications out of the box
> Explore integrations
Real Results
46% Cost Reduction

Reduced monitoring costs by 46% while cutting staff overhead by 67%.

— Leonardo Antunez, Codyas

Zero Pipeline

No data shipping. No central storage costs. Query at the edge.

From Our Users
"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

No Query Language

Point-and-click troubleshooting. No PromQL, no LogQL, no learning curve.

Enterprise Ready
67% Less Staff, 46% Cost Cut

Enterprise efficiency without enterprise complexity—real ROI from day one.

— Leonardo Antunez, Codyas

SOC 2 Type 2 Certified

Zero data egress. Only metadata reaches the cloud. Your metrics stay on your infrastructure.

Full Coverage
800+ Collectors

Auto-discovered and configured. No manual setup required.

Any Notification Channel

Slack, PagerDuty, Teams, email, webhooks—all built-in.

Built for the People Who Get Paged

Because 3am alerts deserve instant answers, not hour-long hunts.

Every Industry Has Rules. We Master Them.

See how healthcare, finance, and government teams cut monitoring costs 90% while staying audit-ready.

Monitor Any Technology. Configure Nothing.

Install the agent. It already knows your stack.
From Our Users
"A Rare Unicorn"

Netdata gives more than you invest in it. A rare unicorn that obeys the Pareto rule.

— Eduard Porquet Mateu, TMB Barcelona

99% Downtime Reduction

Reduced website downtime by 99% and cloud bill by 30% using Netdata alerts.

— Falkland Islands Government

Real Savings
30% Cloud Cost Reduction

Optimized resource allocation based on Netdata alerts cut cloud spending by 30%.

— Falkland Islands Government

46% Cost Cut

Reduced monitoring staff by 67% while cutting operational costs by 46%.

— Codyas

Real Coverage
"Plugin for Everything"

Netdata has agent capacity or a plugin for everything, including Windows and Kubernetes.

— Eduard Porquet Mateu, TMB Barcelona

"Out-of-the-Box"

So many out-of-the-box features! I mostly don't have to develop anything.

— Simon Beginn, LANCOM Systems

Real Speed
Troubleshooting in 30 Seconds

From 2-3 minutes to 30 seconds—instant visibility into any node issue.

— Matthew Artist, Nodecraft

20% Downtime Reduction

20% less downtime and 40% budget optimization from out-of-the-box monitoring.

— Simon Beginn, LANCOM Systems

Pay per Node. Unlimited Everything Else.

One price per node. Unlimited metrics, logs, users, and retention. No per-GB surprises.

Free tier—forever
No metric limits or caps
Retention you control
Cancel anytime
> See pricing plans

What's Your Monitoring Really Costing You?

Most teams overpay by 40-60%. Let's find out why.

Expose hidden metric charges
Calculate tool consolidation
Customers report 30-67% savings
Results in under 60 seconds
> See what you're really paying

Your Infrastructure Is Unique. Let's Talk.

Because monitoring 10 nodes is different from monitoring 10,000.

On-prem & air-gapped deployment
Volume pricing & agreements
Architecture review for your scale
Compliance & security support
> Start a conversation

Monitoring That Sells Itself

Deploy in minutes. Impress clients in hours. Earn recurring revenue for years.

30-second live demos close deals
Zero config = zero support burden
Competitive margins & deal protection
Response in 48 hours
> Apply to partner

Per-Second Metrics at Homelab Prices

Same engine, same dashboards, same ML. Just priced for tinkerers.

Community: Free forever · 5 nodes · non-commercial
Homelab: $90/yr · unlimited nodes · fair usage
> Get the Homelab Plan

$1,000 Per Referral. Unlimited Referrals.

Your colleagues get 10% off. You get 10% commission. Everyone wins.

10% of subscriptions, up to $1,000 each
Track earnings inside Netdata Cloud
PayPal/Venmo payouts in 3-4 weeks
No caps, no complexity
> Get your referral link
Cost Proof
40% Budget Optimization

"Netdata's significant positive impact" — LANCOM Systems

Calculate Your Savings

Compare vs Datadog, Grafana, Dynatrace

Savings Proof
46% Cost Reduction

"Cut costs by 46%, staff by 67%" — Codyas

30% Cloud Bill Savings

"Reduced cloud bill by 30%" — Falkland Islands Gov

Enterprise Proof
"Better Than Combined Alternatives"

"Better observability with Netdata than combining other tools." — TMB Barcelona

Real Engineers, <24h Response

DPA, SLAs, on-prem, volume pricing

Why Partners Win
Demo Live Infrastructure

One command, 30 seconds, real data—no sandbox needed

Zero Tickets, High Margins

Auto-config + per-node pricing = predictable profit

Homelab Ready
Free Video Course

8-episode Netdata tutorial by LearnLinux.tv

76k+ GitHub Stars

3rd most starred monitoring project

Worth Recommending
Product That Delivers

Customers report 40-67% cost cuts, 99% downtime reduction

Zero Risk to Your Rep

Free tier lets them try before they buy

AI Support Assistant, Available 24/7

Nedi has access to all official documentation, source code, and resources. Ask any question about Netdata—responds in your language.

Deployment & configuration
Troubleshooting & sizing
Alerts & notifications
Evidence-based answers
> Ask Nedi now

Never Fight Fires Alone

Docs, community, and expert help—pick your path to resolution.

Learn.netdata.cloud docs
Discord, Forums, GitHub
Premium support available
> Get answers now

60 Seconds to First Dashboard

One command to install. Zero config. 850+ integrations documented.

Linux, Windows, K8s, Docker
Auto-discovers your stack
> Read our documentation

76,000+ Engineers Strong

615+ contributors. 1.5M daily downloads. One mission: simplify observability.

Per-Second. 90% Cheaper. Data Stays Home.

Side-by-side comparisons: costs, real-time granularity, and data sovereignty for every major tool.

See why teams switch from Datadog, Prometheus, Grafana, and more.

> Browse all comparisons
Edge-Native Observability, Born Open Source
Per-second visibility, ML on every metric, and data that never leaves your infrastructure.
Founded in 2016
615+ contributors worldwide
Remote-first, engineering-driven
Open source first
> Read our story
Promises We Publish—and Prove
12 principles backed by open code, independent validation, and measurable outcomes.
Open source, peer-reviewed
Zero config, instant value
Data sovereignty by design
Aligned pricing, no surprises
> See all 12 principles
Edge-Native, AI-Ready, 100% Open
76k+ stars. Full ML, AI, and automation—GPLv3+, not premium add-ons.
76,000+ GitHub stars
GPLv3+ licensed forever
ML on every metric, included
Zero vendor lock-in
> Explore our open source
Build Real-Time Observability for the World
Remote-first team shipping per-second monitoring with ML on every metric.
Remote-first, fully distributed
Open source (76k+ stars)
Challenging technical problems
Your code on millions of systems
> See open roles
Meet the Team Behind Netdata
Conferences, meetups, and tradeshows where you can see Netdata in action and talk to the engineers who build it.
Live demos and deep dives
Book 1-on-1 meetings
Talks and panel sessions
Event recaps and photos
> See all events
Talk to a Netdata Human in <24 Hours
Sales, partnerships, press, or professional services—real engineers, fast answers.
Discuss your observability needs
Pricing and volume discounts
Partnership opportunities
Media and press inquiries
> Book a conversation
Your Data. Your Rules.
On-prem data, cloud control plane, transparent terms.
Trust & Scale
76,000+ GitHub Stars

One of the most popular open-source monitoring projects

SOC 2 Type 2 Certified

Enterprise-grade security and compliance

Data Sovereignty

Your metrics stay on your infrastructure

Validated
University of Amsterdam

"Most energy-efficient monitoring solution" — ICSOC 2023, peer-reviewed

ADASTEC (Autonomous Driving)

"Doesn't miss alerts—mission-critical trust for safety software"

Community Stats
615+ Contributors

Global community improving monitoring for everyone

1.5M+ Downloads/Day

Trusted by teams worldwide

GPLv3+ Licensed

Free forever, fully open source agent

Why Join?
Remote-First

Work from anywhere, async-friendly culture

Impact at Scale

Your work helps millions of systems

$ guides / kubernetes / kubernetes-job-cronjob-troubleshooting

Operations Guides

Kubernetes Job and CronJob troubleshooting: history, backoff, and missed runs

You deployed a CronJob to run every minute, but the last successful run was three hours ago. Or a critical data-processing Job failed with BackoffLimitExceeded after six silent retries, leaving a trail of failed Pods and no clear signal about what broke. Batch workloads fail differently from long-running services: they are time-bound, retry-sensitive, and leave debris in etcd if you do not clean them up. Read the failure signals, distinguish retry storms from control plane delays, and fix the root cause without guessing.

What this means

A Job creates Pods and retries failures until enough succeed. A CronJob wraps a Job template with a schedule. Three mechanisms cause most operator confusion:

  • backoffLimit: The number of Pod failures allowed before the Job is marked Failed with reason BackoffLimitExceeded. The default is 6. Retries are counted differently depending on restartPolicy. With OnFailure, container restarts within the same Pod count. With Never, each failed Pod counts as one retry.
  • History limits: successfulJobsHistoryLimit (default 3) and failedJobsHistoryLimit (default 1) control how many completed Jobs a CronJob retains. Older Jobs are deleted automatically, but if these fields are unset or misconfigured, completed Jobs and their Pods can accumulate and bloat etcd.
  • Missed runs and startingDeadlineSeconds: The CronJob controller evaluates missed schedules from the last scheduled time until now. If more than 100 missed schedules accumulate, the controller stops starting Jobs entirely and logs a warning. The startingDeadlineSeconds field defines a catch-up window; values below 10 seconds often cause silent skips because the controller reconciles on roughly 10-second intervals.

Common causes

CauseWhat it looks likeFirst thing to check
backoffLimit exhaustedJob status Failed with reason BackoffLimitExceeded; multiple failed Pods with increasing restart countsPod logs and events for the most recent failure
History accumulationHundreds of completed Jobs or Pods in the namespace; etcd database size growingsuccessfulJobsHistoryLimit and failedJobsHistoryLimit on the CronJob
Missed runs due to control plane latencyCronJob status shows no recent active Jobs; events mention “Cannot determine if job needs to be started”API server and etcd latency; node readiness
startingDeadlineSeconds too tightCronJob with a sub-10-second window never creates JobsThe CronJob spec.startingDeadlineSeconds value
concurrencyPolicy Forbid masking slownessCronJob skips runs silently because the previous Job is still activeDuration of the active Job versus the CronJob schedule interval
Disruption consuming retry budgetPods evicted or preempted count toward backoffLimit, exhausting retries before the workload runsPod status and events for DisruptionTarget or eviction
Missing cleanupCompleted Jobs and Pods remain indefinitely; node disk or etcd pressure increasesttlSecondsAfterFinished and activeDeadlineSeconds settings

Quick checks

Run these checks in order. All are read-only and safe to run during an incident.

# List Jobs and their status
kubectl get jobs -n <namespace> -o wide

# Inspect a specific Job for conditions and events
kubectl describe job <job-name> -n <namespace>

# Check CronJob schedule, history limits, and last run times
kubectl describe cronjob <cronjob-name> -n <namespace>

# View logs from the most recent failed Pod (use --previous only if the container restarted)
kubectl logs -n <namespace> <pod-name>

# Check for backoff or deadline events
kubectl get events -n <namespace> --field-selector reason=BackoffLimitExceeded

# Verify if CronJob active Jobs are blocking new runs
kubectl get cronjob <cronjob-name> -n <namespace> -o jsonpath='{.status.active}'

# Check cluster-wide Job count to spot accumulation
kubectl get jobs --all-namespaces --no-headers | wc -l

# Check etcd database size if history limits are high or unset
# Adjust certificate paths to match your cluster (example paths below are for kubeadm)
ETCDCTL_API=3 etcdctl --endpoints=https://127.0.0.1:2379 \
  --cacert=/etc/kubernetes/pki/etcd/ca.crt \
  --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt \
  --key=/etc/kubernetes/pki/etcd/healthcheck-client.key \
  endpoint status --cluster -w table

How to diagnose it

Use this flow to narrow down whether the issue is retry exhaustion, scheduling delay, or control plane backlog.

  1. Determine the Job status. Run kubectl get job <name> -o jsonpath='{.status.conditions}'. A condition with type Failed and reason BackoffLimitExceeded means the retry budget is spent. A condition with type Complete means the Job finished but may not have been cleaned up. If there are no conditions and .status.active shows Pods, the Job is still running.
  2. Read the most recent Pod failure. Find the newest Pod owned by the Job (kubectl get pods --selector=job-name=<job-name>). Check its container status state.terminated.reason (or lastState.terminated.reason if the container restarted). OOMKilled, Error, or ContainerCannotRun point to workload issues. Evicted or DisruptionTarget point to cluster-level pressure.
  3. Check for missed run semantics on CronJobs. Look at the CronJob status fields lastScheduleTime and lastSuccessfulTime. If lastScheduleTime is far behind the current time and no Job exists, the controller may have skipped the run. Check events for the message “Cannot determine if job needs to be started. Too many missed start times (>100).” This appears when the controller has given up on catch-up.
  4. Validate startingDeadlineSeconds. If .spec.startingDeadlineSeconds is set, ensure it is at least 10 seconds. Values below 10 seconds are known to cause silent skips because the CronJob controller reconciles on approximately 10-second intervals.
  5. Check concurrency and active Job count. For a CronJob with concurrencyPolicy: Forbid, a long-running Job blocks all subsequent scheduled runs. List Jobs in the namespace and inspect the ACTIVE column for Jobs owned by the CronJob. If a Job has been active longer than the schedule interval, subsequent runs were silently skipped.
  6. Inspect control plane health if runs are missing cluster-wide. Slow API server writes or high etcd fsync latency can delay CronJob creation. Check API server latency and etcd leader stability. See How the Kubernetes control plane works and Kubernetes API server etcd latency.
  7. Audit history and retention. Check successfulJobsHistoryLimit and failedJobsHistoryLimit. If they are set to high values or omitted on custom controllers, completed Jobs accumulate. Check the total Job count across namespaces. High counts correlate with slow LIST operations and etcd bloat.
flowchart TD
    A[CronJob missed or Job failed] --> B{Job status condition}
    B -->|Failed: BackoffLimitExceeded| C[Inspect latest pod failure reason]
    B -->|Complete but retained| D[Check ttlSecondsAfterFinished and history limits]
    B -->|Active / no condition| E{Is it a CronJob?}
    E -->|Yes| F{Check missed runs}
    F -->|Last schedule far behind| G[Check startingDeadlineSeconds and control plane latency]
    F -->|Blocked by active job| H[Check concurrencyPolicy and job duration]
    E -->|No| I[Check resource pressure and pod logs]
    C --> J{Failure type}
    J -->|Resource or app error| K[Fix workload spec or dependencies]
    J -->|Disruption / eviction| L[Add podFailurePolicy or reduce churn]
    G --> M[Recreate CronJob if >100 missed schedules]

Metrics and signals to monitor

SignalWhy it mattersWarning sign
Job completion durationLong-running Jobs block CronJob schedules and consume cluster resourcesp99 completion duration exceeds the CronJob interval
CronJob missed schedulesController could not create a Job on timeIncrease in missed schedule counter or stale lastScheduleTime
Pod restart count for JobsRepeated restarts consume the backoffLimit budget quicklyRestart count approaching backoffLimit before completion
API server mutating request latencySlow writes delay Job and CronJob object creationp99 latency > 1s sustained
etcd WAL fsync latencySlow etcd causes the API server to stall, which delays CronJob reconciliationp99 fsync > 100ms
Controller workqueue depthBacklog in the CronJob or Job controller means reconciliation is falling behindDepth > 0 sustained
Cluster Job object countAccumulated Jobs increase etcd size and LIST latencyTotal Jobs growing without bound over days
Pod eviction eventsEvicted batch Pods waste retries and can trigger backoffLimit exhaustionEviction events correlated with Job Pod names

Fixes

If the cause is backoff exhaustion

Increase backoffLimit only if the workload is legitimately retryable, such as transient network dependencies. If the Job fails because of a code bug or missing ConfigMap, raising the limit will only create noise. Set activeDeadlineSeconds to cap total runtime and prevent runaway retries. On Kubernetes 1.26 and later, JobPodFailurePolicy and PodDisruptionConditions are enabled by default; the feature gate was removed after 1.31, making the behavior built in. Add a podFailurePolicy with restartPolicy: Never to ignore failed Pods carrying the DisruptionTarget condition so node pressure or preemption does not exhaust retries.

If the cause is missed runs or schedule skew

Set startingDeadlineSeconds to a value between 10 and 300 seconds, depending on your tolerance for catch-up. Never set it below 10 seconds. If the CronJob has accumulated more than 100 missed schedules, delete and recreate the CronJob resource. The controller does not recover automatically from this state. If your workload must not overlap runs, use concurrencyPolicy: Forbid, but monitor the Job completion time and alert if it approaches the schedule interval.

If the cause is history accumulation

Set ttlSecondsAfterFinished on the Job template or on standalone Jobs so the TTL controller deletes them automatically. For CronJobs, lower successfulJobsHistoryLimit to 1 or 2 if you only need the last run for debugging, and set failedJobsHistoryLimit to 2 or 3 to retain enough context without hoarding objects. If you use a GitOps workflow, ensure your templating does not override these fields to zero or unset on every sync.

If the cause is control plane latency

Fixing the batch workload spec will not help if the API server or etcd is the bottleneck. Follow the control plane latency troubleshooting path. Reduce etcd object churn by cleaning up completed Jobs and events, and verify that the CronJob controller is not being throttled by API Priority and Fairness queues.

Prevention

  • Set TTL and history limits by default. Every CronJob manifest should specify ttlSecondsAfterFinished, successfulJobsHistoryLimit, and failedJobsHistoryLimit. Every standalone Job should specify ttlSecondsAfterFinished.
  • Monitor schedule skew. Alert when lastScheduleTime on a CronJob is older than 1.5 * schedule_interval. This catches silent skips before they become an outage.
  • Test control plane maintenance windows. During API server upgrades or etcd compaction, CronJob creation can lag. Know whether your critical batch workloads can tolerate a 1-5 minute delay.
  • Align resource requests with peak batch usage. A Job that requests 100m CPU but briefly spikes to 2 cores will be throttled or evicted, wasting retries. Size requests based on observed peak usage.
  • Use Pod failure policies to ignore non-application failures. Explicitly exclude DisruptionTarget and other cluster-level conditions so node maintenance does not burn the retry budget.

How Netdata helps

  • Correlate CronJob missed runs with API server write latency and etcd fsync duration to determine whether the control plane is the bottleneck.
  • Monitor per-node CPU throttling and memory pressure that cause batch Pod evictions, which in turn consume backoffLimit retries.
  • Track controller workqueue depth and Job object counts to spot accumulation trends before etcd bloat becomes critical.
  • Alert on sustained increases in Pod restart counts for Jobs, providing an early signal that a batch workload is failing before backoffLimit is reached.
The Netdata solution

Kubernetes monitoring with Netdata

Netdata monitors Kubernetes with per-second metrics across the control plane, nodes, and every pod, with ML anomaly detection and zero per-pod configuration. Correlate API-server and etcd latency, kubelet PLEG stalls, scheduling pressure, and OOMKills in one place.