Monday, 3 August 2026

Kubernetes Job object

 

A Kubernetes Job is a controller object designed to run a batch task to completion.

Unlike Deployments or ReplicaSets (which keep applications running indefinitely) or CronJobs (which trigger tasks on a schedule), a Job creates one or more Pods, executes the workload, and ensures they terminate cleanly. Once the specified number of Pods complete successfully, the Job itself is marked as complete and stops.

Standard Job Manifest


apiVersion: batch/v1
kind: Job
metadata:
  name: data-migration-job
spec:
  backoffLimit: 4             # Number of retries before marking the job failed
  completions: 1              # Number of successful pod completions required
  parallelism: 1              # How many pods run concurrently
  ttlSecondsAfterFinished: 600 # Clean up job & pods 10 minutes after completion
  template:
    spec:
      containers:
      - name: migration-task
        image: python:3.11-slim
        command: ["python", "-c", "print('Running database migration...'); import time; time.sleep(10); print('Done!')"]
      restartPolicy: OnFailure # Required: OnFailure or Never (Always is invalid)


Core Execution Patterns


Kubernetes Jobs support three primary workload execution models:

  • 1. Non-Parallel Jobs
    • Behavior: Starts a single Pod and waits for it to complete successfully.
    • Use Case: One-off database schema migrations, report generation, or administrative scripts.
  • 2. Parallel Jobs with Fixed Completions
    • Behavior: Runs multiple Pods in parallel until a total number of successful completions (spec.completions) is reached.
    • Use Case: Batch processing where $N$ independent tasks need to be completed.
  • 3. Parallel Jobs with Work Queue
    • Behavior: Pods coordinate via an external message queue (e.g., RabbitMQ, Redis, SQS). Each Pod pulls work until the queue is empty, then exits.
    • Use Case: High-throughput task processing, media transcoding, or distributed data transformation.


Key Configuration Fields


Field
  • Default
  • Description

restartPolicy
  • Required
  • Must be OnFailure (restarts container inside same Pod) or Never (spawns a new Pod on failure).

backoffLimit
  • 6
  • Maximum number of retries before marking the Job as failed.

completions
  • 1
  • Total number of successful Pod terminations needed for Job completion.

parallelism
  • 1
  • Max number of Pods allowed to run concurrently at any given moment.

activeDeadlineSeconds
  • Unlimited
  • Max time allowed for the entire Job (including retries) before terminating all running Pods.

completionMode
  • NonIndexed
  • Set to Indexed to assign each Pod a unique completion index ($0$ to $\text{completions}-1$) via environment variables.

ttlSecondsAfterFinished
  • Disabled
  • Automatically deletes the Job and its underlying Pods after $N$ seconds of finishing.

Essential kubectl Commands


Operation                                 Command
==================         ===========
Create a Job imperatively         kubectl create job my-job --image=busybox -- echo "Hello World"
Get Job status                            kubectl get jobs
Inspect job details                     kubectl describe job my-job
List Pods associated with Job   kubectl get pods --selector=batch.kubernetes.io/job-name=my-job
View logs of Job Pods              kubectl logs job/my-job
Delete Job and its Pods            kubectl delete job my-job


Jobs vs Deployments vs CronJobs        



                    ┌─────────────────────────┐
                    │      Workload Type      │
                    └────────────┬────────────┘
                                 │
           ┌─────────────────────┴─────────────────────┐
           ▼                                           ▼
   Long-Running Services                     Batch Tasks / One-Off
(Deployments, StatefulSets)                      (Jobs & CronJobs)
           │                                           │
  Maintains target Pod                       Executes task, then
  count indefinitely.                        terminates cleanly.
                                                       │
                                      ┌────────────────┴────────────────┐
                                      ▼                                 ▼
                                Single Run                         Scheduled Run
                                  (Job)                              (CronJob)


Kubernetes CronJob


A Kubernetes CronJob creates and manages short-lived Jobs on a scheduled, repeating basis. It is the Kubernetes equivalent of a standard Unix crontab file, making it ideal for periodic tasks like database backups, report generation, or maintenance scripts.

Minimal Example Manifest

apiVersion: batch/v1
kind: CronJob
metadata:
  name: nightly-backup
spec:
  schedule: "0 2 * * *" # Runs every day at 02:00 UTC
  timeZone: "Etc/UTC"   # Optional: set preferred timezone (Kubernetes 1.27+)
  concurrencyPolicy: Forbid
  startingDeadlineSeconds: 100
  successfulJobsHistoryLimit: 3
  failedJobsHistoryLimit: 1
  jobTemplate:
    spec:
      template:
        spec:
          containers:
          - name: backup-task
            image: alpine:latest
            command:
            - /bin/sh
            - -c
            - echo "Running database backup..."; sleep 5
          restartPolicy: OnFailure


Schedule Syntax Quick Reference


The schedule field uses standard cron syntax with 5 fields:

minute hour day-of-month month day-of-week

Schedule Format Interpretation


*    * * * *             Every minute
*/15 * * * *              Every 15 minutes
0    0 * * *              Every day at midnight
0    9 * * 1              Every Monday at 9:00 AM

Critical Settings

  • concurrencyPolicy: Controls how overlapping executions are handled when a previous run hasn't finished:
    • Allow (default): Runs concurrent jobs simultaneously. 
    • Forbid: Skips the new job if the previous one is still running. 
    • Replace: Cancels the currently running job and starts the new one. 
  • startingDeadlineSeconds: The deadline (in seconds) for starting a job if it missed its scheduled time (e.g., cluster was temporarily down).
  • successfulJobsHistoryLimit / failedJobsHistoryLimit: Number of completed or failed Job/Pod records to keep for auditing before automatic cleanup.
  • restartPolicy: Must be set on the pod template spec to either OnFailure or Never (Always is invalid for Jobs).  

Helpful kubectl Commands


Task                                            Command 
====                                           ========
List CronJobs                              kubectl get cronjobs
Inspect configuration                  kubectl describe cronjob <name>
Manually trigger immediately    kubectl create job --from=cronjob/<cronjob-name>                        <manual-job-name>
Pause schedule                           kubectl patch cronjob <name> -p '{"spec":                             {"suspend":true}}'
View logs of latest run               kubectl logs job/<job-name>



CronJobs Inner Mechanism


Under the hood, Kubernetes CronJobs rely on a decentralized control loop pattern. They are not handled by a traditional Linux cron daemon running on a single server, but rather by the Kubernetes Control Plane through cascading controllers.

How CronJobs Are Implemented


The implementation follows a 3-tier hierarchical model:

CronJob Object --> Job Object --> Pod(s)

Rather than running code directly, a CronJob acts as a factory for Job objects, which in turn manage the Pods where your container actually executes


┌─────────────────────────────────────────────────────────┐
│                 kube-controller-manager                 │
│                                                         │
│   ┌─────────────────┐       Creates      ┌─────────┐  │
│   │ CronJob Controller│ ─────────────────> │   Job   │  │
│   └──────────────────┘                    └───┬───┘  │
└──────────────────────────────────────────────────┼──────┘
                                                   │
                                                Creates
                                                   │
                                                   ▼
                                              ┌─────────┐
                                              │   Pod   │
                                              └─────────┘


The Control Loop Mechanism

  1. Synchronization Loop: The CronJob Controller runs inside kube-controller-manager. Every ~10 seconds, it iterates through all CronJob objects defined in the cluster.  
  2. Schedule Checking: The controller parses the schedule field (e.g., 0 * * * *) and compares the current time against the last time the job was executed.  
  3. Job Spawning: If a run is due, the CronJob controller reads the embedded jobTemplate and creates an actual Job resource.  
  4. Execution: The cluster's separate Job Controller detects the newly created Job resource and spawns one or more Pods to execute your container workload to completion.  
  5. Garbage Collection: Depending on successfulJobsHistoryLimit and failedJobsHistoryLimit, the CronJob controller periodically deletes old completed Job objects (and their associated logs/pods).  


Who Controls Them?


Control over CronJobs is split between system components (automation) and users/roles (permissions).

System Component Control

  • kube-controller-manager: The core control plane component where the CronJob controller code actually executes. If this component is down, scheduled triggers will pause until it recovers.  
  • kube-apiserver: Stores the desired state in etcd and validates user manifests.
  • kube-scheduler: Assigns the individual Pods spawned by the resulting Jobs to healthy worker nodes.

User & Permission Control (RBAC)

Human administrators and automated service accounts control CronJobs via Kubernetes Role-Based Access Control (RBAC):

Role / Action              Required API Permissions (batch/v1)
=============     ==============================
Manage Schedules     create, update, patch, delete on cronjobs
View Status                get, list, watch on cronjobs
Manual Trigger          create permissions on jobs (to invoke kubectl create job --from=cronjob/...)


Example RBAC Role for CronJob Operators:


apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  namespace: prod
  name: cronjob-operator
rules:
- apiGroups: ["batch"]
  resources: ["cronjobs", "jobs"]
  verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]


Technical Considerations

  • At-Least-Once Execution: Kubernetes schedules are designed around at-least-once execution semantics. Due to control loop timing or network hiccups, a scheduled job might occasionally run twice or run slightly late. Workloads should always be designed to be idempotent
  • Timezones: Controller clocks default to UTC or the local time of kube-controller-manager unless explicit timezones are passed via spec.timeZone (supported in K8s 1.27+). 


CronJob is a controller object


In Kubernetes, a controller is a control loop that watches the state of your cluster through the API server and makes changes attempting to move the current state toward the desired state.

Here is how a CronJob fits into the controller pattern:

Why CronJob is a Controller

  • Custom Resource / Spec & Status Model: Like Deployment, ReplicaSet, and Job, a CronJob has an API object schema (spec defining desired behavior, status tracking execution state).
  • Control Loop Execution: The CronJob implementation runs as a control loop inside the kube-controller-manager component.
  • Cascading Controller Pattern: CronJob sits at the top of a controller hierarchy:

CronJob Controller --[creates/manages]--> Job Controller --[creates/manages]--> Pods

  • The CronJob Controller reconciles the CronJob spec: it checks the schedule, creates Job resources when a execution is due, cleans up old jobs based on history limits, and handles concurrency policies.
  • The Job Controller reconciles those created Job resources to manage individual Pods to completion.

Summary Table


Controller                        API Group      What it Watches      What it Creates/Manages
========                       =========     =============      ===================
CronJob Controller            batch/v1          CronJob specs            Job objects
Job Controller                    batch/v1          Job specs                    Pod objects
Deployment Controller      apps/v1           Deployment specs      ReplicaSet objects


Grafana Alerting

 


How it works

  • Grafana alerting periodically queries data sources and evaluates the condition defined in the alert rule
  • If the condition is breached, an alert instance fires
  • Firing instances are routed to notification policies based on matching labels
  • Notifications are sent out to the contact points specified in the notification policy

How to set alerts

  • Alert rules: Create an alert rule to query a data source and evaluate the condition defined in the alert rule. 
    • There are two types of alerts rules:
      • Grafana-managed. Examples:
        • APM
        • AWS
        • Data Quality
        • Kubernetes
        • MongoDB
        • Storage
        • Synthetics
      • Data source-managed. Data sources containing configured alerts rules are for example Mimir or Loki data sources where alert rules are stored and evaluated in the data sources itself. In these data sources you can select Manage alerts via Alerting UI to be able to manage these alerts rules in the Grafana UI as well as in the data source where they were configured.
        • Prometheus
        • Mimir
        • Loki
    • Define the condition that must be met before an alert rule fires
  • Route alert notifications either directly to a contact point or through notification policies for more flexibility
    • Contact points: Configure who receives notifications and how they are sent
    • Notification policies: Configure how firing alert instances are routed to contact points
  • Monitor your alert rules using dashboards and visualizations