DevOps Automation & Scripting

Reviewed & published by Brayan K

Master Python automation for infrastructure management, deployment pipelines, monitoring, backups, and production system orchestration

Part of the free Python course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn

Why Python for DevOps Automation?

Python has become the de facto standard for DevOps automation, replacing shell scripts with safer, more maintainable solutions.

FeatureBash ScriptsPython Automation
Error handlingCryptic exit codestry/except with clear messages
Cross-platformLinux/Mac onlyWorks everywhere
API integrationRequires curl hacksNative requests/boto3
MaintainabilityHard to read at scaleClean, testable code

Common Use Cases

File & Directory Automation

Every DevOps workflow involves managing files: rotating logs, cleaning temporary data, synchronizing directories, and organizing backups.

Common Tasks

Real-World Example

A production CI server runs a cleanup script every hour to remove build artifacts older than 7 days, preventing disk space exhaustion. This same pattern applies to log management, cache cleanup, and temporary file handling.

🧑‍🏫 Worked example: a backup-retention job you can actually run

This is a complete, working version of the "keep the newest N, delete the rest" job. It builds its own throwaway sandbox directory first, so running it cannot touch anything you care about. Two habits are built in and both are worth stealing: it does a dry run before deleting anything, and it is idempotent — running it twice does no extra damage. Read the comments, then run it.

# WORKED EXAMPLE: a backup-retention job - "keep the newest N, delete the rest".
# This exact pattern runs nightly on more production servers than any other script.
import tempfile
from pathlib import Path

# --- sandbox so this is completely safe to run ------------------------------
sandbox = Path(tempfile.mkdtemp())        # a throwaway directory, nothing real at risk
for name in [
    "backup-2026-08-21.tar.gz",
    "backup-2026-08-22.tar.gz",
    "backup-2026-08-23.tar.gz",
    "backup-2026-08-24.tar.gz",
    "backup-2026-08-25.tar.gz",
    "notes.txt",                          # NOT a backup - the glob below must ignore it
]:
    (sandbox / name).write_text("pretend backup data")

KEEP = 3                                  # retention policy: keep the 3 newest backups

def find_deletable(directory, keep):
    """Return the backups outside the retention window, oldest first."""
    # ISO date stamps (YYYY-MM-DD) sort oldest -> newest as plain text, which is
    # exactly why backups get named this way.
    backups = sorted(directory.glob("backup-*.tar.gz"))
    if keep == 0:
        return backups                    # guard: backups[:-0] would wrongly be empty
    return backups[:-keep]                # everything except the last `keep` items

def cleanup(directory, keep, dry_run=True):
    """Delete old backups. Returns how many files it acted on."""
    doomed = find_deletable(directory, keep)
    for path in doomed:
        if dry_run:
            print(f"[DRY RUN] would delete {path.name}")
        else:
            path.unlink()                 # unlink() is pathlib's "delete this file"
            print(f"deleted   {path.name}")
    return len(doomed)

total = len(list(sandbox.glob("backup-*.tar.gz")))
print(f"Found {total} backups, policy says keep {KEEP}")

# Rule one of destructive automation: dry-run first, every time.
count = cleanup(sandbox, KEEP, dry_run=True)
print(f"-> {count} file(s) would be removed")
print()

cleanup(sandbox, KEEP, dry_run=False)     # now do it for real
print()
print("Remaining:", sorted(p.name for p in sandbox.iterdir()))

# Idempotent means "safe to run again" - a second run finds nothing left to do.
print("Second run removed:", cleanup(sandbox, KEEP, dry_run=False), "file(s)")

# ✅ Output:
# Found 5 backups, policy says keep 3
# [DRY RUN] would delete backup-2026-08-21.tar.gz
# [DRY RUN] would delete backup-2026-08-22.tar.gz
# -> 2 file(s) would be removed
#
# deleted   backup-2026-08-21.tar.gz
# deleted   backup-2026-08-22.tar.gz
#
# Remaining: ['backup-2026-08-23.tar.gz', 'backup-2026-08-24.tar.gz', 'backup-2026-08-25.tar.gz', 'notes.txt']
# Second run removed: 0 file(s)

System Command Execution

The subprocess module provides safe, controlled execution of system commands with proper error handling and timeout management.

Best Practices

Common Operations

Task Scheduling

Modern DevOps requires more intelligent scheduling than traditional cron. Python provides flexible alternatives.

ToolBest ForComplexity
CronSimple, one-off scriptsLow
APSchedulerIn-process schedulingMedium
Celery BeatDistributed, high-volumeHigh

APScheduler (Python)

More powerful: retry on failure, parallel execution, event-based triggers, state management

Celery Beat

Distributed task queue with advanced scheduling capabilities

Typical Scheduled Tasks

Server Health Monitoring

Proactive monitoring prevents outages. Python can track system resources and alert teams before problems escalate.

Metrics to Monitor

The psutil Library

psutil is the standard for cross-platform system monitoring in Python:

Provides CPU, memory, disk, network, and process information on Linux, macOS, and Windows.

API & Webhook Automation

Modern infrastructure is API-driven. Python integrates seamlessly with CI/CD systems, monitoring tools, and cloud platforms.

Common Integrations

Event-driven deployment

Git push → trigger pipeline → deploy

Automated alerting

High CPU → send Slack alert → scale infrastructure

Self-healing systems

Service down → restart automatically → notify team

Docker Automation

The Docker Python SDK enables comprehensive container lifecycle management from within Python scripts.

Automation Tasks

Production Use Case

A maintenance script runs nightly to clean up stopped containers and dangling images, preventing disk space issues. It also restarts any containers marked as unhealthy by Docker's health checks.

Kubernetes Automation

The Kubernetes Python client allows programmatic cluster management, enabling GitOps-style automation.

Automation Capabilities

Advanced Patterns

• Blue/green deployments - maintain two production environments

• Canary releases - gradually roll out changes to a subset of users

• Automatic rollback - revert on health check failure

• Multi-cluster management - orchestrate across regions

Backup Automation & Data Rotation

Regular, automated backups are essential for disaster recovery. Python orchestrates the entire backup lifecycle.

Backup Strategy

3-2-1 Backup Rule

3 copies of data • 2 different media types • 1 offsite copy

Python scripts can implement this automatically: local disk, network storage, cloud backup.

Log Processing & Automated Alerts

Logs contain critical information about system health, security events, and errors. Automated analysis prevents issues from going unnoticed.

Log Analysis Tasks

Alert Triggers

• Error threshold - Alert when error rate exceeds 1%

• Security events - Failed login attempts, suspicious patterns

• Performance degradation - Response time above threshold

• Service crashes - Application or container restarts

🎯 Your turn: write the alert rule

Log parsing is written for you below. What is missing is the part that matters — the rule that decides whether a human gets woken up. Fill in the three ___ blanks and run it. The last blank is the exit code: cron, systemd timers and CI runners all treat a non-zero exit code as "this job failed", which is how your script actually raises the alarm.

# 🎯 YOUR TURN - finish the log alerting rule
# Replace each ___ below, then press Run.

log_lines = [
    "2026-08-27 09:00:01 INFO  checkout ok",
    "2026-08-27 09:00:04 ERROR payment gateway timeout",
    "2026-08-27 09:00:07 WARN  slow query 1.8s",
    "2026-08-27 09:00:09 ERROR payment gateway timeout",
    "2026-08-27 09:00:12 INFO  checkout ok",
    "2026-08-27 09:00:15 ERROR database connection refused",
]

ERROR_BUDGET = 2          # more errors than this in one window = page someone

counts = {"INFO": 0, "WARN": 0, "ERROR": 0}

for line in log_lines:
    for level in counts:              # looping a dict gives you its keys
        if ___ in line:               # 👉 replace ___ with the variable holding the level name
            counts[level] += 1

print("Log summary:", counts)

if counts["ERROR"] ___ ERROR_BUDGET:  # 👉 replace ___ so this fires only when errors EXCEED the budget
    print(f"ALERT: {counts['ERROR']} errors this window (budget {ERROR_BUDGET})")
    exit_code = ___                   # 👉 replace ___ with the exit code that means "job failed"
else:
    print("OK: error count within budget")
    exit_code = 0

print("exit code:", exit_code)

# ✅ Expected output once the blanks are filled in:
# Log summary: {'INFO': 2, 'WARN': 1, 'ERROR': 3}
# ALERT: 3 errors this window (budget 2)
# exit code: 1
#
# Then raise ERROR_BUDGET to 5 and run again - the alert should go quiet.

Zero-Downtime Deployment

Production deployments must minimize or eliminate downtime. Python orchestrates sophisticated deployment strategies.

Deployment Pipeline

1. Pull latest code from Git

2. Run test suite - abort on failure

3. Build Docker image

4. Push to container registry

5. Update Kubernetes deployment

6. Wait for health checks to pass

7. Rollback automatically if unhealthy

8. Send deployment notification

Safety Mechanisms

Security Considerations

Automation scripts often run with elevated privileges. Security must be a top priority.

✓ Best Practices

✗ Security Anti-Patterns

Building Production-Ready Automation

Professional automation systems require more than working code. They need reliability, observability, and maintainability.

Essential Components

The DevOps Loop

Write automation → Test thoroughly → Deploy → Monitor → Learn from failures → Improve → Repeat

Every automation failure is an opportunity to make the system more resilient.

🎯 Mini Challenge: Self-Healing Watchdog

A watchdog is the script that checks your services and fixes what it can without waking anyone. The rules are simple, and the important one is the escalation rule: a service that keeps dying must stop being restarted and start being escalated, or you build a machine that hides a real fault forever.

The starter below is an outline only — no logic is written for you. Match the printed text exactly and your output will match the expected output.

# 🎯 MINI-CHALLENGE: a self-healing service watchdog

services = {
    "web":    {"healthy": True,  "restarts": 0},
    "worker": {"healthy": False, "restarts": 1},
    "cache":  {"healthy": False, "restarts": 3},
}
MAX_RESTARTS = 3

# 1. Create two counters: restarted = 0 and escalated = 0
#
# 2. Loop over services with:  for name, info in services.items():
#
# 3. For each service apply these rules, in this order:
#      healthy                              -> print(f"OK        {name}")
#      not healthy, restarts < MAX_RESTARTS -> print(f"RESTART   {name} (attempt {info['restarts'] + 1} of {MAX_RESTARTS})")
#                                              and add 1 to restarted
#      otherwise                            -> print(f"ESCALATE  {name} - {info['restarts']} restarts already, paging on-call")
#                                              and add 1 to escalated
#
# 4. Print a blank line, then f"{restarted} restarted, {escalated} escalated"

# your code here

# ✅ Expected output:
# OK        web
# RESTART   worker (attempt 2 of 3)
# ESCALATE  cache - 3 restarts already, paging on-call
#
# 1 restarted, 1 escalated
#
# Stretch goal: make it idempotent-aware. Add "last_action" to each service and skip
# any service you already restarted in this run.

Key Takeaways

📋 Quick Reference — DevOps Automation

Tool / ModuleWhat it does
pathlib.PathModern file and directory manipulation
subprocess.run(cmd, check=True)Run shell commands from Python
shutil.copy2 / shutil.rmtreeHigh-level file operations
docker SDKManage Docker containers from Python
psutilMonitor CPU, memory, and processes

🎉 Great work! You've completed this lesson.

You can now automate deployments, manage infrastructure, and build self-healing systems using Python's DevOps toolkit.

Practice quiz

Which module is the safe, recommended way to run system commands from Python?

  • os.system
  • commands
  • subprocess
  • shlex

Answer: subprocess. subprocess.run gives controlled execution with output capture, timeouts, and return-code checking.

What is a best practice when calling subprocess for security?

  • Pass the command as a list and avoid shell=True to prevent injection
  • Always use shell=True
  • Concatenate user input into a string
  • Disable timeouts

Answer: Pass the command as a list and avoid shell=True to prevent injection. Passing a list like ['ls', '-la'] and avoiding shell=True prevents shell-injection attacks.

Why set a timeout on subprocess calls?

  • To speed up the command
  • To capture stderr
  • It is required syntax
  • To prevent the script hanging on an unresponsive command

Answer: To prevent the script hanging on an unresponsive command. A timeout stops the script from hanging indefinitely if a command never returns.

Which library is the standard for cross-platform system metrics (CPU, memory, disk)?

  • os
  • psutil
  • sys
  • platform

Answer: psutil. psutil provides CPU, memory, disk, network, and process info across Linux, macOS, and Windows.

Which class in the lesson uses the Docker SDK to clean up stopped containers and restart unhealthy ones?

  • DockerAutomation
  • DockerManager
  • ContainerBot
  • DockerClient

Answer: DockerAutomation. The DockerAutomation class wraps docker.from_env() to remove stopped containers, prune images, and restart unhealthy ones.

In the deployment workflow, what happens if the Kubernetes rollout fails its health check?

  • The script ignores it
  • It deletes the deployment
  • It automatically rolls back with kubectl rollout undo and alerts the team
  • It retries forever

Answer: It automatically rolls back with kubectl rollout undo and alerts the team. On a failed rollout the script runs kubectl rollout undo to revert and sends an alert about the failure.

What is the recommended place to store secrets like API keys in automation scripts?

  • Hard-coded in the script
  • Environment variables or a secret manager
  • In the log files
  • In the Git history

Answer: Environment variables or a secret manager. Secrets belong in environment variables or secret managers, never hard-coded in code.

What does the '3-2-1 backup rule' mean?

  • 3 servers, 2 regions, 1 admin
  • 3 daily backups kept for 21 days
  • 3 scripts, 2 schedules, 1 alert
  • 3 copies of data, 2 different media types, 1 offsite copy

Answer: 3 copies of data, 2 different media types, 1 offsite copy. The 3-2-1 rule keeps 3 copies of data on 2 media types with 1 stored offsite.

Why does the lesson favor Python over bash scripts for DevOps automation?

  • Python runs only on Linux
  • Better error handling, cross-platform support, and clean native API integration
  • Bash cannot run commands
  • Python is always faster

Answer: Better error handling, cross-platform support, and clean native API integration. Python offers try/except error handling, works across platforms, and integrates with APIs via libraries like requests and boto3.

What does it mean for an automation script to be 'idempotent'?

  • It runs only once ever
  • It requires root access
  • It can run multiple times safely without causing harm
  • It never writes to disk

Answer: It can run multiple times safely without causing harm. An idempotent script produces the same safe result whether run once or many times, a key property of reliable automation.

Continue this course