多彩编程 多彩编程MZPH · CODE BLOG
ARTICLE DETAIL

文章详情

深耕前端与后端开发技术的一线实战笔记与踩坑复盘。

How to Automate Daily Linux Health Checks with a Bash Script + Cron

How to Automate Daily Linux Health Checks with a Bash Script + Cron In today’s digital landscape, maintaining the health and reliability of Linux servers is critical for businesses and developers alike. Manual health checks—such as monitoring CPU usage, memory leaks, or disk space—are time-consuming, error-prone, and often overlooked. Automating these checks ensures early detection of issues, reduces downtime, and frees up valuable time for other tasks.In this guide, we’ll walk through creating aBash scriptto perform daily Linux health checks and scheduling it withCron(a time-based job scheduler) to run automatically. By the end, you’ll have a robust system to monitor key metrics like CPU, memory, disk space, and more—with alerts for critical thresholds.Feb 24, 20262026-02Table of Contents#PrerequisitesPlanning Your Health ChecksStep 1: Create the Bash Health Check ScriptScript StructureKey Health Check ComponentsStep 2: Test the Script ManuallyStep 3: Schedule the Script with CronTroubleshooting Common IssuesConclusionReferencesPrerequisites#Before getting started, ensure you have:A Linux system (Ubuntu, CentOS, Debian, etc.).Basic familiarity with the Linux command line (e.g.,cd,chmod,nano).Sudo privileges (to install tools and edit Cron jobs).Optional:sysstatpackage (for advanced metrics like disk I/O; install withsudo apt install sysstatorsudo yum install sysstat).Planning Your Health Checks#What metrics should you monitor? Focus on critical components that impact system stability:MetricWhy It MattersTools/CommandsCPU UsageHigh CPU can slow down applications.top,mpstat(fromsysstat)Memory UsageLow memory can cause crashes or swapping.free,vmstatDisk SpaceFull disks block writes (e.g., logs, databases).df -hSystem LoadMeasures pending processes (CPU, memory, I/O).uptimeCritical ProcessesEnsure essential services (e.g.,nginx,sshd) are running.ps,systemctlDisk I/OSlow disk I/O can bottleneck performance.iostat(fromsysstat)Network ConnectivityVerify internet/intranet access.pingStep 1: Create the Bash Health Check Script#Let’s build a script that logs these metrics and alerts on critical thresholds (e.g., 90% disk usage). We’ll store logs in/var/log/system-health/for easy access.Script Structure#The script will:Define variables (log path, thresholds).Create a log directory (if missing).Log a timestamp for each check.Run individual health checks (CPU, memory, etc.).Alert (log) if thresholds are exceeded.Key Health Check Components#Create a file namedsystem_health_check.shusingnanoor your preferred editor:#!/bin/bash # --------------------------# System Health Check Script# Author: Your Name# Date: [Insert Date]# Description: Automates daily health checks for Linux systems.# -------------------------- # --------------------------# Variables (Customize These!)# --------------------------LOG_DIR/var/log/system-healthLOG_FILE${LOG_DIR}/health_check_$(date %Y%m%d).logTHRESHOLD_CPU90 # CPU usage % (alert if exceeded)THRESHOLD_MEM90 # Memory usage % (alert if exceeded)THRESHOLD_DISK90 # Disk usage % (alert if exceeded)THRESHOLD_LOAD5 # System load (15-min avg; alert if exceeded)CRITICAL_PROCESSES(sshd nginx docker) # Add your critical servicesCHECK_NETWORK8.8.8.8 # Test connectivity to this IP (e.g., Google DNS) # --------------------------# Initialize Log Directory# --------------------------if [ ! -d $LOG_DIR ]; then sudo mkdir -p $LOG_DIR sudo chmod 755 $LOG_DIRfi # --------------------------# Log Header with Timestamp# --------------------------echo $LOG_FILEecho System Health Check - $(date %Y-%m-%d %H:%M:%S) $LOG_FILEecho $LOG_FILEecho $LOG_FILE # --------------------------# 1. CPU Usage Check# --------------------------echo CPU Usage Check $LOG_FILEcpu_usage$(top -bn1 | grep Cpu(s) | awk {print $2 $4}) # User System CPUecho Current CPU Usage: ${cpu_usage}% $LOG_FILE if (( $(echo $cpu_usage $THRESHOLD_CPU | bc -l) )); then echo ALERT: CPU usage exceeds threshold (${THRESHOLD_CPU}%)! $LOG_FILEfiecho $LOG_FILE # --------------------------# 2. Memory Usage Check# --------------------------echo Memory Usage Check $LOG_FILEmem_total$(free -m | awk /Mem:/ {print $2})mem_used$(free -m | awk /Mem:/ {print $3})mem_usage$(( (mem_used * 100) / mem_total )) # % used echo Total Memory: ${mem_total}MB | Used: ${mem_used}MB (${mem_usage}%) $LOG_FILE if [ $mem_usage -gt $THRESHOLD_MEM ]; then echo ALERT: Memory usage exceeds threshold (${THRESHOLD_MEM}%)! $LOG_FILEfiecho $LOG_FILE # --------------------------# 3. Disk Space Check (Root Partition)# --------------------------echo Disk Space Check $LOG_FILEdisk_usage$(df -h / | awk /\// {print $5} | sed s/%//) # % used on / echo Root Partition Usage: ${disk_usage}% $LOG_FILE if [ $disk_usage -gt $THRESHOLD_DISK ]; then echo ALERT: Disk space exceeds threshold (${THRESHOLD_DISK}%)! $LOG_FILEfiecho $LOG_FILE # --------------------------# 4. System Load Check (15-minute average)# --------------------------echo System Load Check $LOG_FILEload_avg$(uptime | awk -F load average: {print $2} | cut -d , -f3 | sed s/ //) echo 15-Minute Load Average: ${load_avg} $LOG_FILE if (( $(echo $load_avg $THRESHOLD_LOAD | bc -l) )); then echo ALERT: System load exceeds threshold (${THRESHOLD_LOAD})! $LOG_FILEfiecho $LOG_FILE # --------------------------# 5. Critical Processes Check# --------------------------echo Critical Processes Check $LOG_FILEfor process in ${CRITICAL_PROCESSES[]}; do if ! pgrep -x $process /dev/null; then echo ALERT: Critical process $process is NOT running! $LOG_FILE else echo Process $process is running. $LOG_FILE fidoneecho $LOG_FILE # --------------------------# 6. Disk I/O Check (Optional: Requires sysstat)# --------------------------echo Disk I/O Check $LOG_FILEif command -v iostat /dev/null; then iostat 1 2 | grep -A 1 Device $LOG_FILE # 2-second sampleelse echo iostat not found (install sysstat for Disk I/O metrics). $LOG_FILEfiecho $LOG_FILE # --------------------------# 7. Network Connectivity Check# --------------------------echo Network Connectivity Check $LOG_FILEif ping -c 2 -W 5 $CHECK_NETWORK /dev/null; then echo Network reachable to ${CHECK_NETWORK}. $LOG_FILEelse echo ALERT: Network unreachable to ${CHECK_NETWORK}! $LOG_FILEfiecho $LOG_FILE # --------------------------# Log Footer# --------------------------echo ---------------------------------------------- $LOG_FILEecho Check completed. Log saved to: $LOG_FILECustomization Tips#Thresholds: AdjustTHRESHOLD_CPU,THRESHOLD_MEM, etc., based on your system’s needs (e.g., a high-traffic server may tolerate 95% CPU).Critical Processes: UpdateCRITICAL_PROCESSESto include services likemysql,apache2, or custom apps.Network Check: Replace8.8.8.8with your internal gateway or a critical service IP.Step 2: Test the Script Manually#Before scheduling with Cron, test the script to ensure it works:Make the script executable:chmod x system_health_check.shRun it manually:sudo ./system_health_check.shVerify the log file (e.g.,/var/log/system-health/health_check_20240520.log). It should include all checks and alerts if thresholds are exceeded.Step 3: Schedule the Script with Cron#Cron lets you run the script daily (or at any interval). Here’s how to set it up:Edit the Crontab#Open the crontab editor for the root user (to ensure full system access):sudo crontab -eAdd a line to run the script daily at 2:00 AM (adjust the time as needed):0 2 * * * /path/to/system_health_check.sh /var/log/system-health/cron.log 210 2 * * *: Runs at 2:00 AM every day./path/to/system_health_check.sh: Replace with the full path to your script (e.g.,/home/user/scripts/system_health_check.sh). /var/log/system-health/cron.log 21: Logs Cron output (for debugging).Verify Cron is Running#Ensure the Cron service is active:sudo systemctl status cron # Ubuntu/Debian# orsudo systemctl status crond # CentOS/RHELIf inactive, start it with:sudo systemctl start cron sudo systemctl enable cronTroubleshooting Common Issues#Script Not Running:Check permissions: Ensure the script is executable (chmod x).Use absolute paths in the script (e.g.,/usr/bin/topinstead oftop) to avoid Cron’s limitedPATH.Logs Not Generating:Verify theLOG_DIRexists and has write permissions (runsudo chmod 755 /var/log/system-health).False Alerts:Adjust thresholds (e.g., if memory usage spikes temporarily, increaseTHRESHOLD_MEM).Discover moreScriptsLinuxnetworkingConclusion#By automating daily Linux health checks with a Bash script and Cron, you’ve built a proactive monitoring system. This setup ensures you catch issues like high CPU, low disk space, or failed services before they cause downtime.For advanced use cases, extend the script to send email/SMS alerts (usingmailorcurl), integrate with monitoring tools like Prometheus, or add more metrics (e.g., GPU usage, user logins).References#Bash Scripting GuideCron Documentationsysstat Tools (iostat, mpstat)Linux Performance Monitoring Commands
返回列表