Slurm Command Reference
Quick reference guide for commonly used Slurm commands.
Job Submission
Submit Batch Job
# Submit job script
sbatch job.sh
# Submit with command-line options (override script settings)
sbatch --partition=ngs --cpus-per-task=8 job.sh
# Submit job array
sbatch --array=1-100 job.sh
Interactive Session
# Request interactive session
srun --partition=normal --cpus-per-task=4 --mem=16G --pty bash
# Interactive session with GPU
srun --partition=normal --cpus-per-task=1 --gres=gpu:1 --mem=16G --pty bash
# Interactive session with time limit
srun --partition=normal --cpus-per-task=4 --time=2:00:00 --pty bash
Monitoring Jobs
View Jobs in Queue
# Show your jobs
squeue -u $USER
# Show your jobs with detailed format
squeue -u $USER --long
# Show all jobs
squeue
# Show jobs in specific partition
squeue --partition=normal
squeue --partition=ngs
# Show only running jobs
squeue -u $USER --state=RUNNING
# Show only pending jobs
squeue -u $USER --state=PENDING
# Custom format
squeue -u $USER --format="%.18i %.9P %.30j %.8T %.10M %.6D %R"
Job Details
# Show detailed job information
scontrol show job JOBID
# Show information for all your jobs
scontrol show job -u $USER
# Show node information
scontrol show node node01
Watch Queue in Real-Time
# Update every 2 seconds
watch -n 2 squeue -u $USER
# More compact view
watch -n 2 'squeue -u $USER --format="%.10i %.9P %.20j %.8T %.10M %.6D %R"'
Job Management
Cancel Jobs
# Cancel specific job
scancel JOBID
# Cancel all your jobs
scancel -u $USER
# Cancel all pending jobs
scancel -u $USER --state=PENDING
# Cancel jobs by name
scancel --name=job_name
# Cancel job array
scancel JOBID_[1-100]
# Cancel specific tasks in array
scancel JOBID_[10-20]
Hold and Release
# Hold (prevent from running)
scontrol hold JOBID
# Release (allow to run)
scontrol release JOBID
# Hold all your pending jobs
scontrol hold -u $USER
# Release all your held jobs
scontrol release -u $USER
Modify Pending Job
# Change time limit (only for pending jobs)
scontrol update job=JOBID TimeLimit=48:00:00
# Change partition
scontrol update job=JOBID Partition=ngs
# Change job name
scontrol update job=JOBID JobName=new_name
Job History and Accounting
Recent Jobs
# Show your jobs from today
sacct -u $USER
# Show jobs from last 7 days
sacct -u $USER --starttime=now-7days
# Show jobs from specific date
sacct -u $USER --starttime=2025-10-01
# Show specific job
sacct -j JOBID
# Show job array tasks
sacct -j JOBID --format=JobID,JobName,State,Elapsed,MaxRSS
Job Statistics
# Standard format with useful info
sacct -j JOBID --format=JobID,JobName,Partition,State,Elapsed,CPUTime,MaxRSS,ReqMem
# Memory usage
sacct -j JOBID --format=JobID,JobName,State,MaxRSS,MaxVMSize,ReqMem
# CPU usage
sacct -j JOBID --format=JobID,JobName,State,Elapsed,TotalCPU,CPUTime,ReqCPUS
# Full details
sacct -j JOBID --long
# Array job summary
sacct -j JOBID --format=JobID,JobName,State,Elapsed,MaxRSS --allocations
Custom sacct Formats
# Compact summary
sacct --format=JobID%15,JobName%20,State%12,Elapsed%12,MaxRSS%12
# With node info
sacct --format=JobID,JobName,NodeList,State,Elapsed,MaxRSS
# Time-focused
sacct --format=JobID,JobName,Submit,Start,End,Elapsed,State
Cluster Information
Partition Information
# Show all partitions
sinfo
# Show specific partition
sinfo --partition=normal
# Detailed partition info
scontrol show partition normal
scontrol show partition ngs
# Show nodes by partition
sinfo -N --partition=normal
Node Information
# Show all nodes
sinfo -N
# Show node details
scontrol show node node01
# Show available resources
sinfo -o "%N %C %m %G"
# Output: NodeName CPUs(A/I/O/T) Memory GPUs
GPU Information
# Show GPU availability
sinfo -o "%N %G"
# Show nodes with free GPUs
sinfo -N --state=idle -o "%N %G"
Monitoring Output
Watch Job Output Files
# Follow output file
tail -f job_12345.out
# Follow error file
tail -f job_12345.err
# Show last 50 lines and follow
tail -n 50 -f job_12345.out
# Watch both output and error
tail -f job_12345.out job_12345.err
Check Job Progress
# Show recent lines from output
tail -n 20 job_12345.out
# Search output for specific text
grep "COMPLETED" job_12345.out
# Count lines in output (rough progress indicator)
wc -l job_12345.out
Useful One-Liners
Job Submission
# Quick CPU job
sbatch --partition=normal --cpus-per-task=4 --mem=16G --time=24:00:00 --wrap="your_command"
# Quick GPU job
sbatch --partition=normal --cpus-per-task=1 --gres=gpu:1 --mem=16G --time=24:00:00 --wrap="your_gpu_command"
# Submit multiple similar jobs
for i in {1..10}; do sbatch --export=ALL,ID=$i job.sh; done
Monitoring
# Count your running jobs
squeue -u $USER --state=RUNNING | wc -l
# Count your pending jobs
squeue -u $USER --state=PENDING | wc -l
# Show jobs and their memory usage
squeue -u $USER --format="%.10i %.20j %.10m %.6D %R"
# Find job by name
squeue -u $USER --name=job_name
Analysis
# List your completed jobs from today
sacct -u $USER --state=COMPLETED --starttime=today --format=JobID,JobName,Elapsed
# List your failed jobs from last week
sacct -u $USER --state=FAILED --starttime=now-7days --format=JobID,JobName,State,ExitCode
# Calculate total CPU time used today
sacct -u $USER --starttime=today --format=CPUTime --noheader | awk '{total+=$1} END {print total}'
Environment Variables in Jobs
When your job runs, Slurm sets several useful environment variables:
#!/bin/bash
#SBATCH ...
# Job information
echo "Job ID: $SLURM_JOB_ID"
echo "Job name: $SLURM_JOB_NAME"
echo "Node: $SLURM_NODELIST"
echo "Partition: $SLURM_JOB_PARTITION"
# Resource allocation
echo "CPUs: $SLURM_CPUS_PER_TASK"
echo "Memory: $SLURM_MEM_PER_NODE"
echo "GPUs: $SLURM_GPUS"
# Working directory
echo "Submit dir: $SLURM_SUBMIT_DIR"
cd $SLURM_SUBMIT_DIR
# For job arrays
echo "Array job ID: $SLURM_ARRAY_JOB_ID"
echo "Array task ID: $SLURM_ARRAY_TASK_ID"
Common Environment Variables
| Variable | Description |
|---|---|
$SLURM_JOB_ID |
Job ID |
$SLURM_JOB_NAME |
Job name |
$SLURM_SUBMIT_DIR |
Directory where job was submitted |
$SLURM_NODELIST |
List of nodes allocated |
$SLURM_CPUS_PER_TASK |
CPUs per task |
$SLURM_MEM_PER_NODE |
Memory per node (MB) |
$SLURM_GPUS |
Number of GPUs |
$SLURM_ARRAY_JOB_ID |
Array job ID |
$SLURM_ARRAY_TASK_ID |
Array task ID |
$SLURM_JOB_PARTITION |
Partition name |
Useful Aliases
Add these to your ~/.bashrc for convenience:
# Job monitoring
alias sq='squeue -u $USER'
alias sqw='watch -n 2 squeue -u $USER'
alias sqa='squeue'
# Job details
alias sj='scontrol show job'
alias sa='sacct -u $USER --starttime=now-7days'
# Job management
alias sc='scancel'
alias scall='scancel -u $USER'
# Cluster info
alias si='sinfo'
alias sin='sinfo -N'
After adding to ~/.bashrc, reload:
source ~/.bashrc
Tips and Tricks
1. Use Job Dependencies
Run jobs in sequence:
# Submit first job
job1=$(sbatch --parsable job1.sh)
# Submit second job that depends on first
sbatch --dependency=afterok:$job1 job2.sh
# Alternative: after any completion (ok or failed)
sbatch --dependency=afterany:$job1 job2.sh
2. Email Notifications
Get notified when jobs complete:
#SBATCH --mail-type=END,FAIL
#SBATCH --mail-user=your.email@unimi.it
3. Test Jobs Quickly
Use short time limits for testing:
# Test with 10 minute limit
sbatch --time=00:10:00 --partition=normal test_job.sh
4. Find Your Job Files
# Find all your job output files from last 7 days
find . -name "job_*.out" -mtime -7 -user $USER
# Find job error files with content (potential errors)
find . -name "job_*.err" -size +0 -mtime -7
5. Job Script Templates
Keep template scripts for common job types:
~/job_templates/
├── cpu_job.sh
├── gpu_job.sh
├── array_job.sh
└── interactive.sh
Copy and modify as needed:
cp ~/job_templates/gpu_job.sh my_analysis.sh
Common Error Exit Codes
When jobs fail, sacct shows exit codes:
| Exit Code | Meaning |
|---|---|
| 0 | Success |
| 1 | General error |
| 2 | Misuse of shell command |
| 126 | Command cannot execute |
| 127 | Command not found |
| 137 | Job killed (usually OOM) |
| 143 | Job terminated (SIGTERM) |
Check exit code:
sacct -j JOBID --format=JobID,State,ExitCode
Additional Resources
- Submitting Jobs - Detailed job submission guide
- GPU Jobs - GPU-specific commands and examples
- Monitoring & Troubleshooting - Detailed troubleshooting
- Official Slurm documentation: https://slurm.schedmd.com/
Support
For help with Slurm commands:
- Uliano Guerrini: uliano.guerrini@unimi.it
- Omar Ben Mariem: omar.benmariem@unimi.it
Last Updated: October 2025