Skip to content

Slurm Command Reference

Quick reference guide for commonly used Slurm commands.


Job Submission

Submit Batch Job

# Submit job script
sbatch job.sh

# Submit with command-line options (override script settings)
sbatch --partition=ngs --cpus-per-task=8 job.sh

# Submit job array
sbatch --array=1-100 job.sh

Interactive Session

# Request interactive session
srun --partition=normal --cpus-per-task=4 --mem=16G --pty bash

# Interactive session with GPU
srun --partition=normal --cpus-per-task=1 --gres=gpu:1 --mem=16G --pty bash

# Interactive session with time limit
srun --partition=normal --cpus-per-task=4 --time=2:00:00 --pty bash

Monitoring Jobs

View Jobs in Queue

# Show your jobs
squeue -u $USER

# Show your jobs with detailed format
squeue -u $USER --long

# Show all jobs
squeue

# Show jobs in specific partition
squeue --partition=normal
squeue --partition=ngs

# Show only running jobs
squeue -u $USER --state=RUNNING

# Show only pending jobs
squeue -u $USER --state=PENDING

# Custom format
squeue -u $USER --format="%.18i %.9P %.30j %.8T %.10M %.6D %R"

Job Details

# Show detailed job information
scontrol show job JOBID

# Show information for all your jobs
scontrol show job -u $USER

# Show node information
scontrol show node node01

Watch Queue in Real-Time

# Update every 2 seconds
watch -n 2 squeue -u $USER

# More compact view
watch -n 2 'squeue -u $USER --format="%.10i %.9P %.20j %.8T %.10M %.6D %R"'

Job Management

Cancel Jobs

# Cancel specific job
scancel JOBID

# Cancel all your jobs
scancel -u $USER

# Cancel all pending jobs
scancel -u $USER --state=PENDING

# Cancel jobs by name
scancel --name=job_name

# Cancel job array
scancel JOBID_[1-100]

# Cancel specific tasks in array
scancel JOBID_[10-20]

Hold and Release

# Hold (prevent from running)
scontrol hold JOBID

# Release (allow to run)
scontrol release JOBID

# Hold all your pending jobs
scontrol hold -u $USER

# Release all your held jobs
scontrol release -u $USER

Modify Pending Job

# Change time limit (only for pending jobs)
scontrol update job=JOBID TimeLimit=48:00:00

# Change partition
scontrol update job=JOBID Partition=ngs

# Change job name
scontrol update job=JOBID JobName=new_name

Job History and Accounting

Recent Jobs

# Show your jobs from today
sacct -u $USER

# Show jobs from last 7 days
sacct -u $USER --starttime=now-7days

# Show jobs from specific date
sacct -u $USER --starttime=2025-10-01

# Show specific job
sacct -j JOBID

# Show job array tasks
sacct -j JOBID --format=JobID,JobName,State,Elapsed,MaxRSS

Job Statistics

# Standard format with useful info
sacct -j JOBID --format=JobID,JobName,Partition,State,Elapsed,CPUTime,MaxRSS,ReqMem

# Memory usage
sacct -j JOBID --format=JobID,JobName,State,MaxRSS,MaxVMSize,ReqMem

# CPU usage
sacct -j JOBID --format=JobID,JobName,State,Elapsed,TotalCPU,CPUTime,ReqCPUS

# Full details
sacct -j JOBID --long

# Array job summary
sacct -j JOBID --format=JobID,JobName,State,Elapsed,MaxRSS --allocations

Custom sacct Formats

# Compact summary
sacct --format=JobID%15,JobName%20,State%12,Elapsed%12,MaxRSS%12

# With node info
sacct --format=JobID,JobName,NodeList,State,Elapsed,MaxRSS

# Time-focused
sacct --format=JobID,JobName,Submit,Start,End,Elapsed,State

Cluster Information

Partition Information

# Show all partitions
sinfo

# Show specific partition
sinfo --partition=normal

# Detailed partition info
scontrol show partition normal
scontrol show partition ngs

# Show nodes by partition
sinfo -N --partition=normal

Node Information

# Show all nodes
sinfo -N

# Show node details
scontrol show node node01

# Show available resources
sinfo -o "%N %C %m %G"
# Output: NodeName CPUs(A/I/O/T) Memory GPUs

GPU Information

# Show GPU availability
sinfo -o "%N %G"

# Show nodes with free GPUs
sinfo -N --state=idle -o "%N %G"

Monitoring Output

Watch Job Output Files

# Follow output file
tail -f job_12345.out

# Follow error file
tail -f job_12345.err

# Show last 50 lines and follow
tail -n 50 -f job_12345.out

# Watch both output and error
tail -f job_12345.out job_12345.err

Check Job Progress

# Show recent lines from output
tail -n 20 job_12345.out

# Search output for specific text
grep "COMPLETED" job_12345.out

# Count lines in output (rough progress indicator)
wc -l job_12345.out

Useful One-Liners

Job Submission

# Quick CPU job
sbatch --partition=normal --cpus-per-task=4 --mem=16G --time=24:00:00 --wrap="your_command"

# Quick GPU job
sbatch --partition=normal --cpus-per-task=1 --gres=gpu:1 --mem=16G --time=24:00:00 --wrap="your_gpu_command"

# Submit multiple similar jobs
for i in {1..10}; do sbatch --export=ALL,ID=$i job.sh; done

Monitoring

# Count your running jobs
squeue -u $USER --state=RUNNING | wc -l

# Count your pending jobs
squeue -u $USER --state=PENDING | wc -l

# Show jobs and their memory usage
squeue -u $USER --format="%.10i %.20j %.10m %.6D %R"

# Find job by name
squeue -u $USER --name=job_name

Analysis

# List your completed jobs from today
sacct -u $USER --state=COMPLETED --starttime=today --format=JobID,JobName,Elapsed

# List your failed jobs from last week
sacct -u $USER --state=FAILED --starttime=now-7days --format=JobID,JobName,State,ExitCode

# Calculate total CPU time used today
sacct -u $USER --starttime=today --format=CPUTime --noheader | awk '{total+=$1} END {print total}'

Environment Variables in Jobs

When your job runs, Slurm sets several useful environment variables:

#!/bin/bash
#SBATCH ...

# Job information
echo "Job ID: $SLURM_JOB_ID"
echo "Job name: $SLURM_JOB_NAME"
echo "Node: $SLURM_NODELIST"
echo "Partition: $SLURM_JOB_PARTITION"

# Resource allocation
echo "CPUs: $SLURM_CPUS_PER_TASK"
echo "Memory: $SLURM_MEM_PER_NODE"
echo "GPUs: $SLURM_GPUS"

# Working directory
echo "Submit dir: $SLURM_SUBMIT_DIR"
cd $SLURM_SUBMIT_DIR

# For job arrays
echo "Array job ID: $SLURM_ARRAY_JOB_ID"
echo "Array task ID: $SLURM_ARRAY_TASK_ID"

Common Environment Variables

Variable Description
$SLURM_JOB_ID Job ID
$SLURM_JOB_NAME Job name
$SLURM_SUBMIT_DIR Directory where job was submitted
$SLURM_NODELIST List of nodes allocated
$SLURM_CPUS_PER_TASK CPUs per task
$SLURM_MEM_PER_NODE Memory per node (MB)
$SLURM_GPUS Number of GPUs
$SLURM_ARRAY_JOB_ID Array job ID
$SLURM_ARRAY_TASK_ID Array task ID
$SLURM_JOB_PARTITION Partition name

Useful Aliases

Add these to your ~/.bashrc for convenience:

# Job monitoring
alias sq='squeue -u $USER'
alias sqw='watch -n 2 squeue -u $USER'
alias sqa='squeue'

# Job details
alias sj='scontrol show job'
alias sa='sacct -u $USER --starttime=now-7days'

# Job management
alias sc='scancel'
alias scall='scancel -u $USER'

# Cluster info
alias si='sinfo'
alias sin='sinfo -N'

After adding to ~/.bashrc, reload:

source ~/.bashrc

Tips and Tricks

1. Use Job Dependencies

Run jobs in sequence:

# Submit first job
job1=$(sbatch --parsable job1.sh)

# Submit second job that depends on first
sbatch --dependency=afterok:$job1 job2.sh

# Alternative: after any completion (ok or failed)
sbatch --dependency=afterany:$job1 job2.sh

2. Email Notifications

Get notified when jobs complete:

#SBATCH --mail-type=END,FAIL
#SBATCH --mail-user=your.email@unimi.it

3. Test Jobs Quickly

Use short time limits for testing:

# Test with 10 minute limit
sbatch --time=00:10:00 --partition=normal test_job.sh

4. Find Your Job Files

# Find all your job output files from last 7 days
find . -name "job_*.out" -mtime -7 -user $USER

# Find job error files with content (potential errors)
find . -name "job_*.err" -size +0 -mtime -7

5. Job Script Templates

Keep template scripts for common job types:

~/job_templates/
├── cpu_job.sh
├── gpu_job.sh
├── array_job.sh
└── interactive.sh

Copy and modify as needed:

cp ~/job_templates/gpu_job.sh my_analysis.sh

Common Error Exit Codes

When jobs fail, sacct shows exit codes:

Exit Code Meaning
0 Success
1 General error
2 Misuse of shell command
126 Command cannot execute
127 Command not found
137 Job killed (usually OOM)
143 Job terminated (SIGTERM)

Check exit code:

sacct -j JOBID --format=JobID,State,ExitCode

Additional Resources


Support

For help with Slurm commands:


Last Updated: October 2025