Skip to content

Job Scheduling with Slurm

The XLence cluster uses Slurm (Simple Linux Utility for Resource Management) as the workload manager for submitting and managing computational jobs.


Overview

Slurm is a job scheduler that: - Allocates resources (CPUs, GPUs, memory) on compute nodes - Queues jobs when resources are busy - Manages job priorities and fair resource sharing - Monitors job execution and resource usage


Basic Workflow

  1. Prepare your job: Create a job script or command
  2. Submit to Slurm: Use sbatch (batch jobs) or srun (interactive)
  3. Monitor job: Check status with squeue
  4. Retrieve results: Access output files when job completes

See the detailed guides: - Submitting Jobs - How to submit batch and interactive jobs - GPU Jobs - Running GPU-accelerated computations - Monitoring & Troubleshooting - Track and debug your jobs


Partitions

The cluster has two partitions:

normal (default)

  • 5 compute nodes
  • 10 CPUs per node
  • 128 GB RAM per node
  • 2 GPUs per node

ngs

  • 1 compute node
  • 18 CPUs
  • 128 GB RAM
  • 1 GPU

Specify partition in your job script:

#SBATCH --partition=normal  # or --partition=ngs

Local Conventions (Important!)

To maximize cluster utilization and ensure fair access to both CPU and GPU resources, we ask all users to follow these guidelines:

CPU Job Limits by Partition

Partition normal (10 CPUs, 2 GPUs per node): - When submitting CPU-only jobs, request maximum 8 CPUs per job - This leaves 2 CPUs available (1 per GPU) for GPU jobs

Partition ngs (18 CPUs, 1 GPU): - When submitting CPU-only jobs, request maximum 17 CPUs per job - This leaves 1 CPU available for GPU jobs

Why This Matters

Many modern GPU-accelerated programs (molecular dynamics, deep learning, etc.) require only 1 CPU core per GPU, as most computation happens on the GPU itself. By reserving 1 CPU per GPU on each node, we ensure that GPU jobs can run even when CPU resources are heavily utilized, maximizing overall cluster throughput.

Examples Following Conventions

# CPU job on normal partition (leaves room for 2 GPU jobs)
#SBATCH --partition=normal
#SBATCH --cpus-per-task=8      # Maximum recommended for CPU jobs

# CPU job on ngs partition (leaves room for 1 GPU job)
#SBATCH --partition=ngs
#SBATCH --cpus-per-task=17     # Maximum recommended for CPU jobs

# GPU jobs can use 1 CPU (most common case)
#SBATCH --partition=normal
#SBATCH --cpus-per-task=1
#SBATCH --gres=gpu:1

Resource Limits

The XLence cluster operates on a trust-based model without strict enforcement of time limits or resource quotas.

Our Philosophy

For over five years, our user community has successfully self-regulated resource usage through mutual respect and consideration. As long as users continue to:

  • Be mindful of resource requests
  • Consider other users' needs
  • Follow the local conventions (CPU limits)
  • Communicate about large or long-running jobs

No hard limits will be imposed.

What This Means

  • No time walls: Jobs can run as long as needed
  • No strict quotas: Request resources based on actual needs
  • No rigid limits: Flexibility for legitimate research requirements
  • Community-based: We trust users to be reasonable

If Problems Arise

If the community experiences resource conflicts or abuse, administrators may: 1. Contact users to discuss resource usage 2. Work with research groups to coordinate large jobs 3. Only as a last resort: implement technical limits

We prefer to maintain the current collaborative environment. The success of the past five years demonstrates that our community can manage shared resources responsibly without bureaucratic overhead.

Your Responsibility

  • Estimate resource needs realistically: Don't over-request "just in case"
  • Monitor your jobs: Cancel jobs that fail early to free resources
  • Communicate: If planning large resource usage, coordinate with administrators
  • Be considerate: Remember others are waiting for resources too

Additional Resources


Support

For help with job scheduling or Slurm issues:


Last Updated: October 2025