Skip to content
Contact Support

Tutorial

Evaluating Job Efficiency

!!! time "wibbly wobbly timey wimey"

Objectives

  • review the resource utilisation of a previous job
  • identify areas for improvement in job efficiency

Prerequisites

This tutorial assumes familiarity with Mahuika and the use of HPCs. You should be able to:

  • use the terminal to navigate a filesystem
  • find and load software
  • write a SLURM batch script
  • submit batch jobs to SLURM
  • check on the status of queued and completed SLURM jobs

These tasks are covered in the Introduction to HPC tutorials.

Summary and Setup

This tutorial aims to give you hands on experience evaluating and improving job efficiency. To do that, we are going to work with an example workflow from genomics, but no genomics knowledge is needed for the tutorial. The example used here is based on materials from the Data Carpentry Data Wrangling and Processing for Genomics and Genomics Aotearoa Intro to Bash Scripting and HPC Scheduler workshops if you want more information. There are 3 broad steps in this workflow:

  1. Quality control
  2. Indexing the reference genome
  3. Alignment and variant calling

To begin, we will be working with some example scripts that can be found in /opt/nesi/examples/intermediate_hpc. You will want to have these files in your nobackup directory. Let's make a directory to work from and copy the files in.

cd /nesi/nobackup/<project_id>/
mkdir -p intermed_hpc_<username>
cd intermed_hpc_<username>
cp -r /opt/nesi/examples/intermediate_hpc .

Looking at a previous job and its efficiency

sacct

01_script.sl
#!/bin/bash -e

#SBATCH --job-name=intermed-hpc-01
#SBATCH --output=log/%x_%j.out
#SBATCH --error=log/%x_%j.err
#SBATCH --time=1:00:00
#SBATCH --mem=5G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --profile=task

module purge

# Quality control
module load FastQC/0.12.1

cd untrimmed_fastq
fastqc *.fastq* # fastqc can take multiple files as input, so we can run it on all fastq files in the current directory

mkdir -p ../trimmed
cd ../trimmed
module load Trimmomatic/0.39-Java-1.8.0_144

for infile in ../untrimmed_fastq/*_1.fastq.gz
do
    base=$(basename ${infile} _1.fastq.gz)
    trimmomatic PE ${infile} ../untrimmed_fastq/${base}_2.fastq.gz \
                    ${base}_1.trim.fastq.gz ${base}_1un.trim.fastq.gz \
                    ${base}_2.trim.fastq.gz ${base}_2un.trim.fastq.gz \
                    SLIDINGWINDOW:4:20 MINLEN:25 ILLUMINACLIP:/opt/nesi/CS400_centos7_bdw/Trimmomatic/0.39-Java-1.8.0_144/adapters/NexteraPE-PE.fa:2:40:15 
done


cd ..
mkdir -p results/sam results/bam results/bcf results/vcf
gunzip trimmed/*.fastq.gz

module load bwa-mem2/2.3-GCC-12.3.0
module load SAMtools/1.22-GCC-12.3.0
module load BCFtools/1.22-GCC-12.3.0
# Reference genome indexing
gunzip -k ref_genome/ecoli_rel606.fasta.gz
bwa-mem2 index ref_genome/ecoli_rel606.fasta
# Alignment and variant calling
for infile in trimmed/*_1.trim.fastq
do
    base=$(basename ${infile} _1.trim.fastq)
    bwa-mem2 mem -t 4 ref_genome/ecoli_rel606.fasta ${infile} trimmed/${base}_2.trim.fastq > results/sam/${base}.aligned.sam
    samtools view -S -b results/sam/${base}.aligned.sam > results/bam/${base}.aligned.bam
    samtools sort results/bam/${base}.aligned.bam -o results/bam/${base}.aligned.sorted.bam
    bcftools mpileup -O b -o results/bcf/${base}.bcf -f ref_genome/ecoli_rel606.fasta results/bam/${base}.aligned.sorted.bam
    bcftools call --ploidy 1 -m -v -o results/vcf/${base}.vcf results/bcf/${base}.bcf
    vcfutils.pl varFilter results/vcf/${base}.vcf > results/vcf/${base}.var.vcf
done

The example script 01_script.sl has already been run as a batch job and has the job ID 9296168. Let's take a look at the status of that job before we even start looking at the script and see what we can learn.

sacct -j 9296168
JobID           JobName    Elapsed     AveCPU     MinCPU   TotalCPU Al NT     MaxRSS      State 
------------ ---------- ---------- ---------- ---------- ---------- -- -- ---------- ---------- 
9296168      intermed-+   00:29:10                        29:56.822  8                COMPLETED 
9296168.bat+      batch   00:29:10   00:29:57   00:29:57  29:56.822  8  1   1439444K  COMPLETED 
9296168.ext+     extern   00:29:10                         00:00:00  8  1             COMPLETED

Changing what sacct shows you by default

format flags that might be helpful

stick this in bash profile

seff

The basic sacct tells us a little about the job: the runtime, how many CPUs were allocated, the final job state. But this information doesn't really help us evaluate how well the job ran. To look at the efficiency of our job, we can use the command seff (short for slurm efficiency):

seff 9296168
Job ID: 9296168
State: COMPLETED
Cores: 4
Job Wall-time:         12%  00:29:10 of 04:00:00 time limit
Avg CPU Utilisation:   26%  00:29:57 of 01:56:40 core-walltime
Peak Mem Utilisation:  27%  1.37 GB of 5.00 GB

seff gives us a lot of the same information but with a bit more context.

What resources would you request if resubmitting the job based only on this information?

You aren't working with a lot of information yet, but we can take a first stab at how to improve our job efficiency just by adjusting our requested resources.

Resources for next time

We can probably lower the memory and CPUs requested, but if we are doing so, it might be best to leave the job time alone so if the job runs a little slower it doesn't timeout.

This is a good first step, we noticed that we are requesting more resources than we need and can make some quick adjustments. But this assumes that the resources being used are fairly stable over the script. The CPU utilisation is just an average, so we don't know if there was a portion of the job that did use all the CPUs we allocated to the job.

Job profiling

SLURM has the ability to conduct 'profiling' on jobs being submitted. This stores extra data about the resources the job uses at various times throughout the job and can let you assess your efficiency in more detail. To enable profiling in a batch job, you need to add the following to your SLURM header:

#SBATCH --profile=task

Luckily for us, this was included in our script when it was run previously, so we can access this data about our job. We can run profile_plot 9296168 to produce a PNG with plots of the CPU, memory and I/O utilisation over the course of our job.

Profile plot for job ID 9296168

Now we can see things in a bit more detail.

What stands out?

What more can we learn from these plots? Does this change your thoughts on how to adjust the requested resources?

Reviewing our job script

Now let's actually take a look at what this job was running. Let's open 01_script.sl and poke around.

Looking at the script based on the job ID

If we add the flag -B to our sacct command, sacct will return the script that was called for the job we are interested in.

sacct -B -j 9296168

Ideally you remember what script was run, but this can be helpful if you aren't sure what you changed since you last ran a job or if you aren't sure which script was actually used.

As mentioned above, this script is doing 3 major steps which are indicated with comments in the script:

  1. Quality control
  2. Indexing the reference genome
  3. Alignment and variant calling

Splitting tasks by resource needs

Let's identify different sections of the script (01_script.sl) that have significantly different resource needs and split them into separate jobs with appropriate resource requests for each.

Updated scripts

There is no one correct answer here! But here is one option. Looking at the profile plot, we can try to identify the steps of the job.

Annotated profile plot for 9296168

There are two major sections in the profile plot, and one little blip in the middle that we can guess is the reference genome indexing occurring between the quality control and alignment/variant calling steps.

So we can create three scripts to better assign resources:

Quality control:01_script_a.sl requests reduced CPUs and time, but the same amount of memory.

#!/bin/bash -e

#SBATCH --job-name=intermed-hpc-01a
#SBATCH --output=log/%x_%j.out
#SBATCH --error=log/%x_%j.err
#SBATCH --time=00:30:00
#SBATCH --mem=5G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1
#SBATCH --profile=task

module purge
# Quality control
module load FastQC/0.12.1
cd untrimmed_fastq
fastqc *.fastq* # fastqc can take multiple files as input, so we can run it on all fastq files in the current directory

mkdir -p ../trimmed
cd ../trimmed
module load Trimmomatic/0.39-Java-1.8.0_144

for infile in ../untrimmed_fastq/*_1.fastq.gz
do
    base=$(basename ${infile} _1.fastq.gz)
    trimmomatic PE ${infile} ../untrimmed_fastq/${base}_2.fastq.gz \
                    ${base}_1.trim.fastq.gz ${base}_1un.trim.fastq.gz \
                    ${base}_2.trim.fastq.gz ${base}_2un.trim.fastq.gz \
                    SLIDINGWINDOW:4:20 MINLEN:25 ILLUMINACLIP:/opt/nesi/CS400_centos7_bdw/Trimmomatic/0.39-Java-1.8.0_144/adapters/NexteraPE-PE.fa:2:40:15 
done

Reference genome indexing: 01_script_b.sl requests less CPUs, memory, and time.

#!/bin/bash -e

#SBATCH --job-name=intermed-hpc-01b
#SBATCH --output=log/%x_%j.out
#SBATCH --error=log/%x_%j.err
#SBATCH --time=00:30:00
#SBATCH --mem=1G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1
#SBATCH --profile=task

module load bwa-mem2/2.3-GCC-12.3.0
# Reference genome indexing
gunzip -k ref_genome/ecoli_rel606.fasta.gz
bwa-mem2 index ref_genome/ecoli_rel606.fasta

Alignment and variant calling: 01_script_c.sl requests the same CPUs but less memory and time.

#!/bin/bash -e

#SBATCH --job-name=intermed-hpc-01c
#SBATCH --output=log/%x_%j.out
#SBATCH --error=log/%x_%j.err
#SBATCH --time=00:30:00
#SBATCH --mem=1G
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --profile=task

mkdir -p results/sam results/bam results/bcf results/vcf
gunzip trimmed/*.fastq.gz

module load bwa-mem2/2.3-GCC-12.3.0
module load SAMtools/1.22-GCC-12.3.0
module load BCFtools/1.22-GCC-12.3.0
# Alignment and variant calling
for infile in trimmed/*_1.trim.fastq
do
    base=$(basename ${infile} _1.trim.fastq)
    bwa-mem2 mem -t 4 ref_genome/ecoli_rel606.fasta ${infile} trimmed/${base}_2.trim.fastq > results/sam/${base}.aligned.sam
    samtools view -S -b results/sam/${base}.aligned.sam > results/bam/${base}.aligned.bam
    samtools sort results/bam/${base}.aligned.bam -o results/bam/${base}.aligned.sorted.bam
    bcftools mpileup -O b -o results/bcf/${base}.bcf -f ref_genome/ecoli_rel606.fasta results/bam/${base}.aligned.sorted.bam
    bcftools call --ploidy 1 -m -v -o results/vcf/${base}.vcf results/bcf/${base}.bcf
    vcfutils.pl varFilter results/vcf/${base}.vcf > results/vcf/${base}.var.vcf
done

Keypoints

  • There are multiple tools for evaluating job efficiency
  • Different processes take different resources

We've assessed our job on a general scale and made some changes without editing the script at all. The next step is to work on improving the efficiency of the script itself. There are many ways to approach improving our script, but let's start by looking at job arrays.