Skip to content

BLAST+ 2.17.0 - User Guide

Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari) Date: October 12, 2025 Installation: /sw/blast/2.17.0/


πŸ“¦ Version Information

BLAST+: 2.17.0 Release Date: July 1, 2025 Type: Precompiled binaries (x86_64 Linux)


πŸš€ Loading the Module

module load blast/2.17.0

# Verify installation
blastn -version

πŸ” BLAST Programs

Core BLAST Programs

Program Description Query β†’ Database
blastn Nucleotide BLAST Nucleotide β†’ Nucleotide
blastp Protein BLAST Protein β†’ Protein
blastx Translated BLAST Translated Nucleotide β†’ Protein
tblastn Translated database Protein β†’ Translated Nucleotide
tblastx Both translated Translated Nucleotide β†’ Translated Nucleotide

Additional Tools

  • makeblastdb: Create BLAST databases from FASTA files
  • update_blastdb.pl: Download preformatted NCBI databases
  • blast_formatter: Reformat BLAST archive files
  • blastdbcmd: Extract sequences from BLAST databases

πŸ“– Basic Usage Examples

1. Nucleotide BLAST (blastn)

module load blast/2.17.0

# Simple query against database
blastn -query sequences.fasta -db nt -out results.txt

# With output format and E-value threshold
blastn -query sequences.fasta -db nt -evalue 1e-10 -outfmt 6 -out results.tsv

# With multiple threads
blastn -query sequences.fasta -db nt -num_threads 8 -out results.txt

2. Protein BLAST (blastp)

# Search protein sequence against protein database
blastp -query protein.fasta -db nr -evalue 1e-5 -out results.txt

# Tabular output with custom fields
blastp -query protein.fasta -db swissprot \
  -outfmt "6 qseqid sseqid pident length evalue bitscore stitle" \
  -out results.tsv

3. Translated BLAST (blastx)

# Translate nucleotide query, search against protein database
blastx -query nucleotide.fasta -db nr -evalue 1e-10 -out results.txt

πŸ—„οΈ Working with Databases

Download NCBI Databases

module load blast/2.17.0

# List available databases
update_blastdb.pl --showall

# Download a database (e.g., nt, nr, swissprot)
update_blastdb.pl --decompress nt

# Download to specific directory
update_blastdb.pl --decompress nr --blastdb_version 5

Note: Databases are large! Example sizes: - nt (nucleotide): ~100 GB - nr (non-redundant protein): ~250 GB - swissprot: ~500 MB

Create Custom Database

# From FASTA file
makeblastdb -in mysequences.fasta -dbtype nucl -out mydb

# For protein sequences
makeblastdb -in proteins.fasta -dbtype prot -out proteindb

# With taxonomy information
makeblastdb -in sequences.fasta -dbtype nucl -out mydb \
  -parse_seqids -taxid_map taxid.map

Set Database Location

# Option 1: Use environment variable (automatically set by module)
export BLASTDB=/sw/blast/2.17.0/db

# Option 2: Specify with -db flag
blastn -query seq.fasta -db /path/to/database/dbname -out results.txt

πŸ“Š Output Formats

Common Output Formats (-outfmt)

Format Description
0 Pairwise alignment (default)
6 Tabular (TSV)
7 Tabular with comment lines
10 CSV
11 BLAST archive (ASN.1)
17 SAM format
18 Organism report

Custom Tabular Output

# Specify custom columns
blastn -query seq.fasta -db nt \
  -outfmt "6 qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore" \
  -out custom_results.tsv

Available fields: - qseqid - Query sequence ID - sseqid - Subject sequence ID - pident - Percentage of identical matches - length - Alignment length - mismatch - Number of mismatches - gapopen - Number of gap openings - qstart qend - Query alignment start/end - sstart send - Subject alignment start/end - evalue - Expect value - bitscore - Bit score - stitle - Subject title


πŸ”§ Advanced Options

Filtering and Thresholds

# E-value threshold
blastn -query seq.fasta -db nt -evalue 1e-10

# Percent identity filter
blastn -query seq.fasta -db nt -perc_identity 90

# Query coverage filter
blastn -query seq.fasta -db nt -qcov_hsp_perc 80

# Maximum number of hits
blastn -query seq.fasta -db nt -max_target_seqs 10

Performance Tuning

# Use multiple threads
blastn -query seq.fasta -db nt -num_threads 16

# Word size (smaller = more sensitive, slower)
blastn -query seq.fasta -db nt -word_size 11

# Use best hit overhang for shorter sequences
blastn -query seq.fasta -db nt -best_hit_overhang 0.1 -best_hit_score_edge 0.1

πŸ“ Example Workflows

#!/bin/bash
module load blast/2.17.0

# 1. Create database from FASTA
makeblastdb -in reference.fasta -dbtype nucl -out refdb

# 2. Run BLAST search
blastn -query queries.fasta -db refdb \
  -outfmt "6 qseqid sseqid pident evalue bitscore" \
  -evalue 1e-10 -num_threads 8 \
  -out blast_results.tsv

# 3. Filter results (e.g., >95% identity)
awk '$3 >= 95' blast_results.tsv > high_identity_hits.tsv
# Search against NCBI nr database remotely
blastp -query protein.fasta -db nr -remote -out results.txt

# Note: Remote searches are slower but don't require local database

Workflow 3: Batch Processing

#!/bin/bash
module load blast/2.17.0

# Process multiple query files
for query in queries/*.fasta; do
  basename=$(basename $query .fasta)
  blastn -query $query -db nt \
    -outfmt 6 -num_threads 4 \
    -out results/${basename}_blast.tsv
done

🎯 BLAST Strategies

High Sensitivity (slow)

blastn -query seq.fasta -db nt \
  -task blastn \
  -word_size 11 \
  -evalue 1e-10 \
  -num_threads 16

High Speed (less sensitive)

blastn -query seq.fasta -db nt \
  -task megablast \
  -word_size 28 \
  -perc_identity 95 \
  -num_threads 16

Short Sequences (<50bp)

blastn -query short_seq.fasta -db nt \
  -task blastn-short \
  -word_size 7 \
  -evalue 1000

πŸ› Troubleshooting

"Database not found"

  • Check $BLASTDB environment variable: echo $BLASTDB
  • Verify database exists: ls $BLASTDB
  • Use full path to database: blastn -db /full/path/to/dbname

"Out of memory"

  • Reduce -num_threads
  • Use smaller database or split queries
  • Increase -max_target_seqs threshold

"Low complexity filter"

  • BLAST filters low-complexity regions by default
  • To disable: blastn -dust no (nucleotide) or blastp -seg no (protein)

Slow searches

  • Use megablast task for similar sequences
  • Increase word size (-word_size)
  • Use local database instead of remote

πŸ“š Documentation

Local: - Installation: /sw/blast/2.17.0/ - Binaries: /sw/blast/2.17.0/bin/ - Databases: /sw/blast/2.17.0/db/ (create with update_blastdb.pl)

Online: - BLAST+ Manual: https://www.ncbi.nlm.nih.gov/books/NBK279690/ - BLAST Web Interface: https://blast.ncbi.nlm.nih.gov/ - Database Download: https://ftp.ncbi.nlm.nih.gov/blast/db/ - BLAST Help: https://blast.ncbi.nlm.nih.gov/doc/blast-help/

Command-line help:

blastn -help
makeblastdb -help
update_blastdb.pl --help

πŸ“ Citation

Camacho C., Coulouris G., Avagyan V., Ma N., Papadopoulos J., Bealer K., Madden T.L. BLAST+: architecture and applications. BMC Bioinformatics 10, 421 (2009)


πŸ”„ Quick Reference

Task Command
Nucleotide vs nucleotide blastn -query seq.fasta -db nt
Protein vs protein blastp -query prot.fasta -db nr
Translated nucl vs protein blastx -query seq.fasta -db nr
Create nucleotide database makeblastdb -in seqs.fasta -dbtype nucl -out dbname
Create protein database makeblastdb -in prots.fasta -dbtype prot -out dbname
Download NCBI database update_blastdb.pl --decompress nt
Tabular output blastn -outfmt 6
XML output blastn -outfmt 5

For questions contact: Cluster administrators Installation: October 12, 2025 Installed by: Claude Code + Uliano Guerrini