BLAST+ 2.17.0 - User Guide
Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari)
Date: October 12, 2025
Installation: /sw/blast/2.17.0/
π¦ Version Information
BLAST+: 2.17.0 Release Date: July 1, 2025 Type: Precompiled binaries (x86_64 Linux)
π Loading the Module
module load blast/2.17.0
# Verify installation
blastn -version
π BLAST Programs
Core BLAST Programs
| Program | Description | Query β Database |
|---|---|---|
| blastn | Nucleotide BLAST | Nucleotide β Nucleotide |
| blastp | Protein BLAST | Protein β Protein |
| blastx | Translated BLAST | Translated Nucleotide β Protein |
| tblastn | Translated database | Protein β Translated Nucleotide |
| tblastx | Both translated | Translated Nucleotide β Translated Nucleotide |
Additional Tools
- makeblastdb: Create BLAST databases from FASTA files
- update_blastdb.pl: Download preformatted NCBI databases
- blast_formatter: Reformat BLAST archive files
- blastdbcmd: Extract sequences from BLAST databases
π Basic Usage Examples
1. Nucleotide BLAST (blastn)
module load blast/2.17.0
# Simple query against database
blastn -query sequences.fasta -db nt -out results.txt
# With output format and E-value threshold
blastn -query sequences.fasta -db nt -evalue 1e-10 -outfmt 6 -out results.tsv
# With multiple threads
blastn -query sequences.fasta -db nt -num_threads 8 -out results.txt
2. Protein BLAST (blastp)
# Search protein sequence against protein database
blastp -query protein.fasta -db nr -evalue 1e-5 -out results.txt
# Tabular output with custom fields
blastp -query protein.fasta -db swissprot \
-outfmt "6 qseqid sseqid pident length evalue bitscore stitle" \
-out results.tsv
3. Translated BLAST (blastx)
# Translate nucleotide query, search against protein database
blastx -query nucleotide.fasta -db nr -evalue 1e-10 -out results.txt
ποΈ Working with Databases
Download NCBI Databases
module load blast/2.17.0
# List available databases
update_blastdb.pl --showall
# Download a database (e.g., nt, nr, swissprot)
update_blastdb.pl --decompress nt
# Download to specific directory
update_blastdb.pl --decompress nr --blastdb_version 5
Note: Databases are large! Example sizes: - nt (nucleotide): ~100 GB - nr (non-redundant protein): ~250 GB - swissprot: ~500 MB
Create Custom Database
# From FASTA file
makeblastdb -in mysequences.fasta -dbtype nucl -out mydb
# For protein sequences
makeblastdb -in proteins.fasta -dbtype prot -out proteindb
# With taxonomy information
makeblastdb -in sequences.fasta -dbtype nucl -out mydb \
-parse_seqids -taxid_map taxid.map
Set Database Location
# Option 1: Use environment variable (automatically set by module)
export BLASTDB=/sw/blast/2.17.0/db
# Option 2: Specify with -db flag
blastn -query seq.fasta -db /path/to/database/dbname -out results.txt
π Output Formats
Common Output Formats (-outfmt)
| Format | Description |
|---|---|
0 |
Pairwise alignment (default) |
6 |
Tabular (TSV) |
7 |
Tabular with comment lines |
10 |
CSV |
11 |
BLAST archive (ASN.1) |
17 |
SAM format |
18 |
Organism report |
Custom Tabular Output
# Specify custom columns
blastn -query seq.fasta -db nt \
-outfmt "6 qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore" \
-out custom_results.tsv
Available fields:
- qseqid - Query sequence ID
- sseqid - Subject sequence ID
- pident - Percentage of identical matches
- length - Alignment length
- mismatch - Number of mismatches
- gapopen - Number of gap openings
- qstart qend - Query alignment start/end
- sstart send - Subject alignment start/end
- evalue - Expect value
- bitscore - Bit score
- stitle - Subject title
π§ Advanced Options
Filtering and Thresholds
# E-value threshold
blastn -query seq.fasta -db nt -evalue 1e-10
# Percent identity filter
blastn -query seq.fasta -db nt -perc_identity 90
# Query coverage filter
blastn -query seq.fasta -db nt -qcov_hsp_perc 80
# Maximum number of hits
blastn -query seq.fasta -db nt -max_target_seqs 10
Performance Tuning
# Use multiple threads
blastn -query seq.fasta -db nt -num_threads 16
# Word size (smaller = more sensitive, slower)
blastn -query seq.fasta -db nt -word_size 11
# Use best hit overhang for shorter sequences
blastn -query seq.fasta -db nt -best_hit_overhang 0.1 -best_hit_score_edge 0.1
π Example Workflows
Workflow 1: Local Database Search
#!/bin/bash
module load blast/2.17.0
# 1. Create database from FASTA
makeblastdb -in reference.fasta -dbtype nucl -out refdb
# 2. Run BLAST search
blastn -query queries.fasta -db refdb \
-outfmt "6 qseqid sseqid pident evalue bitscore" \
-evalue 1e-10 -num_threads 8 \
-out blast_results.tsv
# 3. Filter results (e.g., >95% identity)
awk '$3 >= 95' blast_results.tsv > high_identity_hits.tsv
Workflow 2: Remote BLAST Search
# Search against NCBI nr database remotely
blastp -query protein.fasta -db nr -remote -out results.txt
# Note: Remote searches are slower but don't require local database
Workflow 3: Batch Processing
#!/bin/bash
module load blast/2.17.0
# Process multiple query files
for query in queries/*.fasta; do
basename=$(basename $query .fasta)
blastn -query $query -db nt \
-outfmt 6 -num_threads 4 \
-out results/${basename}_blast.tsv
done
π― BLAST Strategies
High Sensitivity (slow)
blastn -query seq.fasta -db nt \
-task blastn \
-word_size 11 \
-evalue 1e-10 \
-num_threads 16
High Speed (less sensitive)
blastn -query seq.fasta -db nt \
-task megablast \
-word_size 28 \
-perc_identity 95 \
-num_threads 16
Short Sequences (<50bp)
blastn -query short_seq.fasta -db nt \
-task blastn-short \
-word_size 7 \
-evalue 1000
π Troubleshooting
"Database not found"
- Check
$BLASTDBenvironment variable:echo $BLASTDB - Verify database exists:
ls $BLASTDB - Use full path to database:
blastn -db /full/path/to/dbname
"Out of memory"
- Reduce
-num_threads - Use smaller database or split queries
- Increase
-max_target_seqsthreshold
"Low complexity filter"
- BLAST filters low-complexity regions by default
- To disable:
blastn -dust no(nucleotide) orblastp -seg no(protein)
Slow searches
- Use
megablasttask for similar sequences - Increase word size (
-word_size) - Use local database instead of remote
π Documentation
Local:
- Installation: /sw/blast/2.17.0/
- Binaries: /sw/blast/2.17.0/bin/
- Databases: /sw/blast/2.17.0/db/ (create with update_blastdb.pl)
Online: - BLAST+ Manual: https://www.ncbi.nlm.nih.gov/books/NBK279690/ - BLAST Web Interface: https://blast.ncbi.nlm.nih.gov/ - Database Download: https://ftp.ncbi.nlm.nih.gov/blast/db/ - BLAST Help: https://blast.ncbi.nlm.nih.gov/doc/blast-help/
Command-line help:
blastn -help
makeblastdb -help
update_blastdb.pl --help
π Citation
Camacho C., Coulouris G., Avagyan V., Ma N., Papadopoulos J., Bealer K., Madden T.L. BLAST+: architecture and applications. BMC Bioinformatics 10, 421 (2009)
π Quick Reference
| Task | Command |
|---|---|
| Nucleotide vs nucleotide | blastn -query seq.fasta -db nt |
| Protein vs protein | blastp -query prot.fasta -db nr |
| Translated nucl vs protein | blastx -query seq.fasta -db nr |
| Create nucleotide database | makeblastdb -in seqs.fasta -dbtype nucl -out dbname |
| Create protein database | makeblastdb -in prots.fasta -dbtype prot -out dbname |
| Download NCBI database | update_blastdb.pl --decompress nt |
| Tabular output | blastn -outfmt 6 |
| XML output | blastn -outfmt 5 |
For questions contact: Cluster administrators Installation: October 12, 2025 Installed by: Claude Code + Uliano Guerrini