InterProScan 5.75-106.0 - User Guide
Cluster: XLence (UNIMI Dipartimento di Scienze Farmacologiche e Biomolecolari)
Date: October 12, 2025
Installation: /sw/interproscan/5.75-106.0/
π¦ Version Information
- Version: 5.75-106.0
- Build: 64-bit (requires Java 11+)
- Release Date: June 2024
- Size: ~6.7 GB (includes all databases)
π§ Quick Start
Load the Module
module load interproscan
interproscan.sh --version
Basic Usage
# Analyze protein sequences
interproscan.sh -i my_proteins.fasta -f tsv -o results.tsv
# Multiple output formats
interproscan.sh -i sequences.fasta -f tsv,gff3,json -o output_prefix
# Specific applications only (faster)
interproscan.sh -i proteins.fasta -appl Pfam,SMART -f tsv -o pfam_smart.tsv
π Available Databases & Applications
InterProScan integrates these protein signature databases:
| Database | Type | Description |
|---|---|---|
| Pfam | HMM | Protein families database |
| PRINTS | Fingerprint | Protein motif fingerprints |
| ProDom | Domain | Protein domain database |
| SMART | HMM | Simple Modular Architecture Research Tool |
| TIGRFAMs | HMM | Protein families for prokaryotes |
| PIRSF | HMM | Protein classification system |
| SUPERFAMILY | HMM | Structural assignments for proteins |
| Gene3D | HMM | CATH protein domain predictions |
| HAMAP | Profile | High-quality Automated Annotation of Microbial Proteomes |
| ProSitePatterns | Regex | Protein domains, families and sites |
| ProSiteProfiles | Profile | Generalized profiles |
| Coils | Prediction | Coiled coil regions |
| MobiDBLite | Prediction | Protein disorder and mobility |
| SFLD | HMM | Structure-Function Linkage Database |
π‘ Common Use Cases
1. Standard Protein Annotation
# Run all available analyses
interproscan.sh -i input.fasta -f tsv,gff3,html -o annotation
# Output files:
# - annotation.tsv (Tab-separated values)
# - annotation.gff3 (GFF3 genomic format)
# - annotation.html.tar.gz (Visual HTML report)
2. Domain Architecture Analysis
# Focus on domain databases
interproscan.sh -i proteins.fasta \
-appl Pfam,SMART,Gene3D,SUPERFAMILY \
-f tsv -o domain_architecture.tsv
3. Functional Classification
# Get GO terms and pathways
interproscan.sh -i sequences.fasta \
-f tsv -goterms -pathways \
-o functional_annotation.tsv
4. High-Throughput Analysis
# Use multiple CPUs for large datasets
interproscan.sh -i large_proteome.fasta \
-f tsv -cpu 8 \
-T /scratch/interproscan_tmp \
-o proteome_results.tsv
π― Output Formats
InterProScan supports multiple output formats:
- TSV (
-f tsv): Tab-separated values (default) - GFF3 (
-f gff3): Genomic feature format - JSON (
-f json): JSON format - XML (
-f xml): XML format - HTML (
-f html): Visual HTML report (compressed) - SVG (
-f svg): Graphical representation
TSV Output Columns
- Protein accession
- Sequence MD5 digest
- Sequence length
- Analysis method
- Signature accession
- Signature description
- Start location
- Stop location
- E-value
- Match status
- Date
- InterPro accession (if available)
- InterPro description (if available)
- GO annotations (if -goterms used)
- Pathways (if -pathways used)
βοΈ Important Options
Performance Tuning
# Number of parallel threads (default: 8)
-cpu 16
# Temporary directory for intermediate files
-T /scratch/tmp_interproscan
# Disable pre-calculation (faster for few sequences)
-dp
Application Selection
# Run specific applications only
-appl Pfam,SMART,Gene3D
# Exclude specific applications
-exclappl Coils,MobiDBLite
# List all available applications
interproscan.sh -appl
Sequence Input
# FASTA file
-i sequences.fasta
# Skip sequences already in output (resume interrupted run)
-resume
# Input type (default: auto-detect)
-t p # protein
-t n # nucleic acid
π Slurm Job Examples
Single Node Job
#!/bin/bash
#SBATCH --job-name=interproscan
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=32G
#SBATCH --time=24:00:00
#SBATCH --output=interproscan_%j.log
module load interproscan
# Set temp directory to job-specific location
export INTERPROSCAN_TMP=/scratch/$SLURM_JOB_ID
interproscan.sh \
-i my_proteome.fasta \
-f tsv,gff3 \
-goterms -pathways \
-cpu $SLURM_CPUS_PER_TASK \
-T $INTERPROSCAN_TMP \
-o results_$SLURM_JOB_ID.tsv
Array Job for Multiple Files
#!/bin/bash
#SBATCH --job-name=interproscan_array
#SBATCH --array=1-100
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=16G
#SBATCH --time=12:00:00
module load interproscan
# Get input file for this array task
INPUT=$(ls input_files/*.fasta | sed -n ${SLURM_ARRAY_TASK_ID}p)
BASENAME=$(basename $INPUT .fasta)
interproscan.sh \
-i $INPUT \
-f tsv \
-cpu $SLURM_CPUS_PER_TASK \
-o results/${BASENAME}_interpro.tsv
π Troubleshooting
Memory Issues
InterProScan can be memory-intensive. Recommended memory per CPU:
- Small proteins (<500 sequences): 2 GB per CPU
- Medium datasets (500-5000): 4 GB per CPU
- Large proteomes (>5000): 8+ GB per CPU
# Reduce parallelism if running out of memory
interproscan.sh -i large.fasta -cpu 4 -o output.tsv
Disk Space Issues
# Set temp directory to location with more space
export INTERPROSCAN_TMP=/scratch/my_interproscan_tmp
mkdir -p $INTERPROSCAN_TMP
interproscan.sh -i proteins.fasta -T $INTERPROSCAN_TMP -o results.tsv
Resume Interrupted Runs
# InterProScan can resume from interrupted runs
interproscan.sh -i sequences.fasta -resume -o results.tsv
Java Issues
# Check Java version (requires Java 11+)
java -version
# If Java issues occur, try setting JAVA_HOME
export JAVA_HOME=/usr/lib/jvm/java-11-openjdk-amd64
π Performance Tips
-
Choose specific applications instead of running all:
bash # Faster: only essential databases interproscan.sh -i input.fasta -appl Pfam,SMART,Gene3D -o output.tsv -
Use appropriate CPU count:
- Too many CPUs β memory issues
- Too few CPUs β slow runtime
-
Sweet spot: 8-16 CPUs for most datasets
-
Use fast local storage for temp files:
bash -T /local/scratch # faster than network storage -
Split large datasets into smaller chunks:
bash # Process in batches of 1000 sequences split -l 2000 large.fasta chunk_ # 1000 seqs (2 lines each in FASTA)
π Integration with Other Tools
Combine with BLAST Results
# Extract top BLAST hits, then annotate with InterProScan
module load blast interproscan
# 1. BLAST search
blastp -query query.fasta -db nr -out blast.tsv -outfmt 6
# 2. Extract unique subject sequences
cut -f2 blast.tsv | sort -u > subjects.txt
blastdbcmd -db nr -entry_batch subjects.txt -out subjects.fasta
# 3. Annotate with InterProScan
interproscan.sh -i subjects.fasta -f tsv -o annotations.tsv
Parse Results with Python
import pandas as pd
# Read TSV output
df = pd.read_csv('interproscan.tsv', sep='\t', header=None)
df.columns = ['acc', 'md5', 'len', 'db', 'sig_acc', 'sig_desc',
'start', 'stop', 'evalue', 'status', 'date',
'interpro_acc', 'interpro_desc', 'go', 'pathways']
# Filter for Pfam domains
pfam = df[df['db'] == 'Pfam']
# Count domain occurrences
domain_counts = pfam['sig_acc'].value_counts()
print(domain_counts.head())
π References & Resources
- Official Documentation: https://interproscan-docs.readthedocs.io/
- InterPro Website: https://www.ebi.ac.uk/interpro/
- GitHub: https://github.com/ebi-pf-team/interproscan
- Publication: Jones P, et al. (2014) Bioinformatics 30(9):1236-1240
π Support
- Cluster Support: uliano.guerrini@unimi.it
- InterProScan Issues: https://github.com/ebi-pf-team/interproscan/issues
- Module Location:
/opt/modulefiles/interproscan/5.75-106.0.lua
Installation Details:
- Path: /sw/interproscan/5.75-106.0/
- Module: module load interproscan
- Type: Precompiled 64-bit binaries
- Dependencies: Java 11+ (system)
- Database Version: 106.0 (June 2024)