RoseTTAFold2 User Guide
Overview
RoseTTAFold2 (RF2) is an advanced protein structure prediction tool that uses deep learning to predict the 3D structures of proteins and protein complexes from amino acid sequences. It's particularly powerful for:
- Monomer prediction: Predicting single protein structures
- Complex prediction: Predicting multi-chain protein complexes with paired MSAs
- Symmetric assemblies: Predicting symmetric oligomers (homo/hetero-oligomers)
- Cyclic peptides: Predicting structures of cyclic peptides
Installation on XLence cluster:
- Location: /sw/rosettafold/2.0/
- Conda Environment: /sw/rosettafold/2.0/env/
- Scripts: /sw/rosettafold/2.0/RoseTTAFold2/
- Pre-trained weights: /sw/rosettafold/2.0/weights/
- Version: January 2024 release
GPU Requirements
- Minimum GPU Memory: 11 GB (available on all XLence compute nodes)
- Cluster GPUs: NVIDIA RTX 2080 Ti (11GB VRAM)
- CUDA Version: 12.1 (via conda environment)
Quick Start
1. Load the RoseTTAFold2 module
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically
When the module loads, it will automatically set up RoseTTAFold2 paths:
- ROSETTAFOLD2_HOME: /sw/rosettafold/2.0/RoseTTAFold2
- ROSETTAFOLD2_WEIGHTS: /sw/rosettafold/2.0/weights
- ROSETTAFOLD2_DB: /sw/rosettafold/2.0/databases (optional databases)
2. Prepare your input FASTA file
For a monomer, use standard FASTA format:
>protein_name
MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPIL...
For a heterodimer, separate chains with ::
>proteinA:proteinB
MKTAYIAKQRQIS:VLSEGEWQLVLHVWAKVEAD
For a heterotrimer, use multiple ::
>proteinA:proteinB:proteinC
SEQUENCE_A:SEQUENCE_B:SEQUENCE_C
3. Run prediction
Basic monomer prediction:
cd $ROSETTAFOLD2_HOME
./run_RF2.sh my_protein.fasta -o output_dir
Usage Examples
Example 1: Predict a monomer structure
cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7UGF.fasta -o 7UGF
Example 2: Predict a heterodimer with paired MSA
cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_8HBN.fasta --pair -o 8HBN
The --pair flag indicates that the MSA for the complex should be paired (joint MSA search).
Example 3: Predict a heterotrimer with paired MSA
cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7ZLR.fasta --pair -o 7ZLR
Example 4: Predict a C6-symmetric homodimer
cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7YTB.fasta --symm C6 -o 7YTB
Supported symmetry groups: C2, C3, C4, C5, C6, D2, D3, D4, D5, D6, etc.
Example 5: Predict a C3-symmetric heterodimer (A₃B₃)
cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7LAW.fasta --symm C3 --pair -o 7LAW
Example 6: Predict a cyclic peptide
cd $ROSETTAFOLD2_HOME
./run_RF2.sh my_cyclic_peptide.fasta -cyclize -o cyclic_out
Command-Line Options
Usage: run_RF2.sh [-o|--outdir name] [-s|--symm symmgroup] [-p|--pair]
[-h|--hhpred] [-cyclize] input.fasta
Options:
-o, --outdir DIR Output directory (default: rf2out)
-s, --symm GROUP Symmetry group (C2-C6, D2-D6, etc.)
-p, --pair Use paired MSA for complexes
-h, --hhpred Run HHsearch for template search
-cyclize Predict cyclic peptide structure
Output Files
Predictions are saved in <outdir>/models/:
- model_final.pdb: Final predicted structure
- B-factors represent predicted LDDT (confidence scores)
-
Higher B-factors = higher confidence
-
model_final.npz: Detailed prediction data (NumPy format)
- Predicted distance distributions
- Predicted aligned error (PAE) matrix
-
Per-residue confidence scores
-
model_final.json: Summary accuracy metrics
- Global confidence scores
- Per-residue pLDDT scores
Running on Slurm
Single GPU Job
#!/bin/bash
#SBATCH --job-name=rosettafold2
#SBATCH --partition=gpu
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=4:00:00
#SBATCH --output=rf2_%j.out
#SBATCH --error=rf2_%j.err
# Load environment
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically
# Run prediction
cd $ROSETTAFOLD2_HOME
./run_RF2.sh /path/to/input.fasta -o /path/to/output
# Deactivate when done
# Module unloaded automatically
Array job for multiple proteins
#!/bin/bash
#SBATCH --job-name=rf2_array
#SBATCH --partition=gpu
#SBATCH --array=1-10
#SBATCH --nodes=1
#SBATCH --cpus-per-task=8
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=4:00:00
#SBATCH --output=rf2_%A_%a.out
#SBATCH --error=rf2_%A_%a.err
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically
# Get FASTA file for this array task
FASTA_FILE=$(sed -n "${SLURM_ARRAY_TASK_ID}p" fasta_list.txt)
BASENAME=$(basename $FASTA_FILE .fasta)
cd $ROSETTAFOLD2_HOME
./run_RF2.sh $FASTA_FILE -o ${BASENAME}_out
# Module unloaded automatically
Sequence Databases (Optional)
RoseTTAFold2 can use sequence databases for MSA generation to improve predictions. These are optional and not currently installed:
- UniRef30 (~46 GB): Clustered protein sequence database
- BFD (~272 GB): Big Fantastic Database for deep MSAs
- pdb100 (~10 GB): Structure template database
If you need these databases, contact cluster administrators.
Understanding Output Quality
pLDDT Scores (B-factors in PDB)
- >90: Very high confidence (dark blue)
- 80-90: High confidence (light blue)
- 70-80: Low confidence (yellow)
- <70: Very low confidence (orange/red)
PAE (Predicted Aligned Error)
- Low PAE (<5Å): High confidence in relative domain positions
- High PAE (>10Å): Uncertain relative positioning
Troubleshooting
Out of memory errors
- Reduce sequence length by removing flexible linkers
- Request more memory in Slurm job (increase
--mem) - Split large complexes into domains
Poor quality predictions
- Check if input sequence is correct
- For complexes, try with
--pairflag - For symmetric assemblies, specify correct symmetry group
- Consider downloading sequence databases for better MSAs
Environment issues
# Verify environment is activated
echo $ROSETTAFOLD2_HOME # Should show /sw/rosettafold/2.0/RoseTTAFold2
# Check Python version
python --version # Should be Python 3.10.18
# Verify PyTorch installation
python -c "import torch; print(torch.__version__)" # Should be 2.2.0
# Check GPU availability
python -c "import torch; print(torch.cuda.is_available())" # Should be True on compute nodes
Performance Tips
- CPU cores: Use 8 CPUs for MSA generation (
--cpus-per-task=8) - Memory: 32GB is usually sufficient; increase for very large complexes
- Runtime: Typical predictions take 1-4 hours depending on sequence length
- Batch jobs: Use Slurm arrays for multiple predictions
Advanced Features
Custom MSA Input
If you already have an MSA, you can provide it directly:
# Place your MSA file (.a3m format) in the output directory
# RoseTTAFold2 will detect and use it
Template-based modeling
Use the --hhpred flag to search for and use structure templates:
./run_RF2.sh input.fasta --hhpred -o output
Note: Requires pdb100 database (not currently installed).
Citing RoseTTAFold2
If you use RoseTTAFold2 for your research, please cite:
Baek, M., DiMaio, F., Anishchenko, I. et al.
Accurate prediction of protein structures and interactions using a three-track neural network.
Science (2024)
Additional Resources
- GitHub Repository: https://github.com/uw-ipd/RoseTTAFold2
- Original Paper: Check repository for latest publications
- Example Files:
$ROSETTAFOLD2_HOME/examples/ - Cluster Support: Contact administrators for help
Quick Reference
# Activate environment
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically
# Basic prediction
cd $ROSETTAFOLD2_HOME
./run_RF2.sh input.fasta -o output_dir
# Complex with paired MSA
./run_RF2.sh complex.fasta --pair -o output_dir
# Symmetric assembly
./run_RF2.sh protein.fasta --symm C3 -o output_dir
# Cyclic peptide
./run_RF2.sh peptide.fasta -cyclize -o output_dir
# Deactivate environment
# Module unloaded automatically
Environment Variables
When the module is loaded, these variables are automatically set:
ROSETTAFOLD2_HOME:/sw/rosettafold/2.0/RoseTTAFold2ROSETTAFOLD2_WEIGHTS:/sw/rosettafold/2.0/weightsROSETTAFOLD2_DB:/sw/rosettafold/2.0/databases
These are automatically removed when you run module unload rosettafold.
Module Conflicts
The RoseTTAFold2 module is mutually exclusive with other Python environments: - miniforge3 - General-purpose conda/mamba environment - biostar - Bioinformatics tools collection
You cannot load rosettafold together with miniforge3 or biostar. Lmod will prevent this and show an error message.
If you need to switch between environments:
# Switch from rosettafold to miniforge3
module unload rosettafold
module load miniforge3
# Switch from miniforge3 to rosettafold
module unload miniforge3
module load rosettafold/2.0
Why this restriction? Each module provides its own Python interpreter and libraries. Loading multiple Python environments simultaneously would create conflicts and unpredictable behavior.