Skip to content

RoseTTAFold2 User Guide

Overview

RoseTTAFold2 (RF2) is an advanced protein structure prediction tool that uses deep learning to predict the 3D structures of proteins and protein complexes from amino acid sequences. It's particularly powerful for:

  • Monomer prediction: Predicting single protein structures
  • Complex prediction: Predicting multi-chain protein complexes with paired MSAs
  • Symmetric assemblies: Predicting symmetric oligomers (homo/hetero-oligomers)
  • Cyclic peptides: Predicting structures of cyclic peptides

Installation on XLence cluster: - Location: /sw/rosettafold/2.0/ - Conda Environment: /sw/rosettafold/2.0/env/ - Scripts: /sw/rosettafold/2.0/RoseTTAFold2/ - Pre-trained weights: /sw/rosettafold/2.0/weights/ - Version: January 2024 release

GPU Requirements

  • Minimum GPU Memory: 11 GB (available on all XLence compute nodes)
  • Cluster GPUs: NVIDIA RTX 2080 Ti (11GB VRAM)
  • CUDA Version: 12.1 (via conda environment)

Quick Start

1. Load the RoseTTAFold2 module

module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically

When the module loads, it will automatically set up RoseTTAFold2 paths: - ROSETTAFOLD2_HOME: /sw/rosettafold/2.0/RoseTTAFold2 - ROSETTAFOLD2_WEIGHTS: /sw/rosettafold/2.0/weights - ROSETTAFOLD2_DB: /sw/rosettafold/2.0/databases (optional databases)

2. Prepare your input FASTA file

For a monomer, use standard FASTA format:

>protein_name
MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQAPIL...

For a heterodimer, separate chains with ::

>proteinA:proteinB
MKTAYIAKQRQIS:VLSEGEWQLVLHVWAKVEAD

For a heterotrimer, use multiple ::

>proteinA:proteinB:proteinC
SEQUENCE_A:SEQUENCE_B:SEQUENCE_C

3. Run prediction

Basic monomer prediction:

cd $ROSETTAFOLD2_HOME
./run_RF2.sh my_protein.fasta -o output_dir

Usage Examples

Example 1: Predict a monomer structure

cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7UGF.fasta -o 7UGF

Example 2: Predict a heterodimer with paired MSA

cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_8HBN.fasta --pair -o 8HBN

The --pair flag indicates that the MSA for the complex should be paired (joint MSA search).

Example 3: Predict a heterotrimer with paired MSA

cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7ZLR.fasta --pair -o 7ZLR

Example 4: Predict a C6-symmetric homodimer

cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7YTB.fasta --symm C6 -o 7YTB

Supported symmetry groups: C2, C3, C4, C5, C6, D2, D3, D4, D5, D6, etc.

Example 5: Predict a C3-symmetric heterodimer (A₃B₃)

cd $ROSETTAFOLD2_HOME
./run_RF2.sh examples/rcsb_pdb_7LAW.fasta --symm C3 --pair -o 7LAW

Example 6: Predict a cyclic peptide

cd $ROSETTAFOLD2_HOME
./run_RF2.sh my_cyclic_peptide.fasta -cyclize -o cyclic_out

Command-Line Options

Usage: run_RF2.sh [-o|--outdir name] [-s|--symm symmgroup] [-p|--pair]
                   [-h|--hhpred] [-cyclize] input.fasta

Options:
  -o, --outdir DIR     Output directory (default: rf2out)
  -s, --symm GROUP     Symmetry group (C2-C6, D2-D6, etc.)
  -p, --pair           Use paired MSA for complexes
  -h, --hhpred         Run HHsearch for template search
  -cyclize             Predict cyclic peptide structure

Output Files

Predictions are saved in <outdir>/models/:

  • model_final.pdb: Final predicted structure
  • B-factors represent predicted LDDT (confidence scores)
  • Higher B-factors = higher confidence

  • model_final.npz: Detailed prediction data (NumPy format)

  • Predicted distance distributions
  • Predicted aligned error (PAE) matrix
  • Per-residue confidence scores

  • model_final.json: Summary accuracy metrics

  • Global confidence scores
  • Per-residue pLDDT scores

Running on Slurm

Single GPU Job

#!/bin/bash
#SBATCH --job-name=rosettafold2
#SBATCH --partition=gpu
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=4:00:00
#SBATCH --output=rf2_%j.out
#SBATCH --error=rf2_%j.err

# Load environment
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically

# Run prediction
cd $ROSETTAFOLD2_HOME
./run_RF2.sh /path/to/input.fasta -o /path/to/output

# Deactivate when done
# Module unloaded automatically

Array job for multiple proteins

#!/bin/bash
#SBATCH --job-name=rf2_array
#SBATCH --partition=gpu
#SBATCH --array=1-10
#SBATCH --nodes=1
#SBATCH --cpus-per-task=8
#SBATCH --gres=gpu:1
#SBATCH --mem=32G
#SBATCH --time=4:00:00
#SBATCH --output=rf2_%A_%a.out
#SBATCH --error=rf2_%A_%a.err

module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically

# Get FASTA file for this array task
FASTA_FILE=$(sed -n "${SLURM_ARRAY_TASK_ID}p" fasta_list.txt)
BASENAME=$(basename $FASTA_FILE .fasta)

cd $ROSETTAFOLD2_HOME
./run_RF2.sh $FASTA_FILE -o ${BASENAME}_out

# Module unloaded automatically

Sequence Databases (Optional)

RoseTTAFold2 can use sequence databases for MSA generation to improve predictions. These are optional and not currently installed:

  • UniRef30 (~46 GB): Clustered protein sequence database
  • BFD (~272 GB): Big Fantastic Database for deep MSAs
  • pdb100 (~10 GB): Structure template database

If you need these databases, contact cluster administrators.

Understanding Output Quality

pLDDT Scores (B-factors in PDB)

  • >90: Very high confidence (dark blue)
  • 80-90: High confidence (light blue)
  • 70-80: Low confidence (yellow)
  • <70: Very low confidence (orange/red)

PAE (Predicted Aligned Error)

  • Low PAE (<5Å): High confidence in relative domain positions
  • High PAE (>10Å): Uncertain relative positioning

Troubleshooting

Out of memory errors

  • Reduce sequence length by removing flexible linkers
  • Request more memory in Slurm job (increase --mem)
  • Split large complexes into domains

Poor quality predictions

  • Check if input sequence is correct
  • For complexes, try with --pair flag
  • For symmetric assemblies, specify correct symmetry group
  • Consider downloading sequence databases for better MSAs

Environment issues

# Verify environment is activated
echo $ROSETTAFOLD2_HOME  # Should show /sw/rosettafold/2.0/RoseTTAFold2

# Check Python version
python --version  # Should be Python 3.10.18

# Verify PyTorch installation
python -c "import torch; print(torch.__version__)"  # Should be 2.2.0

# Check GPU availability
python -c "import torch; print(torch.cuda.is_available())"  # Should be True on compute nodes

Performance Tips

  1. CPU cores: Use 8 CPUs for MSA generation (--cpus-per-task=8)
  2. Memory: 32GB is usually sufficient; increase for very large complexes
  3. Runtime: Typical predictions take 1-4 hours depending on sequence length
  4. Batch jobs: Use Slurm arrays for multiple predictions

Advanced Features

Custom MSA Input

If you already have an MSA, you can provide it directly:

# Place your MSA file (.a3m format) in the output directory
# RoseTTAFold2 will detect and use it

Template-based modeling

Use the --hhpred flag to search for and use structure templates:

./run_RF2.sh input.fasta --hhpred -o output

Note: Requires pdb100 database (not currently installed).

Citing RoseTTAFold2

If you use RoseTTAFold2 for your research, please cite:

Baek, M., DiMaio, F., Anishchenko, I. et al.
Accurate prediction of protein structures and interactions using a three-track neural network.
Science (2024)

Additional Resources

  • GitHub Repository: https://github.com/uw-ipd/RoseTTAFold2
  • Original Paper: Check repository for latest publications
  • Example Files: $ROSETTAFOLD2_HOME/examples/
  • Cluster Support: Contact administrators for help

Quick Reference

# Activate environment
module load rosettafold/2.0
# RoseTTAFold2 environment loaded automatically

# Basic prediction
cd $ROSETTAFOLD2_HOME
./run_RF2.sh input.fasta -o output_dir

# Complex with paired MSA
./run_RF2.sh complex.fasta --pair -o output_dir

# Symmetric assembly
./run_RF2.sh protein.fasta --symm C3 -o output_dir

# Cyclic peptide
./run_RF2.sh peptide.fasta -cyclize -o output_dir

# Deactivate environment
# Module unloaded automatically

Environment Variables

When the module is loaded, these variables are automatically set:

  • ROSETTAFOLD2_HOME: /sw/rosettafold/2.0/RoseTTAFold2
  • ROSETTAFOLD2_WEIGHTS: /sw/rosettafold/2.0/weights
  • ROSETTAFOLD2_DB: /sw/rosettafold/2.0/databases

These are automatically removed when you run module unload rosettafold.

Module Conflicts

The RoseTTAFold2 module is mutually exclusive with other Python environments: - miniforge3 - General-purpose conda/mamba environment - biostar - Bioinformatics tools collection

You cannot load rosettafold together with miniforge3 or biostar. Lmod will prevent this and show an error message.

If you need to switch between environments:

# Switch from rosettafold to miniforge3
module unload rosettafold
module load miniforge3

# Switch from miniforge3 to rosettafold
module unload miniforge3
module load rosettafold/2.0

Why this restriction? Each module provides its own Python interpreter and libraries. Loading multiple Python environments simultaneously would create conflicts and unpredictable behavior.