Skip to content

Latest commit

 

History

648 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HInt_logo

HInt

HInt accelerates AlphaFold by optimizing computations and parallelizing structure predictions. It is a scalable pipeline for high-throughput identification of homologous proteins and interologues—proteins that maintain functional interactions. HInt enables the discovery of conserved interaction networks that may remain undetected using sequence or structural similarity alone.

1. Installation

1.1. HInt

Conda or Mamba must be installed on your system and should not be activated before running the installation.

wget https://raw.githubusercontent.com/Qrouger/HInt/main/Install_HInt.sh
bash Install_HInt.sh
AlphaFold 3 (optional)

⚠️ Warning
AlphaFold 3 model parameters must be downloaded and provided through Path_AlphaFold_Data. (https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md)

Access to AlphaFold 3 parameters is subject to the DeepMind terms of use.

1.2. DeepLoc2 (Eukaryote)

Download the DeepLoc 2.0 package from: https://services.healthtech.dtu.dk/services/DeepLoc-2.0/

conda activate HInt
cd  deeploc2_package
pip install . torch==2.6.0
pip install triton==3.1.0 nvidia-cudnn-cu12==9.25.0.15

1.3. SignalP5

Download the SignalP5 package from: https://services.healthtech.dtu.dk/services/SignalP-5.0/9-Downloads.php

tar -xvzf signalp-5.0b.Linux.tar.gz
cd signalp-5.0b/
sudo cp bin/signalp /usr/local/bin
sudo cp -r lib/* /usr/local/lib

1.4. CCP4

Download the ccp4 package from: https://www.ccp4.ac.uk/download/#os=linux

tar xvzf ccp4-9-setup.tar.gz
./ccp4-9-setup

2. Download databases

2.1. Download the GPU-indexed MMseqs2 database (~1.9 TB)

For optimal performance, store MMseqs2 databases and MSA on NVMe or SSD storage rather than HDDs.

wget https://raw.githubusercontent.com/sokrypton/ColabFold/main/setup_databases.sh
chmod +x setup_databases.sh 
GPU=1 ./setup_databases.sh ./MMseqs2_GPU_database

2.2. Download AlphaFold 2 database (~2.2 TB)

git clone https://github.com/deepmind/alphafold.git
cd ./alphafold
scripts/download_all_data.sh <DB_DIR> > download.log 2> download_all.log &

2.3. Download AlphaFold 3 database (633G)

git clone https://github.com/google-deepmind/alphafold3.git
cd alphafold3
./fetch_databases.sh <DB_DIR>

3. Input parameters

3.1. Setup HInt.txt

You need to download or copy HInt.txt file example.

A priori information

  • Signal_peptide : Filter proteins based on the presence of a predicted signal peptide (Options : Yes, No or None).

  • DeepLoc : Cellular localisation(s) of the protein. Multiple localizations can be specified, separated by commas. All proteins predicted to be in one of these compartments will be used.

    • Eukaryotes : Cytoplasm, Nucleus, Extracellular, Cell membrane, Mitochondrion, Plastid, Endoplasmic reticulum, Lysosome/Vacuole, Golgi apparatus, Peroxisome.
    • Prokaryotes : Cell wall & surface, Extracellular, Cytoplasmic, Cytoplasmic Membrane, Outer Membrane, Periplasmic.
  • Max_protein_length : Maximum length of the protein you search (integer).

  • Min_protein_length : Minimum length of the protein you search (integer), default set on 20aa.

  • AlphaFold : AlphaFold version (Options : 2 or 3).

  • Homo-oligomer : Known homo-oligomerization state of the protein (integer : 1 to 20), default set on 1 (monomer).

  • Interact_with : Names of proteins expected to interact with the query protein (UniprotID or protein fasta name).

Advanced Interact_with uses and examples

One bait :

Interact_with : UniprotID1

Region of a bait :

Interact_with : UniprotID1(20-200)

Multiple baits : # The first protein must correspond to the primary bait. For now, you can specify a maximum of three different baits

Interact_with : UniprotID1, UniprotID2 

Multimer bait : # Create a unique bait with multiple proteins

Interact_with : [Uniprot1, Uniprot2]

And you can mixed up all of theses examples !

⚠️ Warning
HInt currently does not support multiple regions for bait proteins.

  • Organism : Organism of interest for SignalP5 and DeepLoc (arch, gram+, gram-, euk or None). Enables signal peptide prediction and cleavage.

Tip

If you don’t know an information or want to skip it, you can leave this field blank.


Paths

  • Path_AlphaFold_Data : Path of AlphaFold database (string).

  • Path_ccp4 : Path of CCP4 package (string). Default set on /opt/xtal/ccp4-9.

  • Path_MMseqs2_Data : Path of GPU-indexed MMseqs2 database (string).

Note

This path is not mandatory. If not set also MMseqs2-GPU will not be used.

  • Path_Uniprot_ID : Path to the protein sequence file (string).

  • Path_Pickle_Feature : Path where MSA files will be saved (string).


3.2. Setup protein file

The protein file must contain all UniProt IDs or all sequences in FASTA format for both preys and baits.
This can be an NCBI protein FASTA file, a standard FASTA file, UniProt identifiers, or a combination of these formats.

Tip

The use of UniprotIDs is recommended for pipeline speed.

Protein file examples

Example 1

A0ABD7FQG0,P18004,P15069,P33790

Example 2

>A0ABD7FQG0
MSGDENKLKKYRFPETLTNQSRWFGLPLDELIPAAICIGWGITTSKYLFGIGAAVLVYFGIKKLKKGRGSSWLRDLIYWYMPTALLRGIFHNVPDSCFRQWIK
>P18004
MNNPLEAVTQAVNSLVTALKLPDESAKANEVLGEMSFPQFSRLLPYRDYNQESGLFMNDTTMGFMLEAIPINGANESIVEALDHMLRTKLPRGIPLCIHLMSSQLVGDRIEYGLREFSWSGEQAERFNAITRAYYMKAAATQFPLPEGMNLPLTLRHYRVFISYCSPSKKKSRADILEMENLVKIIRASLQGASITTQTVDAQAFIDIVGEMINHNPDSLYPKRRQLDPYSDLNYQCVEDSFDLKVRADYLTLGLRENGRNSTARILNFHLARNPEIAFLWNMADNYSNLLNPELSISCPFILTLTLVVEDQVKTHSEANLKYMDLEKKSKTSYAKWFPSVEKEAKEWGELRQRLGSGQSSVVSYFLNITAFCKDNNETALEVEQDILNSFRKNGFELISPRFNHMRNFLTCLPFMAGKGLFKQLKEAGVVQRAESFNVANLMPLVADNPLTPAGLLAPTYRNQLAFIDIFFRGMNNTNYNMAVCGTSGAGKTGLIQPLIRSVLDSGGFAVVFDMGDGYKSLCENMGGVYLDGETLRFNPFANITDIDQSAERVRDQLSVMASPNGNLDEVHEGLLLQAVRASWLAKENRARIDDVVDFLKNASDSEQYAESPTIRSRLDEMIVLLDQYTANGTYGQYFNSDEPSLRDDAKMVVLELGGLEDRPSLLVAVMFSLIIYIENRMYRTPRNLKKLNVIDEGWRLLDFKNHKVGEFIEKGYRTARRHTGAYITITQNIVDFDSDKASSAARAAWGNSSYKIILKQSAKEFAKYNQLYPDQFLPLQRDMIGKFGAAKDQWFSSFLLQVENHSSWHRLFVDPLSRAMYSSDGPDFEFVQQKRKEGLSIHEAVWQLAWKKSGPEMASLEAWLEEHEKYRSVA
>P15069
MMPRIKPLLVLCAALLTVTPAASADVNSDMNQFFNKLGFASNTTQPGVWQGQAAGYAYGGSLYARTQVKNVQLISMTLPDINAGCGGIDAYLGSFSFINGEQLQRFVKQIMSNAAGYFFDLALQTTVPEIKTAKDFLQKMASDINSMNLSSCQAAQGIIGGLFPRTQVSQQKVCQDIAGESNIFADWAASRQGCTVGGKSDSVRDKASDKDKERVTKNINIMWNALSKNRMFDGNKELKEFVMTLTGSLVFGPNGEITPLSARTTDRSIIRAMMEGGTAKISHCNDSDKCLKVVADTPVTISRDNALKSQITKLLASIQNKAVSDTPLDDKEKGFISSTTIPVFKYLVDPQMLGVSNSMIYQLTDYIGYDILLQYIQELIQQARAMVATGNYDEAVIGHINDNMNDATRQIAAFQSQVQVQQDALLVVDRQMSYMRQQLSARMLSRYQNNYHFGGSTL
>P33790
MNEVYVIAGGEWLRNNLNAIAAFMGTWTWDSIEKIALTLSVLAVAVMWVQRHNVMDLLGWVAVFVLISLLVNVRTSVQIIDNSDLVKVHRVDNVPVGLAMPLSLTTRIGHAMVASYEMIFTQPDSVTYSKTGMLFGANLIVKSTDFLSRNPEIINLFQDYVQNCVLGDIYLNHKYTLEDLMASADPYTLIFSRPSPLRGVYDNNNNFITCKDASVTLKDRLNLDTKTGGKTWHYYVQQIFGGRPDPDLLFRQLVSDSYSYFYGSSQSASQIMRQNVTMNALKEGITSNAARNGDTASLVSLATTSSMEKQRLAHVSIGHVTMRNLPMVQTILTGIAIGIFPLLILAAVFNKLTLSVLKGYVFALMWLQTWPLLYAILNSAMTFYAKQNGAPVVLSELSQIQLKYSNLASTAGYLSAMIPPLSWMMVKGLGAGFSSVYSHFASSSISPTASAAGSVVDGNYSYGNMQTENVNGFSWSTNSTTSFGQMMYQTGSGATATQTRDGNMVMDASGAMSRLPVGINATRQIAAAQQEMAREASNRAESALHGFSSSIASAWNTLSQFGSNRGSSDSVTGGADSTMSAQDSMMASRMRSAVESYAKAHNISNEQATRELASRSTNASLGLYGDAYAKGHLGISVLGNGGGVGLQAGAKASIDGSDLDSHEASSGSRASHDARHDIDARATQDFKEASDYFTSRKVSESGSHTDNNADSRVDQLSAALNSAKQSYDQYTTNMTRSHEYAEMASRTESMSGQMSEDLSQQFAQYVMKNAPQDVEAILTNTSSPEIAERRRAMAWSFVQEQVQPGVDNTWRESRRDIGKGMESVPSGGGSQDIIADHQGHQAIIEQRTQDSNIRNDVKHQVDNMVTEYRGNIGDTQNSIRGEENIVKGQYSELQNHHKTEALTQNNKYNEEKLAQERIPGADSPKELLEKAKSYQHKE

Example 3

A0ABD7FQG0,P18004
>P15069
MMPRIKPLLVLCAALLTVTPAASADVNSDMNQFFNKLGFASNTTQPGVWQGQAAGYAYGGSLYARTQVKNVQLISMTLPDINAGCGGIDAYLGSFSFINGEQLQRFVKQIMSNAAGYFFDLALQTTVPEIKTAKDFLQKMASDINSMNLSSCQAAQGIIGGLFPRTQVSQQKVCQDIAGESNIFADWAASRQGCTVGGKSDSVRDKASDKDKERVTKNINIMWNALSKNRMFDGNKELKEFVMTLTGSLVFGPNGEITPLSARTTDRSIIRAMMEGGTAKISHCNDSDKCLKVVADTPVTISRDNALKSQITKLLASIQNKAVSDTPLDDKEKGFISSTTIPVFKYLVDPQMLGVSNSMIYQLTDYIGYDILLQYIQELIQQARAMVATGNYDEAVIGHINDNMNDATRQIAAFQSQVQVQQDALLVVDRQMSYMRQQLSARMLSRYQNNYHFGGSTL
>P33790
MNEVYVIAGGEWLRNNLNAIAAFMGTWTWDSIEKIALTLSVLAVAVMWVQRHNVMDLLGWVAVFVLISLLVNVRTSVQIIDNSDLVKVHRVDNVPVGLAMPLSLTTRIGHAMVASYEMIFTQPDSVTYSKTGMLFGANLIVKSTDFLSRNPEIINLFQDYVQNCVLGDIYLNHKYTLEDLMASADPYTLIFSRPSPLRGVYDNNNNFITCKDASVTLKDRLNLDTKTGGKTWHYYVQQIFGGRPDPDLLFRQLVSDSYSYFYGSSQSASQIMRQNVTMNALKEGITSNAARNGDTASLVSLATTSSMEKQRLAHVSIGHVTMRNLPMVQTILTGIAIGIFPLLILAAVFNKLTLSVLKGYVFALMWLQTWPLLYAILNSAMTFYAKQNGAPVVLSELSQIQLKYSNLASTAGYLSAMIPPLSWMMVKGLGAGFSSVYSHFASSSISPTASAAGSVVDGNYSYGNMQTENVNGFSWSTNSTTSFGQMMYQTGSGATATQTRDGNMVMDASGAMSRLPVGINATRQIAAAQQEMAREASNRAESALHGFSSSIASAWNTLSQFGSNRGSSDSVTGGADSTMSAQDSMMASRMRSAVESYAKAHNISNEQATRELASRSTNASLGLYGDAYAKGHLGISVLGNGGGVGLQAGAKASIDGSDLDSHEASSGSRASHDARHDIDARATQDFKEASDYFTSRKVSESGSHTDNNADSRVDQLSAALNSAKQSYDQYTTNMTRSHEYAEMASRTESMSGQMSEDLSQQFAQYVMKNAPQDVEAILTNTSSPEIAERRRAMAWSFVQEQVQPGVDNTWRESRRDIGKGMESVPSGGGSQDIIADHQGHQAIIEQRTQDSNIRNDVKHQVDNMVTEYRGNIGDTQNSIRGEENIVKGQYSELQNHHKTEALTQNNKYNEEKLAQERIPGADSPKELLEKAKSYQHKE


4. Run HInt

You need to be in the directory with HInt.txt file.

HInt --cpu <Integer> --gpu <Integer(s)> --multi_job_per_gpu <Boolean>
Flags description
# Number of CPUs available for computation. Enables CPU parallelization. By default, set to half of the available CPUs.
  --cpu : Integer

# Index(es) of GPU(s) you want to use. Declare multiple GPU allows GPU parallelisation. By default set on GPU 0. 
  --gpu : Integer(s)
  
# Allows multiple jobs to run on a single GPU, reducing time of modeling. By default set on True. 
  --multi_job_per_gpu : Boolean
Initial folder structure
HInt_screen/
  HInt.txt
  sequences.txt

5. Results

Final folder structure
HInt_screen/
  HInt.txt
  sequences.txt
  All_Final_result_HInt.csv
  result_PPI_int/
    P33790_and_P15069
    ...
  msa_feature/
    P33790.pkl
    P33790.a3m
    P33790_coverage.pdf
    P33790_feature_metadata_2026-04-03.json
    P15069.pkl
    P15069.a3m
    P15069_coverage.pdf
    P15069_feature_metadata_2026-04-03.json
    ...
  Interface_fig/
  log_file/
    HInt.log
    Summary_result_HInt.csv
    HInt_report.txt

Global results description

The pipeline generates three main output files:

All_Final_result_HInt.csv

Final curated results after filtering steps.

  • Name: Protein identifier
  • Localization: Predicted subcellular localization
  • Signal_peptide: Signal peptide prediction (presence or not)
  • Score: Final HInt interaction score

Summary_result_HInt.csv

Summary of filtering decisions applied to each entry.

  • Name: Protein identifier
  • Reason_for_filtering: Reason why the entry was filtered or flagged

HInt_report.txt

Comprehensive report of all processed entries and timings.

Includes:

  • All entry information
  • Execution timings for each stage (prediction, scoring, filtering, etc.)

PPI specific results

<PPI>_rest_int.csv
Table of interface residues identified at the protein-protein interface.
Includes interface residues identified using PAE and inter-chain distance criteria (<10 Å).

<PPI>_ranked_0.pdb
Structural model of the predicted complex.
Interface residues can be visualized by coloring the structure using the B-factor field.

Standalone iQ-score Calculation

Compute iQ-score independently of the full HInt workflow..

GitHub

Citations

HInt: interaction-based homology discovery through accelerated genome-scale AlphaFold screening Quentin Rouger, Pierre Paillard, Manon Thomet, Emma Touquet, Gwenaël Rabut, Emmanuel Giudice, Damien F. Meyer, Kévin Macé bioRxiv 2026.07.31.741991; doi: https://doi.org/10.64898/2026.07.31.741991

About

HInt (Homologous Interactions)

Resources

Stars

10 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages