Microbial GAIN domains undergo autoproteolysis and enable release of diverse cell surface associated proteins

Published: 8 July 2026| Version 1 | DOI: 10.17632/sm3rzy6sbx.1
Contributors:
Anna Brogan, David Rudner

Description

Domain annotations of MAIN domain containing proteins using PfamScan and DPAM. Dataset S1 - Counts of Pfam domains associated with MAIN domains and Pfam annotation of MAIN-domain containing proteins Dataset S2 - Counts of ECOD domains associated with MAIN domains and DPAM annotation of MAIN-domain containing proteins

Files

Steps to reproduce

A local PSI-BLAST run was performed on the amino acid sequence of B. subtilis YfkN’s MAIN domain against the RefSeq Select Protein database using an e-value of 0.05 and 5 iterations. Eukaryotic proteins were removed from the results and the remaining protein sequences retrieved. PfamScan The ~15,000 proteins identified by PSI-BLAST were annotated using PfamScan against the Pfam database release 31. The output domains were organized by RefSeq accession number and the resulting domain organizations built (Dataset S1). Counts of each domain found associated with the MAIN domain can be found in Dataset S1. DPAM MAIN domain-containing proteins identified by PSI-BLAST were clustered using MMseqs2 using a minimum sequence identity of 50% and a coverage threshold of 80%. The structures of the resulting 2,843 unique proteins were predicted using colabfold version 1.5.5 with the following parameters: each prediction was run with five models, using three recycles, no dropout, no templates. These runs were performed using a local installation of Colabfold running on NVIDIA A6000 GPUs managed by the HMS O2 computing cluster. MSAs were generated using MMseqs2 API server. The rank 1 prediction (based on pLDDT) from each AlphaFold2 run was used as an input into DPAM. The resulting output domains were organized by RefSeq accession number and the domain organizations built (Dataset S2).

Categories

Microbiology, Bioinformatics

Licence