Reference Citation Analysis: Find an Article, Find a Category, Find a Journal, Find a Scholar

For: Yunes JM, Babbitt PC. Effusion: prediction of protein function from sequence similarity networks. Bioinformatics 2019;35:442-451. [PMID: 30084920 PMCID: PMC6361244 DOI: 10.1093/bioinformatics/bty672] [Citation(s) in RCA: 9] [Impact Index Per Article: 1.5] [Reference Citation Analysis] [What about the content of this article? (0)] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/25/2018] [Revised: 07/24/2018] [Accepted: 07/30/2018] [Indexed: 12/26/2022] Open

For:	Yunes JM, Babbitt PC. Effusion: prediction of protein function from sequence similarity networks. Bioinformatics 2019;35:442-451. [PMID: 30084920 PMCID: PMC6361244 DOI: 10.1093/bioinformatics/bty672] [Citation(s) in RCA: 9] [Impact Index Per Article: 1.5] [Reference Citation Analysis] [What about the content of this article? (0)] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/25/2018] [Revised: 07/24/2018] [Accepted: 07/30/2018] [Indexed: 12/26/2022] Open

Number

Cited by Other Article(s)

Monjot A, Rousseau J, Bittner L, Lepère C. Metatranscriptomes-based sequence similarity networks uncover genetic signatures within parasitic freshwater microbial eukaryotes. MICROBIOME 2025;13:43. [PMID: 39915863 PMCID: PMC11800578 DOI: 10.1186/s40168-024-02027-0] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Download PDF] [Figures] [Subscribe] [Scholar Register] [Received: 06/13/2024] [Accepted: 12/31/2024] [Indexed: 02/09/2025]

Abstract

BACKGROUND

Microbial eukaryotes play a crucial role in biochemical cycles and aquatic trophic food webs. Their taxonomic and functional diversity are increasingly well described due to recent advances in sequencing technologies. However, the vast amount of data produced by -omics approaches require data-driven methodologies to make predictions about these microorganisms' role within ecosystems. Using metatranscriptomics data, we employed a sequence similarity network-based approach to explore the metabolic specificities of microbial eukaryotes with different trophic modes in a freshwater ecosystem (Lake Pavin, France).

RESULTS

A total of 2,165,106 proteins were clustered in connected components enabling analysis of a great number of sequences without any references in public databases. This approach coupled with the use of an in-house trophic modes database improved the number of proteins considered by 42%. Our study confirmed the versatility of mixotrophic metabolisms with a large number of shared protein families among mixotrophic and phototrophic microorganisms as well as mixotrophic and heterotrophic microorganisms. Genetic similarities in proteins of saprotrophs and parasites also suggest that fungi-like organisms from Lake Pavin, such as Chytridiomycota and Oomycetes, exhibit a wide range of lifestyles, influenced by their degree of dependence on a host. This plasticity may occur at a fine taxonomic level (e.g., species level) and likely within a single organism in response to environmental parameters. While we observed a relative functional redundancy of primary metabolisms (e.g., amino acid and carbohydrate metabolism) nearly 130,000 protein families appeared to be trophic mode-specific. We found a particular specificity in obligate parasite-related Specific Protein Clusters, underscoring a high degree of specialization in these organisms.

CONCLUSIONS

Although no universal marker for parasitism was identified, candidate genes can be proposed at a fine taxonomic scale. We notably provide several protein families that could serve as keys to understanding host-parasite interactions representing pathogenicity factors (e.g., involved in hijacking host resources, or associated with immune evasion mechanisms). All these protein families could offer valuable insights for developing antiparasitic treatments in health and economic contexts. Video Abstract.

Collapse

Svedberg D, Winiger RR, Berg A, Sharma H, Tellgren-Roth C, Debrunner-Vossbrinck BA, Vossbrinck CR, Barandun J. Functional annotation of a divergent genome using sequence and structure-based similarity. BMC Genomics 2024;25:6. [PMID: 38166563 PMCID: PMC10759460 DOI: 10.1186/s12864-023-09924-y] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/27/2023] [Accepted: 12/18/2023] [Indexed: 01/04/2024] Open

Abstract

BACKGROUND

Microsporidia are a large taxon of intracellular pathogens characterized by extraordinarily streamlined genomes with unusually high sequence divergence and many species-specific adaptations. These unique factors pose challenges for traditional genome annotation methods based on sequence similarity. As a result, many of the microsporidian genomes sequenced to date contain numerous genes of unknown function. Recent innovations in rapid and accurate structure prediction and comparison, together with the growing amount of data in structural databases, provide new opportunities to assist in the functional annotation of newly sequenced genomes.

RESULTS

In this study, we established a workflow that combines sequence and structure-based functional gene annotation approaches employing a ChimeraX plugin named ANNOTEX (Annotation Extension for ChimeraX), allowing for visual inspection and manual curation. We employed this workflow on a high-quality telomere-to-telomere sequenced tetraploid genome of Vairimorpha necatrix. First, the 3080 predicted protein-coding DNA sequences, of which 89% were confirmed with RNA sequencing data, were used as input. Next, ColabFold was used to create protein structure predictions, followed by a Foldseek search for structural matching to the PDB and AlphaFold databases. The subsequent manual curation, using sequence and structure-based hits, increased the accuracy and quality of the functional genome annotation compared to results using only traditional annotation tools. Our workflow resulted in a comprehensive description of the V. necatrix genome, along with a structural summary of the most prevalent protein groups, such as the ricin B lectin family. In addition, and to test our tool, we identified the functions of several previously uncharacterized Encephalitozoon cuniculi genes.

CONCLUSION

We provide a new functional annotation tool for divergent organisms and employ it on a newly sequenced, high-quality microsporidian genome to shed light on this uncharacterized intracellular pathogen of Lepidoptera. The addition of a structure-based annotation approach can serve as a valuable template for studying other microsporidian or similarly divergent species.

Collapse

Wang J, Chen C, Yao G, Ding J, Wang L, Jiang H. Intelligent Protein Design and Molecular Characterization Techniques: A Comprehensive Review. Molecules 2023;28:7865. [PMID: 38067593 PMCID: PMC10707872 DOI: 10.3390/molecules28237865] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/21/2023] [Revised: 11/13/2023] [Accepted: 11/23/2023] [Indexed: 12/18/2023] Open

Zheng R, Huang Z, Deng L. Large-scale predicting protein functions through heterogeneous feature fusion. Brief Bioinform 2023:bbad243. [PMID: 37401369 DOI: 10.1093/bib/bbad243] [Citation(s) in RCA: 5] [Impact Index Per Article: 2.5] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/12/2023] [Revised: 05/18/2023] [Accepted: 06/12/2023] [Indexed: 07/05/2023] Open

Xu W, Zhao Z, Zhang H, Hu M, Yang N, Wang H, Wang C, Jiao J, Gu L. Deep neural learning based protein function prediction. MATHEMATICAL BIOSCIENCES AND ENGINEERING : MBE 2022;19:2471-2488. [PMID: 35240793 DOI: 10.3934/mbe.2022114] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 06/14/2023]

Affiliation(s)

Wenjun Xu Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China School of Life Sciences, Anhui Agricultural University, Hefei 230036, China
Zihao Zhao School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Hongwei Zhang School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Minglei Hu School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Ning Yang School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Hui Wang School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Chao Wang School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Jun Jiao School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China
Lichuan Gu School of Information and Computer, Anhui Agricultural University, Hefei 230036, China Key Laboratory of Agricultural Electronic Commerce, Ministry of Agriculture, Hefei 230036, China Institute of Intelligent Agriculture, Anhui Agricultural University, Hefei 230036, China School of Life Sciences, Anhui Agricultural University, Hefei 230036, China

Collapse

Hubert CB, de Carvalho LPS. Metabolomic approaches for enzyme function and pathway discovery in bacteria. Methods Enzymol 2022;665:29-47. [DOI: 10.1016/bs.mie.2021.12.001] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 11/24/2022]

Elhaj-Abdou MEM, El-Dib H, El-Helw A, El-Habrouk M. Deep_CNN_LSTM_GO: Protein function prediction from amino-acid sequences. Comput Biol Chem 2021;95:107584. [PMID: 34601431 DOI: 10.1016/j.compbiolchem.2021.107584] [Citation(s) in RCA: 4] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/05/2021] [Revised: 09/08/2021] [Accepted: 09/21/2021] [Indexed: 11/15/2022]

Mota APZ, Fernandez D, Arraes FBM, Petitot AS, de Melo BP, de Sa MEL, Grynberg P, Saraiva MAP, Guimaraes PM, Brasileiro ACM, Albuquerque EVS, Danchin EGJ, Grossi-de-Sa MF. Evolutionarily conserved plant genes responsive to root-knot nematodes identified by comparative genomics. Mol Genet Genomics 2020;295:1063-1078. [PMID: 32333171 DOI: 10.1007/s00438-020-01677-7] [Citation(s) in RCA: 10] [Impact Index Per Article: 2.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/13/2019] [Accepted: 04/04/2020] [Indexed: 01/11/2023]

Abstract

Root-knot nematodes (RKNs, genus Meloidogyne) affect a large number of crops causing severe yield losses worldwide, more specifically in tropical and sub-tropical regions. Several plant species display high resistance levels to Meloidogyne, but a general view of the plant immune molecular responses underlying resistance to RKNs is still lacking. Combining comparative genomics with differential gene expression analysis may allow the identification of widely conserved plant genes involved in RKN resistance. To identify genes that are evolutionary conserved across plant species, we used OrthoFinder to compared the predicted proteome of 22 plant species, including important crops, spanning 214 Myr of plant evolution. Overall, we identified 35,238 protein orthogroups, of which 6,132 were evolutionarily conserved and universal to all the 22 plant species (PLAnts Common Orthogroups-PLACO). To identify host genes responsive to RKN infection, we analyzed the RNA-seq transcriptome data from RKN-resistant genotypes of a peanut wild relative (Arachis stenosperma), coffee (Coffea arabica L.), soybean (Glycine max L.), and African rice (Oryza glaberrima Steud.) challenged by Meloidogyne spp. using EdgeR and DESeq tools, and we found 2,597 (O. glaberrima), 743 (C. arabica), 665 (A. stenosperma), and 653 (G. max) differentially expressed genes (DEGs) during the resistance response to the nematode. DEGs' classification into the previously characterized 35,238 protein orthogroups allowed identifying 17 orthogroups containing at least one DEG of each resistant Arachis, coffee, soybean, and rice genotype analyzed. Orthogroups contain 364 DEGs related to signaling, secondary metabolite production, cell wall-related functions, peptide transport, transcription regulation, and plant defense, thus revealing evolutionarily conserved RKN-responsive genes. Interestingly, the 17 DEGs-containing orthogroups (belonging to the PLACO) were also universal to the 22 plant species studied, suggesting that these core genes may be involved in ancestrally conserved immune responses triggered by RKN infection. The comparative genomic approach that we used here represents a promising predictive tool for the identification of other core plant defense-related genes of broad interest that are involved in different plant-pathogen interactions.

Collapse

Investigation of machine learning techniques on proteomics: A comprehensive survey. PROGRESS IN BIOPHYSICS AND MOLECULAR BIOLOGY 2019;149:54-69. [PMID: 31568792 DOI: 10.1016/j.pbiomolbio.2019.09.004] [Citation(s) in RCA: 7] [Impact Index Per Article: 1.2] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Subscribe] [Scholar Register] [Received: 06/05/2019] [Revised: 09/16/2019] [Accepted: 09/23/2019] [Indexed: 11/21/2022]

Saha S, Chatterjee P, Basu S, Nasipuri M, Plewczynski D. FunPred 3.0: improved protein function prediction using protein interaction network. PeerJ 2019;7:e6830. [PMID: 31198622 PMCID: PMC6535044 DOI: 10.7717/peerj.6830] [Citation(s) in RCA: 10] [Impact Index Per Article: 1.7] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/20/2018] [Accepted: 03/21/2019] [Indexed: 11/23/2022] Open

Abstract

Proteins are the most versatile macromolecules in living systems and perform crucial biological functions. In the advent of the post-genomic era, the next generation sequencing is done routinely at the population scale for a variety of species. The challenging problem is to massively determine the functions of proteins that are yet not characterized by detailed experimental studies. Identification of protein functions experimentally is a laborious and time-consuming task involving many resources. We therefore propose the automated protein function prediction methodology using in silico algorithms trained on carefully curated experimental datasets. We present the improved protein function prediction tool FunPred 3.0, an extended version of our previous methodology FunPred 2, which exploits neighborhood properties in protein–protein interaction network (PPIN) and physicochemical properties of amino acids. Our method is validated using the available functional annotations in the PPIN network of Saccharomyces cerevisiae in the latest Munich information center for protein (MIPS) dataset. The PPIN data of S. cerevisiae in MIPS dataset includes 4,554 unique proteins in 13,528 protein–protein interactions after the elimination of the self-replicating and the self-interacting protein pairs. Using the developed FunPred 3.0 tool, we are able to achieve the mean precision, the recall and the F-score values of 0.55, 0.82 and 0.66, respectively. FunPred 3.0 is then used to predict the functions of unpredicted protein pairs (incomplete and missing functional annotations) in MIPS dataset of S. cerevisiae. The method is also capable of predicting the subcellular localization of proteins along with its corresponding functions. The code and the complete prediction results are available freely at: https://github.com/SovanSaha/FunPred-3.0.git.

Collapse