1
|
Cridland JM, Polston ES, Begun DJ. New perspectives on Drosophila melanogaster de novo gene origination revealed by investigation of ancient African genetic variation. Genetics 2025; 230:iyaf044. [PMID: 40106667 PMCID: PMC12059636 DOI: 10.1093/genetics/iyaf044] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/11/2024] [Accepted: 03/04/2025] [Indexed: 03/22/2025] Open
Abstract
De novo genes can be defined as sequences producing evolutionarily derived transcripts that are not homologous to transcripts produced in an ancestor. While they appear to be taxonomically widespread, there is little agreement regarding their abundance, their persistence times in genomes, the population genetic processes responsible for their spread or loss, or their possible functions. In Drosophila melanogaster, 2 approaches have been used to discover these genes and investigate their properties. One uses traditional comparative approaches and existing genomic resources and annotations. A second approach uses raw transcriptome data to discover unannotated genes for which there is no evidence of presence in related species. Investigations using the second approach have focused on D. melanogaster genotypes from recently established cosmopolitan populations. However, most of the genetic variation in the species is found in African populations, suggesting the possibility that fuller understanding of genetic novelties in the species may follow from studies of these populations. Here, we investigate de novo gene candidates expressed in testis and accessory glands in a sample of flies from Zambia and compare them with candidate de novo genes expressed in North American populations. We report a large number of previously undiscovered de novo gene candidates, most of which are expressed polymorphically. Many are predicted to code for secreted proteins. In spite of much different levels of genomic variation in Zambian and North American populations, they express similar numbers of candidate de novo genes. We find evidence from genetic analysis of Raleigh inbred lines that a fraction of rarely expressed gene candidates in this population represent deleterious transcription promoted by inbreeding depression. Many de novo gene candidates are expressed in multiple tissues and both sexes, raising questions about how they may interact with natural selection. The relative importance of positive and negative selection, however, remains unclear.
Collapse
Affiliation(s)
- Julie M Cridland
- Department of Evolution and Ecology, University of California, Davis, Davis, CA 95616, USA
| | - Elizabeth S Polston
- Department of Evolution and Ecology, University of California, Davis, Davis, CA 95616, USA
| | - David J Begun
- Department of Evolution and Ecology, University of California, Davis, Davis, CA 95616, USA
| |
Collapse
|
2
|
Dohmen E, Aubel M, Eicholt LA, Roginski P, Luria V, Karger A, Grandchamp A. DeNoFo: a file format and toolkit for standardised, comparable de novo gene annotation. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2025:2025.03.31.644673. [PMID: 40236033 PMCID: PMC11996330 DOI: 10.1101/2025.03.31.644673] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 04/17/2025]
Abstract
Motivation De novo genes emerge from previously non-coding regions of the genome, challenging the traditional view that new genes primarily arise through duplication and adaptation of existing ones. Characterised by their rapid evolution and their novel structural properties or functional roles, de novo genes represent a young area of research. Therefore, the field currently lacks established standards and methodologies, leading to inconsistent terminology and challenges in comparing and reproducing results. Results This work presents a standardised annotation format to document the methodology of de novo gene datasets in a reproducible way. We developed DeNoFo, a toolkit to provide easy access to this format that simplifies annotation of datasets and facilitates comparison across studies. Unifying the different protocols and methods in one standardised format, while providing integration into established file formats, such as fasta or gff, ensures comparability of studies and advances new insights in this rapidly evolving field. Availability and Implementation DeNoFo is available through the official Python Package Index (PyPI) and at https://github.com/EDohmen/denofo . All tools have a graphical user interface and a command line interface. The toolkit is implemented in Python3, available for all major platforms and installable with pip and uv.
Collapse
|
3
|
Hannon Bozorgmehr J. The De Novo Emergence of Two Brain Genes in the Human Lineage Appears to be Unsupported. J Mol Evol 2025; 93:3-10. [PMID: 39725692 DOI: 10.1007/s00239-024-10227-3] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/04/2024] [Accepted: 12/10/2024] [Indexed: 12/28/2024]
Abstract
Recently, certain studies have claimed that cognitive features and pathologies unique to humans can be traced to certain changes in the nervous system. These are caused by genes that have likely evolved "from scratch," not having any coding precursors. The translated proteins would not appear outside of the human lineage and any orthologs in other species should be non-coding. This contrasts with research that has identified a decisive role for duplication, and modifications to regulatory sequences, for such phenotypic traits. Closer examination, however, reveals that the inferred lineage-specific emergence of at least two of these genes is likely a misinterpretation owing to a lack of peptide verification, experimental oversights, and insufficient species comparisons. A possible pseudogenic origin is proposed for one of them. The implications of these claims for the study of molecular evolution are discussed.
Collapse
|
4
|
Cherezov RO, Vorontsova JE, Kuvaeva EE, Akishina AA, Zavoloka EL, Simonova OB. The lawc gene emerged de novo from conserved genomic elements and acquired a broad expression pattern in Drosophila. J Genet Genomics 2024:S1673-8527(24)00367-9. [PMID: 39733859 DOI: 10.1016/j.jgg.2024.12.014] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/02/2024] [Revised: 12/17/2024] [Accepted: 12/18/2024] [Indexed: 12/31/2024]
Abstract
It has recently become evident that the de novo emergence of genes is widespread and documented for a variety of organisms. De novo genes frequently emerge in proximity to existing genes, forming gene overlaps. Here, we present an analysis of the evolutionary history of a putative de novo gene, lawc, which overlaps with the conserved Trf2 gene, which encodes a general transcription factor in Drosophila melanogaster. We demonstrate that lawc emerged approximately 68 million years ago in the 5'-untranslated region (UTR) of Trf2 and displays an extensive spatiotemporal expression pattern. One of the most remarkable features of the lawc evolutionary history is that its emergence was facilitated by the engagement of Drosophilidae-specific short, highly conserved regions located in Trf2 introns. This represents a unique example of putative de novo gene birth involving conserved DNA regions localized in introns of conserved genes. The observed lawc expression pattern may be due to the overlap of lawc with the 5'-UTR of Trf2. This study not only enriches our understanding of gene evolution but also highlights the complex interplay between genetic conservation and innovation.
Collapse
Affiliation(s)
- Roman O Cherezov
- Kol'tsov Institute of Developmental Biology, Russian Academy of Sciences, Moscow, 119334, Russia.
| | - Julia E Vorontsova
- Institute of Gene Biology, Russian Academy of Sciences, Moscow, 119334, Russia
| | - Elena E Kuvaeva
- Kol'tsov Institute of Developmental Biology, Russian Academy of Sciences, Moscow, 119334, Russia
| | - Angelina A Akishina
- Kol'tsov Institute of Developmental Biology, Russian Academy of Sciences, Moscow, 119334, Russia
| | - Ekaterina L Zavoloka
- Kol'tsov Institute of Developmental Biology, Russian Academy of Sciences, Moscow, 119334, Russia
| | - Olga B Simonova
- Institute of Gene Biology, Russian Academy of Sciences, Moscow, 119334, Russia
| |
Collapse
|
5
|
Naidu P, Holford M. Microscopic marvels: Decoding the role of micropeptides in innate immunity. Immunology 2024; 173:605-621. [PMID: 39188052 DOI: 10.1111/imm.13850] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/11/2024] [Accepted: 07/30/2024] [Indexed: 08/28/2024] Open
Abstract
The innate immune response is under selection pressures from changing environments and pathogens. While sequence evolution can be studied by comparing rates of amino acid mutations within and between species, how a gene's birth and death contribute to the evolution of immunity is less known. Short open reading frames, once regarded as untranslated or transcriptional noise, can often produce micropeptides of <100 amino acids with a wide array of biological functions. Some micropeptide sequences are well conserved, whereas others have no evolutionary conservation, potentially representing new functional compounds that arise from species-specific adaptations. To date, few reports have described the discovery of novel micropeptides of the innate immune system. The diversity of immune-related micropeptides is a blind spot for gene and functional annotation. Immune-related micropeptides represent a potential reservoir of untapped compounds for understanding and treating disease. This review consolidates what is currently known about the evolution and function of innate immune-related micropeptides to facilitate their investigation.
Collapse
Affiliation(s)
- Praveena Naidu
- Graduate Center, Programs in Biology, Biochemistry, Chemistry, City University of New York, New York, New York, USA
- Department of Chemistry and Biochemistry, City University of New York, Hunter College, Belfer Research Building, New York, New York, USA
| | - Mandë Holford
- Graduate Center, Programs in Biology, Biochemistry, Chemistry, City University of New York, New York, New York, USA
- Department of Chemistry and Biochemistry, City University of New York, Hunter College, Belfer Research Building, New York, New York, USA
- American Museum of Natural History, Invertebrate Zoology, Sackler Institute for Comparative Genomics, New York, New York, USA
- Weill Cornell Medicine, Department of Biochemistry, New York, New York, USA
| |
Collapse
|
6
|
Cao Y, Hong J, Zhao Y, Li X, Feng X, Wang H, Zhang L, Lin M, Cai Y, Han Y. De novo gene integration into regulatory networks via interaction with conserved genes in peach. HORTICULTURE RESEARCH 2024; 11:uhae252. [PMID: 39664695 PMCID: PMC11630308 DOI: 10.1093/hr/uhae252] [Citation(s) in RCA: 11] [Impact Index Per Article: 11.0] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Figures] [Subscribe] [Scholar Register] [Received: 05/21/2024] [Accepted: 08/29/2024] [Indexed: 12/13/2024]
Abstract
De novo genes can evolve "from scratch" from noncoding sequences, acquiring novel functions in organisms and integrating into regulatory networks during evolution to drive innovations in important phenotypes and traits. However, identifying de novo genes is challenging, as it requires high-quality genomes from closely related species. According to the comparison with nine closely related Prunus genomes, we determined at least 178 de novo genes in P. persica "baifeng". The distinct differences were observed between de novo and conserved genes in gene characteristics and expression patterns. Gene ontology enrichment analysis suggested that Type I de novo genes originated from sequences related to plastid modification functions, while Type II genes were inferred to have derived from sequences related to reproductive functions. Finally, transcriptome sequencing across different tissues and developmental stages suggested that de novo genes have been evolutionarily recruited into existing regulatory networks, playing important roles in plant growth and development, which was also supported by WGCNA analysis and quantitative trait loci data. This study lays the groundwork for future research on the origins and functions of genes in Prunus and related taxa.
Collapse
Affiliation(s)
- Yunpeng Cao
- State Key Laboratory of Plant Diversity and Specialty Crops, Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan 430074, China
| | - Jiayi Hong
- College of Life Sciences, Anhui Agricultural University, Hefei 230036, China
| | - Yun Zhao
- State Key Laboratory of Plant Diversity and Specialty Crops, Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan 430074, China
| | - Xiaoxu Li
- Beijing Life Science Academy, Beijing 102209, China
| | - Xiaofeng Feng
- College of Life Sciences, Anhui Agricultural University, Hefei 230036, China
| | - Han Wang
- Key Laboratory of Horticultural Crop Germplasm Innovation and Utilization (Co-construction by Ministry and Province), Institute of Horticulture, Anhui Academy of Agricultural Sciences, Hefei 230000, China
| | - Lin Zhang
- Hubei Shizhen Laboratory, Hubei Key Laboratory of Theory and Application Research of Liver and Kidney in Traditional Chinese Medicine, School of Basic Medical Sciences, Hubei University of Chinese Medicine, Wuhan 430065, China
| | - Mengfei Lin
- Jiangxi Provincial Key Laboratory of Plantation and High Valued Utilization of Specialty Fruit Tree and Tea, Institute of Biological Resources, Jiangxi Academy of Sciences, Nanchang 330224 Jiangxi, China
| | - Yongping Cai
- College of Life Sciences, Anhui Agricultural University, Hefei 230036, China
| | - Yuepeng Han
- State Key Laboratory of Plant Diversity and Specialty Crops, Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan 430074, China
| |
Collapse
|
7
|
Houghton CJ, Coelho NC, Chiang A, Hedayati S, Parikh SB, Ozbaki-Yagan N, Wacholder A, Iannotta J, Berger A, Carvunis AR, O’Donnell AF. Cellular processing of beneficial de novo emerging proteins. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2024:2024.08.28.610198. [PMID: 39257767 PMCID: PMC11384008 DOI: 10.1101/2024.08.28.610198] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 09/12/2024]
Abstract
Novel proteins can originate de novo from non-coding DNA and contribute to species-specific adaptations. It is challenging to conceive how de novo emerging proteins may integrate pre-existing cellular systems to bring about beneficial traits, given that their sequences are previously unseen by the cell. To address this apparent paradox, we investigated 26 de novo emerging proteins previously associated with growth benefits in yeast. Microscopy revealed that these beneficial emerging proteins preferentially localize to the endoplasmic reticulum (ER). Sequence and structure analyses uncovered a common protein organization among all ER-localizing beneficial emerging proteins, characterized by a short hydrophobic C-terminus immediately preceded by a transmembrane domain. Using genetic and biochemical approaches, we showed that ER localization of beneficial emerging proteins requires the GET and SND pathways, both of which are evolutionarily conserved and known to recognize transmembrane domains to promote post-translational ER insertion. The abundance of ER-localizing beneficial emerging proteins was regulated by conserved proteasome- and vacuole-dependent processes, through mechanisms that appear to be facilitated by the emerging proteins' C-termini. Consequently, we propose that evolutionarily conserved pathways can convergently govern the cellular processing of de novo emerging proteins with unique sequences, likely owing to common underlying protein organization patterns.
Collapse
Affiliation(s)
- Carly J. Houghton
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Nelson Castilho Coelho
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Annette Chiang
- Department of Biological Sciences, University of Pittsburgh, Pittsburgh, PA 15260, United States
| | - Stefanie Hedayati
- Department of Biological Sciences, University of Pittsburgh, Pittsburgh, PA 15260, United States
| | - Saurin B. Parikh
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Nejla Ozbaki-Yagan
- Department of Biological Sciences, University of Pittsburgh, Pittsburgh, PA 15260, United States
| | - Aaron Wacholder
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - John Iannotta
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Alexis Berger
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Anne-Ruxandra Carvunis
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
| | - Allyson F. O’Donnell
- Pittsburgh Center for Evolutionary Biology and Medicine (CEBaM), Department of Computational and Systems Biology, School of Medicine, University of Pittsburgh, Pittsburgh, PA, 15213, United States
- Department of Biological Sciences, University of Pittsburgh, Pittsburgh, PA 15260, United States
| |
Collapse
|
8
|
Roginski P, Grandchamp A, Quignot C, Lopes A. De Novo Emerged Gene Search in Eukaryotes with DENSE. Genome Biol Evol 2024; 16:evae159. [PMID: 39212967 PMCID: PMC11363675 DOI: 10.1093/gbe/evae159] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Accepted: 07/07/2024] [Indexed: 09/04/2024] Open
Abstract
The discovery of de novo emerged genes, originating from previously noncoding DNA regions, challenges traditional views of species evolution. Indeed, the hypothesis of neutrally evolving sequences giving rise to functional proteins is highly unlikely. This conundrum has sparked numerous studies to quantify and characterize these genes, aiming to understand their functional roles and contributions to genome evolution. Yet, no fully automated pipeline for their identification is available. Therefore, we introduce DENSE (DE Novo emerged gene SEarch), an automated Nextflow pipeline based on two distinct steps: detection of taxonomically restricted genes (TRGs) through phylostratigraphy, and filtering of TRGs for de novo emerged genes via genome comparisons and synteny search. DENSE is available as a user-friendly command-line tool, while the second step is accessible through a web server upon providing a list of TRGs. Highly flexible, DENSE provides various strategy and parameter combinations, enabling users to adapt to specific configurations or define their own strategy through a rational framework, facilitating protocol communication, and study interoperability. We apply DENSE to seven model organisms, exploring the impact of its strategies and parameters on de novo gene predictions. This thorough analysis across species with different evolutionary rates reveals useful metrics for users to define input datasets, identify favorable/unfavorable conditions for de novo gene detection, and control potential biases in genome annotations. Additionally, predictions made for the seven model organisms are compiled into a requestable database, which we hope will serve as a reference for de novo emerged gene lists generated with specific criteria combinations.
Collapse
Affiliation(s)
- Paul Roginski
- Institute for Integrative Biology of the Cell (I2BC), Université Paris-Saclay, CEA, CNRS, 91198 Gif-sur-Yvette, France
| | - Anna Grandchamp
- Institute for Evolution and Biodiversity, University of Münster, 48149 Münster, Germany
| | - Chloé Quignot
- Institute for Integrative Biology of the Cell (I2BC), Université Paris-Saclay, CEA, CNRS, 91198 Gif-sur-Yvette, France
| | - Anne Lopes
- Institute for Integrative Biology of the Cell (I2BC), Université Paris-Saclay, CEA, CNRS, 91198 Gif-sur-Yvette, France
| |
Collapse
|
9
|
Middendorf L, Ravi Iyengar B, Eicholt LA. Sequence, Structure, and Functional Space of Drosophila De Novo Proteins. Genome Biol Evol 2024; 16:evae176. [PMID: 39212966 PMCID: PMC11363682 DOI: 10.1093/gbe/evae176] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Accepted: 07/29/2024] [Indexed: 09/04/2024] Open
Abstract
During de novo emergence, new protein coding genes emerge from previously nongenic sequences. The de novo proteins they encode are dissimilar in composition and predicted biochemical properties to conserved proteins. However, functional de novo proteins indeed exist. Both identification of functional de novo proteins and their structural characterization are experimentally laborious. To identify functional and structured de novo proteins in silico, we applied recently developed machine learning based tools and found that most de novo proteins are indeed different from conserved proteins both in their structure and sequence. However, some de novo proteins are predicted to adopt known protein folds, participate in cellular reactions, and to form biomolecular condensates. Apart from broadening our understanding of de novo protein evolution, our study also provides a large set of testable hypotheses for focused experimental studies on structure and function of de novo proteins in Drosophila.
Collapse
Affiliation(s)
- Lasse Middendorf
- Institute for Evolution and Biodiversity, University of Muenster, Huefferstrasse 1, 48149 Muenster, Germany
| | - Bharat Ravi Iyengar
- Institute for Evolution and Biodiversity, University of Muenster, Huefferstrasse 1, 48149 Muenster, Germany
| | - Lars A Eicholt
- Institute for Evolution and Biodiversity, University of Muenster, Huefferstrasse 1, 48149 Muenster, Germany
| |
Collapse
|
10
|
Chen J, Li Q, Xia S, Arsala D, Sosa D, Wang D, Long M. The Rapid Evolution of De Novo Proteins in Structure and Complex. Genome Biol Evol 2024; 16:evae107. [PMID: 38753069 PMCID: PMC11149777 DOI: 10.1093/gbe/evae107] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Accepted: 05/10/2024] [Indexed: 06/06/2024] Open
Abstract
Recent studies in the rice genome-wide have established that de novo genes, evolving from noncoding sequences, enhance protein diversity through a stepwise process. However, the pattern and rate of their evolution in protein structure over time remain unclear. Here, we addressed these issues within a surprisingly short evolutionary timescale (<1 million years for 97% of Oryza de novo genes) with comparative approaches to gene duplicates. We found that de novo genes evolve faster than gene duplicates in the intrinsically disordered regions (such as random coils), secondary structure elements (such as α helix and β strand), hydrophobicity, and molecular recognition features. In de novo proteins, specifically, we observed an 8% to 14% decay in random coils and intrinsically disordered region lengths and a 2.3% to 6.5% increase in structured elements, hydrophobicity, and molecular recognition features, per million years on average. These patterns of structural evolution align with changes in amino acid composition over time as well. We also revealed higher positive charges but smaller molecular weights for de novo proteins than duplicates. Tertiary structure predictions showed that most de novo proteins, though not typically well folded on their own, readily form low-energy and compact complexes with other proteins facilitated by extensive residue contacts and conformational flexibility, suggesting a faster-binding scenario in de novo proteins to promote interaction. These analyses illuminate a rapid evolution of protein structure in de novo genes in rice genomes, originating from noncoding sequences, highlighting their quick transformation into active, protein complex-forming components within a remarkably short evolutionary timeframe.
Collapse
Affiliation(s)
- Jianhai Chen
- Department of Ecology and Evolution, The University of Chicago, Chicago, IL 60637, USA
| | - Qingrong Li
- Division of Pharmaceutical Sciences, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego, La Jolla, CA 92093, USA
- Department of Cellular & Molecular Medicine, School of Medicine, University of California San Diego, La Jolla, CA 92093, USA
| | - Shengqian Xia
- Department of Ecology and Evolution, The University of Chicago, Chicago, IL 60637, USA
| | - Deanna Arsala
- Department of Ecology and Evolution, The University of Chicago, Chicago, IL 60637, USA
| | - Dylan Sosa
- Department of Ecology and Evolution, The University of Chicago, Chicago, IL 60637, USA
| | - Dong Wang
- Division of Pharmaceutical Sciences, Skaggs School of Pharmacy and Pharmaceutical Sciences, University of California San Diego, La Jolla, CA 92093, USA
- Department of Cellular & Molecular Medicine, School of Medicine, University of California San Diego, La Jolla, CA 92093, USA
| | - Manyuan Long
- Department of Ecology and Evolution, The University of Chicago, Chicago, IL 60637, USA
| |
Collapse
|
11
|
Lee U, Mozeika SM, Zhao L. A Synergistic, Cultivator Model of De Novo Gene Origination. Genome Biol Evol 2024; 16:evae103. [PMID: 38748819 PMCID: PMC11152449 DOI: 10.1093/gbe/evae103] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Accepted: 05/12/2024] [Indexed: 06/07/2024] Open
Abstract
The origin and fixation of evolutionarily young genes is a fundamental question in evolutionary biology. However, understanding the origins of newly evolved genes arising de novo from noncoding genomic sequences is challenging. This is partly due to the low likelihood that several neutral or nearly neutral mutations fix prior to the appearance of an important novel molecular function. This issue is particularly exacerbated in large effective population sizes where the effect of drift is small. To address this problem, we propose a regulation-focused, cultivator model for de novo gene evolution. This cultivator-focused model posits that each step in a novel variant's evolutionary trajectory is driven by well-defined, selectively advantageous functions for the cultivator genes, rather than solely by the de novo genes, emphasizing the critical role of genome organization in the evolution of new genes.
Collapse
Affiliation(s)
- UnJin Lee
- Laboratory of Evolutionary Genetics and Genomics, The Rockefeller University, New York, NY, USA
| | - Shawn M Mozeika
- Laboratory of Evolutionary Genetics and Genomics, The Rockefeller University, New York, NY, USA
| | - Li Zhao
- Laboratory of Evolutionary Genetics and Genomics, The Rockefeller University, New York, NY, USA
| |
Collapse
|
12
|
Liu X, Xiao C, Xu X, Zhang J, Mo F, Chen JY, Delihas N, Zhang L, An NA, Li CY. Origin of functional de novo genes in humans from "hopeful monsters". WILEY INTERDISCIPLINARY REVIEWS. RNA 2024; 15:e1845. [PMID: 38605485 DOI: 10.1002/wrna.1845] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Received: 08/24/2023] [Revised: 03/13/2024] [Accepted: 03/18/2024] [Indexed: 04/13/2024]
Abstract
For a long time, it was believed that new genes arise only from modifications of preexisting genes, but the discovery of de novo protein-coding genes that originated from noncoding DNA regions demonstrates the existence of a "motherless" origination process for new genes. However, the features, distributions, expression profiles, and origin modes of these genes in humans seem to support the notion that their origin is not a purely "motherless" process; rather, these genes arise preferentially from genomic regions encoding preexisting precursors with gene-like features. In such a case, the gene loci are typically not brand new. In this short review, we will summarize the definition and features of human de novo genes and clarify their process of origination from ancestral non-coding genomic regions. In addition, we define the favored precursors, or "hopeful monsters," for the origin of de novo genes and present a discussion of the functional significance of these young genes in brain development and tumorigenesis in humans. This article is categorized under: RNA Evolution and Genomics > RNA and Ribonucleoprotein Evolution.
Collapse
Affiliation(s)
- Xiaoge Liu
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
| | - Chunfu Xiao
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
| | - Xinwei Xu
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
| | - Jie Zhang
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
| | - Fan Mo
- State Key Laboratory of Stem Cell and Reproductive Biology, Institute of Stem Cell and Regeneration, Institute of Zoology, Chinese Academy of Sciences, Beijing, China
| | - Jia-Yu Chen
- State Key Laboratory of Pharmaceutical Biotechnology, School of Life Sciences, Chemistry and Biomedicine Innovation Center (ChemBIC), Nanjing University, Nanjing, China
| | - Nicholas Delihas
- Department of Microbiology and Immunology, Renaissance School of Medicine, Stony Brook University, Stony Brook, New York, USA
| | - Li Zhang
- Chinese Institute for Brain Research, Beijing, China
| | - Ni A An
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
| | - Chuan-Yun Li
- State Key Laboratory of Protein and Plant Gene Research, Laboratory of Bioinformatics and Genomic Medicine, Institute of Molecular Medicine, College of Future Technology, Peking University, Beijing, China
- Chinese Institute for Brain Research, Beijing, China
- Southwest United Graduate School, Kunming, China
| |
Collapse
|
13
|
Martin NS, Camargo CQ, Louis AA. Bias in the arrival of variation can dominate over natural selection in Richard Dawkins's biomorphs. PLoS Comput Biol 2024; 20:e1011893. [PMID: 38536880 PMCID: PMC10971585 DOI: 10.1371/journal.pcbi.1011893] [Citation(s) in RCA: 1] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/23/2023] [Accepted: 02/02/2024] [Indexed: 11/12/2024] Open
Abstract
Biomorphs, Richard Dawkins's iconic model of morphological evolution, are traditionally used to demonstrate the power of natural selection to generate biological order from random mutations. Here we show that biomorphs can also be used to illustrate how developmental bias shapes adaptive evolutionary outcomes. In particular, we find that biomorphs exhibit phenotype bias, a type of developmental bias where certain phenotypes can be many orders of magnitude more likely than others to appear through random mutations. Moreover, this bias exhibits a strong preference for simpler phenotypes with low descriptional complexity. Such bias towards simplicity is formalised by an information-theoretic principle that can be intuitively understood from a picture of evolution randomly searching in the space of algorithms. By using population genetics simulations, we demonstrate how moderately adaptive phenotypic variation that appears more frequently upon random mutations can fix at the expense of more highly adaptive biomorph phenotypes that are less frequent. This result, as well as many other patterns found in the structure of variation for the biomorphs, such as high mutational robustness and a positive correlation between phenotype evolvability and robustness, closely resemble findings in molecular genotype-phenotype maps. Many of these patterns can be explained with an analytic model based on constrained and unconstrained sections of the genome. We postulate that the phenotype bias towards simplicity and other patterns biomorphs share with molecular genotype-phenotype maps may hold more widely for developmental systems.
Collapse
Affiliation(s)
- Nora S. Martin
- Rudolf Peierls Centre for Theoretical Physics, University of Oxford, Oxford, United Kingdom
| | - Chico Q. Camargo
- Rudolf Peierls Centre for Theoretical Physics, University of Oxford, Oxford, United Kingdom
- College of Engineering, Mathematics and Physical Sciences, University of Exeter, Exeter, United Kingdom
| | - Ard A. Louis
- Rudolf Peierls Centre for Theoretical Physics, University of Oxford, Oxford, United Kingdom
| |
Collapse
|
14
|
Alvarez S, Nartey CM, Mercado N, de la Paz JA, Huseinbegovic T, Morcos F. In vivo functional phenotypes from a computational epistatic model of evolution. Proc Natl Acad Sci U S A 2024; 121:e2308895121. [PMID: 38285950 PMCID: PMC10861889 DOI: 10.1073/pnas.2308895121] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/26/2023] [Accepted: 12/19/2023] [Indexed: 01/31/2024] Open
Abstract
Computational models of evolution are valuable for understanding the dynamics of sequence variation, to infer phylogenetic relationships or potential evolutionary pathways and for biomedical and industrial applications. Despite these benefits, few have validated their propensities to generate outputs with in vivo functionality, which would enhance their value as accurate and interpretable evolutionary algorithms. We demonstrate the power of epistasis inferred from natural protein families to evolve sequence variants in an algorithm we developed called sequence evolution with epistatic contributions (SEEC). Utilizing the Hamiltonian of the joint probability of sequences in the family as fitness metric, we sampled and experimentally tested for in vivo [Formula: see text]-lactamase activity in Escherichia coli TEM-1 variants. These evolved proteins can have dozens of mutations dispersed across the structure while preserving sites essential for both catalysis and interactions. Remarkably, these variants retain family-like functionality while being more active than their wild-type predecessor. We found that depending on the inference method used to generate the epistatic constraints, different parameters simulate diverse selection strengths. Under weaker selection, local Hamiltonian fluctuations reliably predict relative changes to variant fitness, recapitulating neutral evolution. SEEC has the potential to explore the dynamics of neofunctionalization, characterize viral fitness landscapes, and facilitate vaccine development.
Collapse
Affiliation(s)
- Sophia Alvarez
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX75080
| | - Charisse M. Nartey
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX75080
| | - Nicholas Mercado
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX75080
| | | | - Tea Huseinbegovic
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX75080
| | - Faruck Morcos
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX75080
- Department of Bioengineering, University of Texas at Dallas, Richardson, TX75080
- Center for Systems Biology, University of Texas at Dallas, Richardson, TX75080
| |
Collapse
|
15
|
Hannon Bozorgmehr J. Four classic "de novo" genes all have plausible homologs and likely evolved from retro-duplicated or pseudogenic sequences. Mol Genet Genomics 2024; 299:6. [PMID: 38315248 DOI: 10.1007/s00438-023-02090-6] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/27/2023] [Accepted: 10/15/2023] [Indexed: 02/07/2024]
Abstract
Despite being previously regarded as extremely unlikely, the idea that entirely novel protein-coding genes can emerge from non-coding sequences has gradually become accepted over the past two decades. Examples of "de novo origination", resulting in lineage-specific "orphan" genes, lacking coding orthologs, are now produced every year. However, many are likely cases of duplicates that are difficult to recognize. Here, I re-examine the claims and show that four very well-known examples of genes alleged to have emerged completely "from scratch"- FLJ33706 in humans, Goddard in fruit flies, BSC4 in baker's yeast and AFGP2 in codfish-may have plausible evolutionary ancestors in pre-existing genes. The first two are likely highly diverged retrogenes coding for regulatory proteins that have been misidentified as orphans. The antifreeze glycoprotein, moreover, may not have evolved from repetitive non-genic sequences but, as in several other related cases, from an apolipoprotein that could have become pseudogenized before later being reactivated. These findings detract from various claims made about de novo gene birth and show there has been a tendency not to invest the necessary effort in searching for homologs outside of a very limited syntenic or phylostratigraphic methodology. A robust approach is used for improving detection that draws upon similarities, not just in terms of statistical sequence analysis, but also relating to biochemistry and function, to obviate notable failures to identify homologs.
Collapse
|
16
|
Mönttinen HAM, Frilander MJ, Löytynoja A. Generation of de novo miRNAs from template switching during DNA replication. Proc Natl Acad Sci U S A 2023; 120:e2310752120. [PMID: 38019864 PMCID: PMC10710096 DOI: 10.1073/pnas.2310752120] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 06/26/2023] [Accepted: 11/01/2023] [Indexed: 12/01/2023] Open
Abstract
The mechanisms generating novel genes and genetic information are poorly known, even for microRNA (miRNA) genes with an extremely constrained design. All miRNA primary transcripts need to fold into a stem-loop structure to yield short gene products ([Formula: see text]22 nt) that bind and repress their mRNA targets. While a substantial number of miRNA genes are ancient and highly conserved, short secondary structures coding for entirely novel miRNA genes have been shown to emerge in a lineage-specific manner. Template switching is a DNA-replication-related mutation mechanism that can introduce complex changes and generate perfect base pairing for entire hairpin structures in a single event. Here, we show that the template-switching mutations (TSMs) have participated in the emergence of over 6,000 suitable hairpin structures in the primate lineage to yield at least 18 new human miRNA genes, that is 26% of the miRNAs inferred to have arisen since the origin of primates. While the mechanism appears random, the TSM-generated miRNAs are enriched in introns where they can be expressed with their host genes. The high frequency of TSM events provides raw material for evolution. Being orders of magnitude faster than other mechanisms proposed for de novo creation of genes, TSM-generated miRNAs enable near-instant rewiring of genetic information and rapid adaptation to changing environments.
Collapse
Affiliation(s)
- Heli A. M. Mönttinen
- Institute of Biotechnology, Helsinki Institute of Life Science, University of Helsinki, HelsinkiFI-000, Finland
| | - Mikko J. Frilander
- Institute of Biotechnology, Helsinki Institute of Life Science, University of Helsinki, HelsinkiFI-000, Finland
| | - Ari Löytynoja
- Institute of Biotechnology, Helsinki Institute of Life Science, University of Helsinki, HelsinkiFI-000, Finland
| |
Collapse
|
17
|
Kore H, Datta KK, Nagaraj SH, Gowda H. Protein-coding potential of non-canonical open reading frames in human transcriptome. Biochem Biophys Res Commun 2023; 684:149040. [PMID: 37897910 DOI: 10.1016/j.bbrc.2023.09.068] [Citation(s) in RCA: 1] [Impact Index Per Article: 0.5] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/13/2023] [Revised: 09/09/2023] [Accepted: 09/23/2023] [Indexed: 10/30/2023]
Abstract
In recent years, proteogenomics and ribosome profiling studies have identified a large number of proteins encoded by noncoding regions in the human genome. They are encoded by small open reading frames (sORFs) in the untranslated regions (UTRs) of mRNAs and long non-coding RNAs (lncRNAs). These sORF encoded proteins (SEPs) are often <150AA and show poor evolutionary conservation. A subset of them have been functionally characterized and shown to play an important role in fundamental biological processes including cardiac and muscle function, DNA repair, embryonic development and various human diseases. How many novel protein-coding regions exist in the human genome and what fraction of them are functionally important remains a mystery. In this review, we discuss current progress in unraveling SEPs, approaches used for their identification, their limitations and reliability of these identifications. We also discuss functionally characterized SEPs and their involvement in various biological processes and diseases. Lastly, we provide insights into their distinctive features compared to canonical proteins and challenges associated with annotating these in protein reference databases.
Collapse
Affiliation(s)
- Hitesh Kore
- Centre for Genomics and Personalised Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia; Cancer Precision Medicine Group, QIMR Berghofer Medical Research Institute, 300 Herston Road, Herston, Queensland, 4006, Australia; Faculty of Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia.
| | - Keshava K Datta
- Proteomics and Metabolomics Platform, La Trobe University, Melbourne, VIC, 3083, Australia
| | - Shivashankar H Nagaraj
- Centre for Genomics and Personalised Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia; Faculty of Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia
| | - Harsha Gowda
- Centre for Genomics and Personalised Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia; Cancer Precision Medicine Group, QIMR Berghofer Medical Research Institute, 300 Herston Road, Herston, Queensland, 4006, Australia; Faculty of Health, Queensland University of Technology, Brisbane, Queensland, 4059, Australia; Faculty of Medicine, The University of Queensland, Queensland, 4072, Australia.
| |
Collapse
|
18
|
Hlouchova K. Toxin rescue by a random sequence. Nat Ecol Evol 2023; 7:1963-1964. [PMID: 37945945 DOI: 10.1038/s41559-023-02252-0] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 11/12/2023]
Affiliation(s)
- Klara Hlouchova
- Department of Cell Biology, Faculty of Science, Charles University, BIOCEV, Prague, Czech Republic.
- Institute of Organic Chemistry and Biochemistry, Czech Academy of Sciences, Prague, Czech Republic.
| |
Collapse
|
19
|
Frumkin I, Laub MT. Selection of a de novo gene that can promote survival of Escherichia coli by modulating protein homeostasis pathways. Nat Ecol Evol 2023; 7:2067-2079. [PMID: 37945946 PMCID: PMC10697842 DOI: 10.1038/s41559-023-02224-4] [Citation(s) in RCA: 12] [Impact Index Per Article: 6.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 02/11/2023] [Accepted: 09/12/2023] [Indexed: 11/12/2023]
Abstract
Cellular novelty can emerge when non-functional loci become functional genes in a process termed de novo gene birth. But how proteins with random amino acid sequences beneficially integrate into existing cellular pathways remains poorly understood. We screened ~108 genes, generated from random nucleotide sequences and devoid of homology to natural genes, for their ability to rescue growth arrest of Escherichia coli cells producing the ribonuclease toxin MazF. We identified ~2,000 genes that could promote growth, probably by reducing transcription from the promoter driving toxin expression. Additionally, one random protein, named Random antitoxin of MazF (RamF), modulated protein homeostasis by interacting with chaperones, leading to MazF proteolysis and a consequent loss of its toxicity. Finally, we demonstrate that random proteins can improve during evolution by identifying beneficial mutations that turned RamF into a more efficient inhibitor. Our work provides a mechanistic basis for how de novo gene birth can produce functional proteins that effectively benefit cells evolving under stress.
Collapse
Affiliation(s)
- Idan Frumkin
- Department of Biology, Massachusetts Institute of Technology, Cambridge, MA, USA
| | - Michael T Laub
- Department of Biology, Massachusetts Institute of Technology, Cambridge, MA, USA.
- Howard Hughes Medical Institute, Cambridge, MA, USA.
| |
Collapse
|
20
|
Yamamoto S, Kono F, Nakatani K, Hirose M, Horii K, Hippo Y, Tamada T, Suenaga Y, Matsuo T. Structural characterization of human de novo protein NCYM and its complex with a newly identified DNA aptamer using atomic force microscopy and small-angle X-ray scattering. Front Oncol 2023; 13:1213678. [PMID: 38074684 PMCID: PMC10701690 DOI: 10.3389/fonc.2023.1213678] [Citation(s) in RCA: 5] [Impact Index Per Article: 2.5] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 04/28/2023] [Accepted: 11/03/2023] [Indexed: 02/04/2025] Open
Abstract
NCYM, a Homininae-specific oncoprotein, is the first de novo gene product experimentally shown to have oncogenic functions. NCYM stabilizes MYCN and β-catenin via direct binding and inhibition of GSK3β and promotes cancer progression in various tumors. Thus, the identification of compounds that binds to NCYM and structural characterization of the complex of such compounds with NCYM are required to deepen our understanding of the molecular mechanism of NCYM function and eventually to develop anticancer drugs against NCYM. In this study, the DNA aptamer that specifically binds to NCYM and enhances interaction between NCYM and GSK3β were identified for the first time using systematic evolution of ligands by exponential enrichment (SELEX). The structural properties of the complex of the aptamer and NCYM were investigated using atomic force microscopy (AFM) in combination with truncation and mutation of DNA sequence, pointing to the regions on the aptamer required for NCYM binding. Further analysis was carried out by small-angle X-ray scattering (SAXS). Structural modeling based on SAXS data revealed that when isolated, NCYM shows high flexibility, though not as a random coil, while the DNA aptamer exists as a dimer in solution. In the complex state, models in which NCYM was bound to a region close to an edge of the aptamer reproduced the SAXS data. Therefore, using a combination of SELEX, AFM, and SAXS, the present study revealed the structural properties of NCYM in its functionally active form, thus providing useful information for the possible future design of novel anti-cancer drugs targeting NCYM.
Collapse
Affiliation(s)
- Seigi Yamamoto
- Laboratory of Evolutionary Oncology, Chiba Cancer Center Research Institute, Chiba, Japan
| | - Fumiaki Kono
- Institute for Quantum Life Science, National Institutes for Quantum Science and Technology, Chiba, Japan
| | - Kazuma Nakatani
- Laboratory of Evolutionary Oncology, Chiba Cancer Center Research Institute, Chiba, Japan
- Graduate School of Medical and Pharmaceutical Sciences, Chiba University, Chiba, Japan
- Innovative Medicine CHIBA Doctoral WISE Program, Chiba University, Chiba, Japan
- All Directional Innovation Creator Ph.D. Project, Chiba University, Chiba, Japan
| | - Miwako Hirose
- Digital Healthcare Business Development Office, NEC Solution Innovators, Ltd., Tokyo, Japan
| | - Katsunori Horii
- Digital Healthcare Business Development Office, NEC Solution Innovators, Ltd., Tokyo, Japan
| | - Yoshitaka Hippo
- Laboratory of Evolutionary Oncology, Chiba Cancer Center Research Institute, Chiba, Japan
- Graduate School of Medical and Pharmaceutical Sciences, Chiba University, Chiba, Japan
- Laboratory of Precision Tumor Model Systems, Chiba Cancer Center Research Institute, Chiba, Japan
| | - Taro Tamada
- Institute for Quantum Life Science, National Institutes for Quantum Science and Technology, Chiba, Japan
- Graduate School of Science, Chiba University, Chiba, Japan
| | - Yusuke Suenaga
- Laboratory of Evolutionary Oncology, Chiba Cancer Center Research Institute, Chiba, Japan
| | - Tatsuhito Matsuo
- Institute for Quantum Life Science, National Institutes for Quantum Science and Technology, Chiba, Japan
| |
Collapse
|
21
|
Weisman CM. The permissive binding theory of cancer. Front Oncol 2023; 13:1272981. [PMID: 38023252 PMCID: PMC10666763 DOI: 10.3389/fonc.2023.1272981] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/04/2023] [Accepted: 10/20/2023] [Indexed: 12/01/2023] Open
Abstract
The later stages of cancer, including the invasion and colonization of new tissues, are actively mysterious compared to earlier stages like primary tumor formation. While we lack many details about both, we do have an apparently successful explanatory framework for the earlier stages: one in which genetic mutations hold ultimate causal and explanatory power. By contrast, on both empirical and conceptual grounds, it is not currently clear that mutations alone can explain the later stages of cancer. Can a different type of molecular change do better? Here, I introduce the "permissive binding theory" of cancer, which proposes that novel protein binding interactions are the key causal and explanatory entity in invasion and metastasis. It posits that binding is more abundant at baseline than we observe because it is restricted in normal physiology; that any large perturbation to physiological state revives this baseline abundance, unleashing many new binding interactions; and that a subset of these cause the cellular functions at the heart of oncogenesis, especially invasion and metastasis. Significant physiological perturbations occur in cancer cells in very early stages, and generally become more extreme with progression, providing interactions that continually fuel invasion and metastasis. The theory is compatible with, but not limited to, causal roles for the diverse molecular changes observed in cancer (e.g. gene expression or epigenetic changes), as these generally act causally upstream of proteins, and so may exert their effects by changing the protein binding interactions that occur in the cell. This admits the possibility that molecular changes that appear quite different may actually converge in creating the same few protein complexes, simplifying our picture of invasion and metastasis. If correct, the theory offers a concrete therapeutic strategy: targeting the key novel complexes. The theory is straightforwardly testable by large-scale identification of protein interactions in different cancers.
Collapse
Affiliation(s)
- Caroline M. Weisman
- Lewis-Sigler Institute for Integrative Genomics, Princeton University, Princeton, NJ, United States
| |
Collapse
|
22
|
Wang Z, Wang YW, Kasuga T, Lopez-Giraldez F, Zhang Y, Zhang Z, Wang Y, Dong C, Sil A, Trail F, Yarden O, Townsend JP. Lineage-specific genes are clustered with HET-domain genes and respond to environmental and genetic manipulations regulating reproduction in Neurospora. PLoS Genet 2023; 19:e1011019. [PMID: 37934795 PMCID: PMC10684091 DOI: 10.1371/journal.pgen.1011019] [Citation(s) in RCA: 1] [Impact Index Per Article: 0.5] [Reference Citation Analysis] [Abstract] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/23/2023] [Revised: 11/28/2023] [Accepted: 10/16/2023] [Indexed: 11/09/2023] Open
Abstract
Lineage-specific genes (LSGs) have long been postulated to play roles in the establishment of genetic barriers to intercrossing and speciation. In the genome of Neurospora crassa, most of the 670 Neurospora LSGs that are aggregated adjacent to the telomeres are clustered with 61% of the HET-domain genes, some of which regulate self-recognition and define vegetative incompatibility groups. In contrast, the LSG-encoding proteins possess few to no domains that would help to identify potential functional roles. Possible functional roles of LSGs were further assessed by performing transcriptomic profiling in genetic mutants and in response to environmental alterations, as well as examining gene knockouts for phenotypes. Among the 342 LSGs that are dynamically expressed during both asexual and sexual phases, 64% were detectable on unusual carbon sources such as furfural, a wildfire-produced chemical that is a strong inducer of sexual development, and the structurally-related furan 5-hydroxymethyl furfural (HMF). Expression of a significant portion of the LSGs was sensitive to light and temperature, factors that also regulate the switch from asexual to sexual reproduction. Furthermore, expression of the LSGs was significantly affected in the knockouts of adv-1 and pp-1 that regulate hyphal communication, and expression of more than one quarter of the LSGs was affected by perturbation of the mating locus. These observations encouraged further investigation of the roles of clustered lineage-specific and HET-domain genes in ecology and reproduction regulation in Neurospora, especially the regulation of the switch from the asexual growth to sexual reproduction, in response to dramatic environmental conditions changes.
Collapse
Affiliation(s)
- Zheng Wang
- Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut, United States of America
| | - Yen-Wen Wang
- Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut, United States of America
| | - Takao Kasuga
- College of Biological Sciences, University of California, Davis, California, United States of America
| | | | - Yang Zhang
- National Genomics Data Center, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China
| | - Zhang Zhang
- National Genomics Data Center, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China
| | - Yaning Wang
- Institute of Microbiology, Chinese Academy of Sciences, Beijing, China
| | - Caihong Dong
- Institute of Microbiology, Chinese Academy of Sciences, Beijing, China
| | - Anita Sil
- Department of Microbiology and Immunology, University of California, San Francisco, California, United States of America
| | - Frances Trail
- Department of Plant, Soil and Microbial Sciences, Michigan State University, East Lansing, Michigan, United States of America
| | - Oded Yarden
- Department of Plant Pathology and Microbiology, The Robert H. Smith Faculty of Agriculture, Food and Environment, The Hebrew University of Jerusalem, Rehovot, Israel
| | - Jeffrey P. Townsend
- Department of Biostatistics, Yale School of Public Health, New Haven, Connecticut, United States of America
- Department of Ecology and Evolutionary Biology, Program in Microbiology, and Program in Computational Biology and Bioinformatics, Yale University, New Haven, Connecticut, United States of America
| |
Collapse
|
23
|
Ardern Z. Alternative Reading Frames are an Underappreciated Source of Protein Sequence Novelty. J Mol Evol 2023; 91:570-580. [PMID: 37326679 DOI: 10.1007/s00239-023-10122-3] [Citation(s) in RCA: 8] [Impact Index Per Article: 4.0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/19/2022] [Accepted: 05/31/2023] [Indexed: 06/17/2023]
Abstract
Protein-coding DNA sequences can be translated into completely different amino acid sequences if the nucleotide triplets used are shifted by a non-triplet amount on the same DNA strand or by translating codons from the opposite strand. Such "alternative reading frames" of protein-coding genes are a major contributor to the evolution of novel protein products. Recent studies demonstrating this include examples across the three domains of cellular life and in viruses. These sequences increase the number of trials potentially available for the evolutionary invention of new genes and also have unusual properties which may facilitate gene origin. There is evidence that the structure of the standard genetic code contributes to the features and gene-likeness of some alternative frame sequences. These findings have important implications across diverse areas of molecular biology, including for genome annotation, structural biology, and evolutionary genomics.
Collapse
|
24
|
Alvarez S, Nartey CM, Mercado N, de la Paz A, Huseinbegovic T, Morcos F. In vivo functional phenotypes from a computational epistatic model of evolution. BIORXIV : THE PREPRINT SERVER FOR BIOLOGY 2023:2023.05.24.542176. [PMID: 37292895 PMCID: PMC10245989 DOI: 10.1101/2023.05.24.542176] [Citation(s) in RCA: 1] [Impact Index Per Article: 0.5] [Reference Citation Analysis] [Abstract] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 06/10/2023]
Abstract
Computational models of evolution are valuable for understanding the dynamics of sequence variation, to infer phylogenetic relationships or potential evolutionary pathways and for biomedical and industrial applications. Despite these benefits, few have validated their propensities to generate outputs with in vivo functionality, which would enhance their value as accurate and interpretable evolutionary algorithms. We demonstrate the power of epistasis inferred from natural protein families to evolve sequence variants in an algorithm we developed called Sequence Evolution with Epistatic Contributions. Utilizing the Hamiltonian of the joint probability of sequences in the family as fitness metric, we sampled and experimentally tested for in vivo β -lactamase activity in E. coli TEM-1 variants. These evolved proteins can have dozens of mutations dispersed across the structure while preserving sites essential for both catalysis and interactions. Remarkably, these variants retain family-like functionality while being more active than their WT predecessor. We found that depending on the inference method used to generate the epistatic constraints, different parameters simulate diverse selection strengths. Under weaker selection, local Hamiltonian fluctuations reliably predict relative changes to variant fitness, recapitulating neutral evolution. SEEC has the potential to explore the dynamics of neofunctionalization, characterize viral fitness landscapes and facilitate vaccine development.
Collapse
Affiliation(s)
- Sophia Alvarez
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX 75080, USA
| | - Charisse M. Nartey
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX 75080, USA
| | - Nicholas Mercado
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX 75080, USA
| | - Alberto de la Paz
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX 75080, USA
| | - Tea Huseinbegovic
- School of Natural Sciences and Mathematics, University of Texas at Dallas, Richardson, TX 75080, USA
| | - Faruck Morcos
- Department of Biological Sciences, University of Texas at Dallas, Richardson, TX 75080, USA
- Department of Bioengineering, University of Texas at Dallas, Richardson, TX 75080, USA
- Center for Systems Biology, University of Texas at Dallas, Richardson, TX 75080, USA
| |
Collapse
|
25
|
Ardern Z, Uz-Zaman MH. Between noise and function: Toward a taxonomy of the non-canonical translatome. Cell Syst 2023; 14:343-345. [PMID: 37201506 DOI: 10.1016/j.cels.2023.04.004] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 04/17/2023] [Accepted: 04/17/2023] [Indexed: 05/20/2023]
Abstract
Eukaryotic genomes are pervasively translated, but the properties of translated sequences outside of canonical genes are poorly understood. A new study in Cell Systems reveals a large translatome that is not under significant evolutionary constraint but is still an active part of diverse cellular systems.
Collapse
Affiliation(s)
- Zachary Ardern
- Parasites and Microbes Programme, Wellcome Sanger Institute, Hinxton, Cambridgeshire, UK.
| | - Md Hassan Uz-Zaman
- Department of Molecular Biosciences, University of Texas at Austin, Austin, TX, USA.
| |
Collapse
|
26
|
Tornini VA, Miao L, Lee HJ, Gerson T, Dube SE, Schmidt V, Kroll F, Tang Y, Du K, Kuchroo M, Vejnar CE, Bazzini AA, Krishnaswamy S, Rihel J, Giraldez AJ. linc-mipep and linc-wrb encode micropeptides that regulate chromatin accessibility in vertebrate-specific neural cells. eLife 2023; 12:e82249. [PMID: 37191016 PMCID: PMC10188112 DOI: 10.7554/elife.82249] [Citation(s) in RCA: 2] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/28/2022] [Accepted: 04/14/2023] [Indexed: 05/17/2023] Open
Abstract
Thousands of long intergenic non-coding RNAs (lincRNAs) are transcribed throughout the vertebrate genome. A subset of lincRNAs enriched in developing brains have recently been found to contain cryptic open-reading frames and are speculated to encode micropeptides. However, systematic identification and functional assessment of these transcripts have been hindered by technical challenges caused by their small size. Here, we show that two putative lincRNAs (linc-mipep, also called lnc-rps25, and linc-wrb) encode micropeptides with homology to the vertebrate-specific chromatin architectural protein, Hmgn1, and demonstrate that they are required for development of vertebrate-specific brain cell types. Specifically, we show that NMDA receptor-mediated pathways are dysregulated in zebrafish lacking these micropeptides and that their loss preferentially alters the gene regulatory networks that establish cerebellar cells and oligodendrocytes - evolutionarily newer cell types that develop postnatally in humans. These findings reveal a key missing link in the evolution of vertebrate brain cell development and illustrate a genetic basis for how some neural cell types are more susceptible to chromatin disruptions, with implications for neurodevelopmental disorders and disease.
Collapse
Affiliation(s)
| | - Liyun Miao
- Department of Genetics, Yale UniversityNew HavenUnited States
| | - Ho-Joon Lee
- Department of Genetics, Yale UniversityNew HavenUnited States
- Yale Center for Genome Analysis, Yale UniversityNew HavenUnited States
| | - Timothy Gerson
- Department of Genetics, Yale UniversityNew HavenUnited States
| | - Sarah E Dube
- Department of Genetics, Yale UniversityNew HavenUnited States
| | - Valeria Schmidt
- Department of Genetics, Yale UniversityNew HavenUnited States
| | - François Kroll
- Department of Cell and Developmental Biology, University College LondonLondonUnited Kingdom
| | - Yin Tang
- Department of Genetics, Yale UniversityNew HavenUnited States
| | - Katherine Du
- Department of Genetics, Yale UniversityNew HavenUnited States
- Department of Computer Science, Yale UniversityNew HavenUnited States
| | - Manik Kuchroo
- Department of Genetics, Yale UniversityNew HavenUnited States
- Department of Computer Science, Yale UniversityNew HavenUnited States
| | | | - Ariel Alejandro Bazzini
- Stowers Institute for Medical ResearchKansas CityUnited States
- Department of Molecular & Integrative Physiology, University of Kansas School of MedicineKansas CityUnited States
| | - Smita Krishnaswamy
- Department of Genetics, Yale UniversityNew HavenUnited States
- Department of Computer Science, Yale UniversityNew HavenUnited States
| | - Jason Rihel
- Department of Cell and Developmental Biology, University College LondonLondonUnited Kingdom
| | - Antonio J Giraldez
- Department of Genetics, Yale UniversityNew HavenUnited States
- Yale Stem Cell Center, Yale University School of MedicineNew HavenUnited States
- Yale Cancer Center, Yale University School of MedicineNew HavenUnited States
| |
Collapse
|
27
|
Kozlov AP. Carcino-Evo-Devo, A Theory of the Evolutionary Role of Hereditary Tumors. Int J Mol Sci 2023; 24:ijms24108611. [PMID: 37239953 DOI: 10.3390/ijms24108611] [Citation(s) in RCA: 2] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/16/2023] [Revised: 05/08/2023] [Accepted: 05/09/2023] [Indexed: 05/28/2023] Open
Abstract
A theory of the evolutionary role of hereditary tumors, or the carcino-evo-devo theory, is being developed. The main hypothesis of the theory, the hypothesis of evolution by tumor neofunctionalization, posits that hereditary tumors provided additional cell masses during the evolution of multicellular organisms for the expression of evolutionarily novel genes. The carcino-evo-devo theory has formulated several nontrivial predictions that have been confirmed in the laboratory of the author. It also suggests several nontrivial explanations of biological phenomena previously unexplained by the existing theories or incompletely understood. By considering three major types of biological development-individual, evolutionary, and neoplastic development-within one theoretical framework, the carcino-evo-devo theory has the potential to become a unifying biological theory.
Collapse
Affiliation(s)
- Andrei P Kozlov
- Vavilov Institute of General Genetics, Russian Academy of Sciences, 3 Gubkina Street, 117971 Moscow, Russia
- Peter the Great St. Petersburg Polytechnic University, 29 Polytekhnicheskaya Street, 195251 St. Petersburg, Russia
| |
Collapse
|
28
|
Aubel M, Eicholt L, Bornberg-Bauer E. Assessing structure and disorder prediction tools for de novo emerged proteins in the age of machine learning. F1000Res 2023; 12:347. [PMID: 37113259 PMCID: PMC10126731 DOI: 10.12688/f1000research.130443.1] [Citation(s) in RCA: 11] [Impact Index Per Article: 5.5] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Submit a Manuscript] [Subscribe] [Scholar Register] [Accepted: 03/17/2023] [Indexed: 03/31/2023] Open
Abstract
Background: De novo protein coding genes emerge from scratch in the non-coding regions of the genome and have, per definition, no homology to other genes. Therefore, their encoded de novo proteins belong to the so-called "dark protein space". So far, only four de novo protein structures have been experimentally approximated. Low homology, presumed high disorder and limited structures result in low confidence structural predictions for de novo proteins in most cases. Here, we look at the most widely used structure and disorder predictors and assess their applicability for de novo emerged proteins. Since AlphaFold2 is based on the generation of multiple sequence alignments and was trained on solved structures of largely conserved and globular proteins, its performance on de novo proteins remains unknown. More recently, natural language models of proteins have been used for alignment-free structure predictions, potentially making them more suitable for de novo proteins than AlphaFold2. Methods: We applied different disorder predictors (IUPred3 short/long, flDPnn) and structure predictors, AlphaFold2 on the one hand and language-based models (Omegafold, ESMfold, RGN2) on the other hand, to four de novo proteins with experimental evidence on structure. We compared the resulting predictions between the different predictors as well as to the existing experimental evidence. Results: Results from IUPred, the most widely used disorder predictor, depend heavily on the choice of parameters and differ significantly from flDPnn which has been found to outperform most other predictors in a comparative assessment study recently. Similarly, different structure predictors yielded varying results and confidence scores for de novo proteins. Conclusions: We suggest that, while in some cases protein language model based approaches might be more accurate than AlphaFold2, the structure prediction of de novo emerged proteins remains a difficult task for any predictor, be it disorder or structure.
Collapse
Affiliation(s)
- Margaux Aubel
- Institute for Evolution and Bidiversity, University of Muenster, Muenster, 48149, Germany
| | - Lars Eicholt
- Institute for Evolution and Bidiversity, University of Muenster, Muenster, 48149, Germany
| | - Erich Bornberg-Bauer
- Institute for Evolution and Bidiversity, University of Muenster, Muenster, 48149, Germany
- Department Protein Evolution, Max Planck-Institute for Biology, Tuebingen, 72076, Germany
| |
Collapse
|
29
|
Evolution and implications of de novo genes in humans. Nat Ecol Evol 2023:10.1038/s41559-023-02014-y. [PMID: 36928843 DOI: 10.1038/s41559-023-02014-y] [Citation(s) in RCA: 33] [Impact Index Per Article: 16.5] [Reference Citation Analysis] [Abstract] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/24/2022] [Accepted: 02/06/2023] [Indexed: 03/18/2023]
Abstract
Genes and translated open reading frames (ORFs) that emerged de novo from previously non-coding sequences provide species with opportunities for adaptation. When aberrantly activated, some human-specific de novo genes and ORFs have disease-promoting properties-for instance, driving tumour growth. Thousands of putative de novo coding sequences have been described in humans, but we still do not know what fraction of those ORFs has readily acquired a function. Here, we discuss the challenges and controversies surrounding the detection, mechanisms of origin, annotation, validation and characterization of de novo genes and ORFs. Through manual curation of literature and databases, we provide a thorough table with most de novo genes reported for humans to date. We re-evaluate each locus by tracing the enabling mutations and list proposed disease associations, protein characteristics and supporting evidence for translation and protein detection. This work will support future explorations of de novo genes and ORFs in humans.
Collapse
|
30
|
The Theory of Carcino-Evo-Devo and Its Non-Trivial Predictions. Genes (Basel) 2022; 13:genes13122347. [PMID: 36553613 PMCID: PMC9777766 DOI: 10.3390/genes13122347] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/08/2022] [Revised: 12/04/2022] [Accepted: 12/08/2022] [Indexed: 12/15/2022] Open
Abstract
To explain the sources of additional cell masses in the evolution of multicellular organisms, the theory of carcino-evo-devo, or evolution by tumor neofunctionalization, has been developed. The important demand for a new theory in experimental science is the capability to formulate non-trivial predictions which can be experimentally confirmed. Several non-trivial predictions were formulated using carcino-evo-devo theory, four of which are discussed in the present paper: (1) The number of cellular oncogenes should correspond to the number of cell types in the organism. The evolution of oncogenes, tumor suppressor and differentiation gene classes should proceed concurrently. (2) Evolutionarily new and evolving genes should be specifically expressed in tumors (TSEEN genes). (3) Human orthologs of fish TSEEN genes should acquire progressive functions connected with new cell types, tissues and organs. (4) Selection of tumors for new functions in the organism is possible. Evolutionarily novel organs should recapitulate tumor features in their development. As shown in this paper, these predictions have been confirmed by the laboratory of the author. Thus, we have shown that carcino-evo-devo theory has predictive power, fulfilling a fundamental requirement for a new theory.
Collapse
|
31
|
Parikh SB, Houghton C, Van Oss SB, Wacholder A, Carvunis A. Origins, evolution, and physiological implications of de novo genes in yeast. Yeast 2022; 39:471-481. [PMID: 35959631 PMCID: PMC9544372 DOI: 10.1002/yea.3810] [Citation(s) in RCA: 4] [Impact Index Per Article: 1.3] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/20/2022] [Revised: 08/08/2022] [Accepted: 08/09/2022] [Indexed: 12/03/2022] Open
Abstract
De novo gene birth is the process by which new genes emerge in sequences that were previously noncoding. Over the past decade, researchers have taken advantage of the power of yeast as a model and a tool to study the evolutionary mechanisms and physiological implications of de novo gene birth. We summarize the mechanisms that have been proposed to explicate how noncoding sequences can become protein-coding genes, highlighting the discovery of pervasive translation of the yeast transcriptome and its presumed impact on evolutionary innovation. We summarize current best practices for the identification and characterization of de novo genes. Crucially, we explain that the field is still in its nascency, with the physiological roles of most young yeast de novo genes identified thus far still utterly unknown. We hope this review inspires researchers to investigate the true contribution of de novo gene birth to cellular physiology and phenotypic diversity across yeast strains and species.
Collapse
Affiliation(s)
- Saurin B. Parikh
- Department of Computational and Systems Biology, School of Medicine, Pittsburgh Center for Evolutionary Biology and EvolutionUniversity of PittsburghPittsburghPennsylvaniaUSA
| | - Carly Houghton
- Department of Computational and Systems Biology, School of Medicine, Pittsburgh Center for Evolutionary Biology and EvolutionUniversity of PittsburghPittsburghPennsylvaniaUSA
| | - S. Branden Van Oss
- Department of Computational and Systems Biology, School of Medicine, Pittsburgh Center for Evolutionary Biology and EvolutionUniversity of PittsburghPittsburghPennsylvaniaUSA
| | - Aaron Wacholder
- Department of Computational and Systems Biology, School of Medicine, Pittsburgh Center for Evolutionary Biology and EvolutionUniversity of PittsburghPittsburghPennsylvaniaUSA
| | - Anne‐Ruxandra Carvunis
- Department of Computational and Systems Biology, School of Medicine, Pittsburgh Center for Evolutionary Biology and EvolutionUniversity of PittsburghPittsburghPennsylvaniaUSA
| |
Collapse
|
32
|
Pajic P, Shen S, Qu J, May AJ, Knox S, Ruhl S, Gokcumen O. A mechanism of gene evolution generating mucin function. SCIENCE ADVANCES 2022; 8:eabm8757. [PMID: 36026444 PMCID: PMC9417175 DOI: 10.1126/sciadv.abm8757] [Citation(s) in RCA: 15] [Impact Index Per Article: 5.0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Received: 10/18/2021] [Accepted: 07/12/2022] [Indexed: 05/12/2023]
Abstract
How novel gene functions evolve is a fundamental question in biology. Mucin proteins, a functionally but not evolutionarily defined group of proteins, allow the study of convergent evolution of gene function. By analyzing the genomic variation of mucins across a wide range of mammalian genomes, we propose that exonic repeats and their copy number variation contribute substantially to the de novo evolution of new gene functions. By integrating bioinformatic, phylogenetic, proteomic, and immunohistochemical approaches, we identified 15 undescribed instances of evolutionary convergence, where novel mucins originated by gaining densely O-glycosylated exonic repeat domains. Our results suggest that secreted proteins rich in proline are natural precursors for acquiring mucin function. Our findings have broad implications for understanding the role of exonic repeats in the parallel evolution of new gene functions, especially those involving protein glycosylation.
Collapse
Affiliation(s)
- Petar Pajic
- Department of Biological Sciences, University at Buffalo, The State University of New York, Buffalo, NY 14260, USA
- Department of Oral Biology, School of Dental Medicine, University at Buffalo, The State University of New York, Buffalo, NY 14214, USA
| | - Shichen Shen
- Department of Pharmaceutical Sciences, University at Buffalo, The State University of New York, Buffalo, NY 14214, USA
- Center of Excellence in Bioinformatics and Life Science, Buffalo, NY 14203, USA
| | - Jun Qu
- Department of Pharmaceutical Sciences, University at Buffalo, The State University of New York, Buffalo, NY 14214, USA
- Center of Excellence in Bioinformatics and Life Science, Buffalo, NY 14203, USA
| | - Alison J. May
- Program in Craniofacial Biology, Department of Cell and Tissue Biology, School of Dentistry, University of California, San Francisco, San Francisco, CA 94143, USA
| | - Sarah Knox
- Program in Craniofacial Biology, Department of Cell and Tissue Biology, School of Dentistry, University of California, San Francisco, San Francisco, CA 94143, USA
| | - Stefan Ruhl
- Department of Oral Biology, School of Dental Medicine, University at Buffalo, The State University of New York, Buffalo, NY 14214, USA
| | - Omer Gokcumen
- Department of Biological Sciences, University at Buffalo, The State University of New York, Buffalo, NY 14260, USA
| |
Collapse
|
33
|
Foley S, Vlasova A, Marcet-Houben M, Gabaldón T, Hinman VF. Evolutionary analyses of genes in Echinodermata offer insights towards the origin of metazoan phyla. Genomics 2022; 114:110431. [PMID: 35835427 PMCID: PMC9552553 DOI: 10.1016/j.ygeno.2022.110431] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/03/2021] [Revised: 05/10/2022] [Accepted: 07/06/2022] [Indexed: 11/24/2022]
Abstract
Despite recent studies discussing the evolutionary impacts of gene duplications and losses among metazoans, the genomic basis for the evolution of phyla remains enigmatic. Here, we employ phylogenomic approaches to search for orthologous genes without known functions among echinoderms, and subsequently use them to guide the identification of their homologs across other metazoans. Our final set of 14 genes was obtained via a suite of homology prediction tools, gene expression data, gene ontology, and generating the Strongylocentrotus purpuratus phylome. The gene set was subjected to selection pressure analyses, which indicated that they are highly conserved and under negative selection. Their presence across broad taxonomic depths suggests that genes required to form a phylum are ancestral to that phylum. Therefore, rather than de novo gene genesis, we posit that evolutionary forces such as selection on existing genomic elements over large timescales may drive divergence and contribute to the emergence of phyla.
Collapse
Affiliation(s)
- Saoirse Foley
- Department of Biological Sciences, Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213, USA; Echinobase #6-46, Mellon Institute, 4400 Fifth Ave, Pittsburgh, PA 15213, USA.
| | - Anna Vlasova
- Barcelona Supercomputing Centre (BSC-CNS), Jordi Girona, 29, 08034 Barcelona, Spain; Institute for Research in Biomedicine (IRB Barcelona), The Barcelona Institute of Science and Technology, Baldiri Reixac, 10, 08028 Barcelona, Spain
| | - Marina Marcet-Houben
- Barcelona Supercomputing Centre (BSC-CNS), Jordi Girona, 29, 08034 Barcelona, Spain; Institute for Research in Biomedicine (IRB Barcelona), The Barcelona Institute of Science and Technology, Baldiri Reixac, 10, 08028 Barcelona, Spain
| | - Toni Gabaldón
- Barcelona Supercomputing Centre (BSC-CNS), Jordi Girona, 29, 08034 Barcelona, Spain; Institute for Research in Biomedicine (IRB Barcelona), The Barcelona Institute of Science and Technology, Baldiri Reixac, 10, 08028 Barcelona, Spain; Catalan Institution for Research and Advanced Studies (ICREA), Barcelona, Spain
| | - Veronica F Hinman
- Department of Biological Sciences, Carnegie Mellon University, 5000 Forbes Ave, Pittsburgh, PA 15213, USA; Echinobase #6-46, Mellon Institute, 4400 Fifth Ave, Pittsburgh, PA 15213, USA
| |
Collapse
|