Compositional biases of bacterial genomes and evolutionary implications
- PMID: 9190805
- PMCID: PMC179198
- DOI: 10.1128/jb.179.12.3899-3913.1997
Compositional biases of bacterial genomes and evolutionary implications
Abstract
We compare and contrast genome-wide compositional biases and distributions of short oligonucleotides across 15 diverse prokaryotes that have substantial genomic sequence collections. These include seven complete genomes (Escherichia coli, Haemophilus influenzae, Mycoplasma genitalium, Mycoplasma pneumoniae, Synechocystis sp. strain PCC6803, Methanococcus jannaschii, and Pyrobaculum aerophilum). A key observation concerns the constancy of the dinucleotide relative abundance profiles over multiple 50-kb disjoint contigs within the same genome. (The profile is rhoXY* = fXY*/fX*fY* for all XY, where fX* denotes the frequency of the nucleotide X and fY* denotes the frequency of the dinucleotide XY, both computed from the sequence concatenated with its inverted complementary sequence.) On the basis of this constancy, we refer to the collection [rhoXY*] as the genome signature. We establish that the differences between [rhoXY*] vectors of 50-kb sample contigs of different genomes virtually always exceed the differences between those of the same genomes. Various di- and tetranucleotide biases are identified. In particular, we find that the dinucleotide CpG=CG is underrepresented in many thermophiles (e.g., M. jannaschii, Sulfolobus sp., and M. thermoautotrophicum) but overrepresented in halobacteria. TA is broadly underrepresented in prokaryotes and eukaryotes, but normal counts appear in Sulfolobus and P. aerophilum sequences. More than for any other bacterial genome, palindromic tetranucleotides are underrepresented in H. influenzae. The M. jannaschii sequence is unprecedented in its extreme underrepresentation of CTAG tetranucleotides and in the anomalous distribution of CTAG sites around the genome. Comparative analysis of numbers of long tetranucleotide microsatellites distinguishes H. influenzae. Dinucleotide relative abundance differences between bacterial sequences are compared. For example, in these assessments of differences, the cyanobacteria Synechocystis, Synechococcus, and Anabaena do not form a coherent group and are as far from each other as general gram-negative sequences are from general gram-positive sequences. The difference of M. jannaschii from low-G+C gram-positive proteobacteria is one-half of the difference from gram-negative proteobacteria. Interpretations and hypotheses center on the role of the genome signature in highlighting similarities and dissimilarities across different classes of prokaryotic species, possible mechanisms underlying the genome signature, the form and level of genome compositional flux, the use of the genome signature as a chronometer of molecular phylogeny, and implications with respect to the three putative eubacterial, archaeal, and eukaryote domains of life and to the origin and early evolution of eukaryotes.
Similar articles
-
Microbial genome analyses: global comparisons of transport capabilities based on phylogenies, bioenergetics and substrate specificities.J Mol Biol. 1998 Apr 3;277(3):573-92. doi: 10.1006/jmbi.1998.1609. J Mol Biol. 1998. PMID: 9533881
-
Comparison of archaeal and bacterial genomes: computer analysis of protein sequences predicts novel functions and suggests a chimeric origin for the archaea.Mol Microbiol. 1997 Aug;25(4):619-37. doi: 10.1046/j.1365-2958.1997.4821861.x. Mol Microbiol. 1997. PMID: 9379893
-
Frequent oligonucleotides and peptides of the Haemophilus influenzae genome.Nucleic Acids Res. 1996 Nov 1;24(21):4263-72. doi: 10.1093/nar/24.21.4263. Nucleic Acids Res. 1996. PMID: 8932382 Free PMC article.
-
Comparative DNA analysis across diverse genomes.Annu Rev Genet. 1998;32:185-225. doi: 10.1146/annurev.genet.32.1.185. Annu Rev Genet. 1998. PMID: 9928479 Review.
-
Global dinucleotide signatures and analysis of genomic heterogeneity.Curr Opin Microbiol. 1998 Oct;1(5):598-610. doi: 10.1016/s1369-5274(98)80095-7. Curr Opin Microbiol. 1998. PMID: 10066522 Review.
Cited by
-
Depletion of CpG dinucleotides in bacterial genomes may represent an adaptation to high temperatures.NAR Genom Bioinform. 2024 Jul 27;6(3):lqae088. doi: 10.1093/nargab/lqae088. eCollection 2024 Sep. NAR Genom Bioinform. 2024. PMID: 39071851 Free PMC article.
-
Synsor: a tool for alignment-free detection of engineered DNA sequences.Front Bioeng Biotechnol. 2024 Jul 12;12:1375626. doi: 10.3389/fbioe.2024.1375626. eCollection 2024. Front Bioeng Biotechnol. 2024. PMID: 39070163 Free PMC article.
-
Environment and taxonomy shape the genomic signature of prokaryotic extremophiles.Sci Rep. 2023 Sep 26;13(1):16105. doi: 10.1038/s41598-023-42518-y. Sci Rep. 2023. PMID: 37752120 Free PMC article.
-
Viral community composition of hypersaline lakes.Virus Evol. 2023 Aug 30;9(2):vead057. doi: 10.1093/ve/vead057. eCollection 2023. Virus Evol. 2023. PMID: 37692898 Free PMC article.
-
Genomic Signature in Evolutionary Biology: A Review.Biology (Basel). 2023 Feb 16;12(2):322. doi: 10.3390/biology12020322. Biology (Basel). 2023. PMID: 36829597 Free PMC article. Review.
References
Publication types
MeSH terms
Substances
Grants and funding
LinkOut - more resources
Full Text Sources
Other Literature Sources