Home LiteratureArticle Details
PMID: 17355171 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

The Sorcerer II Global Ocean Sampling expedition: expanding the universe of protein families.

PLoS biology ·Vol. 5 ·No. 3 ·2007-03-00 ·Pages e16

Yooseph S, Sutton G, Rusch DB, Halpern AL, Williamson SJ, Remington K, Eisen JA, Heidelberg KB, Manning G, Li W, Jaroszewski L, Cieplak P, Miller CS, Li H, Mashiyama ST, Joachimiak MP, van Belle C, Chandonia JM, Soergel DA, Zhai Y, Natarajan K, Lee S, Raphael BJ, Bafna V, Friedman R, Brenner SE, Godzik A, Eisenberg D, Dixon JE, Taylor SS, Strausberg RL, Frazier M, Venter JC

Abstract

Metagenomics projects based on shotgun sequencing of populations of micro-organisms yield insight into protein families. We used sequence similarity clustering to explore proteins with a comprehensive dataset consisting of sequences from available databases together with 6.12 million proteins predicted from an assembly of 7.7 million Global Ocean Sampling (GOS) sequences. The GOS dataset covers nearly all known prokaryotic protein families. A total of 3,995 medium- and large-sized clusters consisting of only GOS sequences are identified, out of which 1,700 have no detectable homology to known families. The GOS-only clusters contain a higher than expected proportion of sequences of viral origin, thus reflecting a poor sampling of viral diversity until now. Protein domain distributions in the GOS dataset and current protein databases show distinct biases. Several protein domains that were previously categorized as kingdom specific are shown to have GOS examples in other kingdoms. About 6,000 sequences (ORFans) from the literature that heretofore lacked similarity to known proteins have matches in the GOS data. The GOS dataset is also used to improve remote homology detection. Overall, besides nearly doubling the number of current proteins, the predicted GOS proteins also add a great deal of diversity to known protein families and shed light on their evolution. These observations are illustrated using several protein families, including phosphatases, proteases, ultraviolet-irradiation DNA damage repair enzymes, glutamine synthetase, and RuBisCO. The diversity added by GOS data has implications for choosing targets for experimental structure characterization as part of structural genomics efforts. Our analysis indicates that new families are being discovered at a rate that is linear or almost linear with the addition of new sequences, implying that we are still far from discovering all protein families in nature.

MeSH Terms
Expressed Sequence Tags Oceans and Seas Proteins/chemistry,genetics Water Microbiology
Chemicals
Proteins
Authors & Affiliations
33 authors, click to expand affiliations / ORCID
Yooseph Shibu
J. Craig Venter Institute, Rockville, Maryland, United States of America. Shibu.Yooseph@venterinstitute.org
Sutton Granger
Rusch Douglas B
Halpern Aaron L
Williamson Shannon J
Remington Karin
Eisen Jonathan A
Heidelberg Karla B
Manning Gerard
Li Weizhong
Jaroszewski Lukasz
Cieplak Piotr
Miller Christopher S
Li Huiying
Mashiyama Susan T
Joachimiak Marcin P
van Belle Christopher
Chandonia John-Marc
Soergel David A
Zhai Yufeng
Natarajan Kannan
Lee Shaun
Raphael Benjamin J
Bafna Vineet
Friedman Robert
Brenner Steven E
Godzik Adam
Eisenberg David
Dixon Jack E
Taylor Susan S
Strausberg Robert L
Frazier Marvin
Venter J Craig
References (141)
141 references, click to expand
  1. Orphans as taxonomically restricted and ecologically important genes.
    Microbiology. 2005 Aug;151(Pt 8):2499-501 PMID: 16079329
  2. Prolinks: a database of protein functional linkages derived from coevolution.
    Genome Biol. 2004;5(5):R35 PMID: 15128449
  3. Myriads of protein families, and still counting.
    Genome Biol. 2003;4(2):401 PMID: 12620116
  4. Ensembl 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D468-70 PMID: 14681459
  5. Codon-substitution models for heterogeneous selection pressure at amino acid sites.
    Genetics. 2000 May;155(1):431-49 PMID: 10790415
  6. Distinguishing the ORFs from the ELFs: short bacterial genes and the annotation of genomes.
    Trends Genet. 2002 Jul;18(7):335-7 PMID: 12127765
  7. Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes.
    J Mol Biol. 2001 Jan 19;305(3):567-80 PMID: 11152613
  8. Protein structure prediction and structural genomics.
    Science. 2001 Oct 5;294(5540):93-6 PMID: 11588250
  9. Structural biology. Structural genomics, round 2.
    Science. 2005 Mar 11;307(5715):1554-8 PMID: 15761136
  10. The frequency distribution of gene family sizes in complete genomes.
    Mol Biol Evol. 1998 May;15(5):583-9 PMID: 9580988
  11. QuickTree: building huge Neighbour-Joining trees of protein sequences.
    Bioinformatics. 2002 Nov;18(11):1546-7 PMID: 12424131
  12. Weighted neighbor joining: a likelihood-based approach to distance-based phylogeny reconstruction.
    Mol Biol Evol. 2000 Jan;17(1):189-97 PMID: 10666718
  13. IDO expression by dendritic cells: tolerance and tryptophan catabolism.
    Nat Rev Immunol. 2004 Oct;4(10):762-74 PMID: 15459668
  14. Emergence of scaling in random networks
    Science. 1999 Oct 15;286(5439):509-12 PMID: 10521342
  15. Protein folds, functions and evolution.
    J Mol Biol. 1999 Oct 22;293(2):333-42 PMID: 10529349
  16. Characterization of a eukaryotic type serine/threonine protein kinase and protein phosphatase of Streptococcus pneumoniae and identification of kinase substrates.
    FEBS J. 2005 Mar;272(5):1243-54 PMID: 15720398
  17. Network biology: understanding the cell's functional organization.
    Nat Rev Genet. 2004 Feb;5(2):101-13 PMID: 14735121
  18. SWISS-PROT: connecting biomolecular knowledge via a protein database.
    Curr Issues Mol Biol. 2001 Jul;3(3):47-55 PMID: 11488411
  19. Basic charge clusters and predictions of membrane protein topology.
    J Chem Inf Comput Sci. 2002 May-Jun;42(3):620-32 PMID: 12086524
  20. Protein phosphatases--a phylogenetic perspective.
    Chem Rev. 2001 Aug;101(8):2291-312 PMID: 11749374
  21. The number of protein folds and their distribution over families in nature.
    Proteins. 2004 Feb 15;54(3):491-9 PMID: 14747997
  22. The EMBL Nucleotide Sequence Database: major new developments.
    Nucleic Acids Res. 2003 Jan 1;31(1):17-22 PMID: 12519939
  23. InterPro, progress and status in 2005.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D201-5 PMID: 15608177
  24. An overview of Ensembl.
    Genome Res. 2004 May;14(5):925-8 PMID: 15078858
  25. Protein kinases and phosphatases: the yin and yang of protein phosphorylation and signaling.
    Cell. 1995 Jan 27;80(2):225-36 PMID: 7834742
  26. Protein phosphatase 2Calpha inhibits the human stress-responsive p38 and JNK MAPK pathways.
    EMBO J. 1998 Aug 17;17 (16):4744-52 PMID: 9707433
  27. Inhibition of experimental asthma by indoleamine 2,3-dioxygenase.
    J Clin Invest. 2004 Jul;114(2):270-9 PMID: 15254594
  28. Finding families for genomic ORFans.
    Bioinformatics. 1999 Sep;15(9):759-62 PMID: 10498776
  29. Plant PP2C phosphatases: emerging functions in stress signaling.
    Trends Plant Sci. 2004 May;9(5):236-43 PMID: 15130549
  30. The Protein Information Resource.
    Nucleic Acids Res. 2003 Jan 1;31(1):345-7 PMID: 12520019
  31. Probing the function of conserved residues in the serine/threonine phosphatase PP2Calpha.
    Biochemistry. 2003 Jul 22;42(28):8513-21 PMID: 12859198
  32. Resistance of spores of Bacillus species to ultraviolet light.
    Environ Mol Mutagen. 2001;38(2-3):97-104 PMID: 11746741
  33. Inhibition of indoleamine 2,3-dioxygenase, an immunoregulatory target of the cancer suppression gene Bin1, potentiates cancer chemotherapy.
    Nat Med. 2005 Mar;11(3):312-9 PMID: 15711557
  34. Clustering of highly homologous sequences to reduce the size of large protein databases.
    Bioinformatics. 2001 Mar;17(3):282-3 PMID: 11294794
  35. Community genomics among stratified microbial assemblages in the ocean's interior.
    Science. 2006 Jan 27;311(5760):496-503 PMID: 16439655
  36. Regulation of glutamine synthetase. XII. Electron microscopy of the enzyme from Escherichia coli.
    Biochemistry. 1968 Jun;7(6):2143-52 PMID: 4873173
  37. Stress-induced protein phosphatase 2C is a negative regulator of a mitogen-activated protein kinase.
    J Biol Chem. 2003 May 23;278(21):18945-52 PMID: 12646559
  38. Structural genomics: an overview.
    Prog Biophys Mol Biol. 2000;73(5):289-95 PMID: 11063776
  39. A new ATP-independent DNA endonuclease from Schizosaccharomyces pombe that recognizes cyclobutane pyrimidine dimers and 6-4 photoproducts.
    Nucleic Acids Res. 1994 Aug 11;22(15):3026-32 PMID: 8065916
  40. Characterization of PrpC from Bacillus subtilis, a member of the PPM phosphatase family.
    J Bacteriol. 2000 Oct;182(19):5634-8 PMID: 10986276
  41. Metagenomics: DNA sequencing of environmental samples.
    Nat Rev Genet. 2005 Nov;6(11):805-14 PMID: 16304596
  42. GenBank.
    Nucleic Acids Res. 2003 Jan 1;31(1):23-7 PMID: 12519940
  43. Crystal structure of a RuBisCO-like protein from the green sulfur bacterium Chlorobium tepidum.
    Structure. 2005 May;13(5):779-89 PMID: 15893668
  44. Genomic analysis of uncultured marine viral communities.
    Proc Natl Acad Sci U S A. 2002 Oct 29;99(22):14250-5 PMID: 12384570
  45. A unifold, mesofold, and superfold model of protein fold use.
    Proteins. 2002 Jan 1;46(1):61-71 PMID: 11746703
  46. Crystal structure of the protein serine/threonine phosphatase 2C at 2.0 A resolution.
    EMBO J. 1996 Dec 16;15(24):6798-809 PMID: 9003755
  47. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  48. Phosphoprotein phosphatase of Mycobacterium tuberculosis dephosphorylates serine-threonine kinases PknA and PknB.
    Biochem Biophys Res Commun. 2003 Nov 7;311(1):112-20 PMID: 14575702
  49. The complete genome sequence of Chlorobium tepidum TLS, a photosynthetic, anaerobic, green-sulfur bacterium.
    Proc Natl Acad Sci U S A. 2002 Jul 9;99(14):9509-14 PMID: 12093901
  50. Marine phage genomics: what have we learned?
    Curr Opin Biotechnol. 2005 Jun;16(3):299-307 PMID: 15961031
  51. The protein phosphatase 2C (PP2C) superfamily: detection of bacterial homologues.
    Protein Sci. 1996 Jul;5(7):1421-5 PMID: 8819174
  52. Diversity and population structure of a near-shore marine-sediment viral community.
    Proc Biol Sci. 2004 Mar 22;271(1539):565-74 PMID: 15156913
  53. The TIGRFAMs database of protein families.
    Nucleic Acids Res. 2003 Jan 1;31(1):371-3 PMID: 12520025
  54. DNA Data Bank of Japan (DDBJ) in XML.
    Nucleic Acids Res. 2003 Jan 1;31(1):13-6 PMID: 12519938
  55. Three Prochlorococcus cyanophage genomes: signature features and ecological interpretations.
    PLoS Biol. 2005 May;3(5):e144 PMID: 15828858
  56. MEROPS: the peptidase database.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D270-2 PMID: 16381862
  57. MUSCLE: multiple sequence alignment with high accuracy and high throughput.
    Nucleic Acids Res. 2004 Mar 19;32(5):1792-7 PMID: 15034147
  58. Predictions of gene family distributions in microbial genomes: evolution by gene duplication and modification.
    Phys Rev Lett. 2000 Sep 18;85(12):2641-4 PMID: 10978127
  59. The impact of structural genomics: expectations and outcomes.
    Science. 2006 Jan 20;311(5759):347-51 PMID: 16424331
  60. Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation.
    Bioinformatics. 2003 Jul 1;19(10):1275-83 PMID: 12835272
  61. A whole-genome assembly of Drosophila.
    Science. 2000 Mar 24;287(5461):2196-204 PMID: 10731133
  62. Inference of protein function and protein linkages in Mycobacterium tuberculosis based on prokaryotic genome organization: a combined computational approach.
    Genome Biol. 2003;4(9):R59 PMID: 12952538
  63. The Sorcerer II Global Ocean Sampling expedition: northwest Atlantic through eastern tropical Pacific.
    PLoS Biol. 2007 Mar;5(3):e77 PMID: 17355176
  64. ProClust: improved clustering of protein sequences with an extended graph-based approach.
    Bioinformatics. 2002;18 Suppl 2:S182-91 PMID: 12386002
  65. Evidence of a large novel gene pool associated with prokaryotic genomic islands.
    PLoS Genet. 2005 Nov;1(5):e62 PMID: 16299586
  66. Finishing a whole-genome shotgun: release 3 of the Drosophila melanogaster euchromatic genome sequence.
    Genome Biol. 2002;3(12):RESEARCH0079 PMID: 12537568
  67. Structural genomics.
    Methods Biochem Anal. 2003;44:591-612 PMID: 12647406
  68. Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.
    J Mol Biol. 2001 Nov 2;313(4):673-81 PMID: 11697896
  69. Community structure and metabolism through reconstruction of microbial genomes from the environment.
    Nature. 2004 Mar 4;428(6978):37-43 PMID: 14961025
  70. A ribulose-1,5-bisphosphate carboxylase/oxygenase (RubisCO)-like protein from Chlorobium tepidum that is involved with sulfur metabolism and the response to oxidative stress.
    Proc Natl Acad Sci U S A. 2001 Apr 10;98(8):4397-402 PMID: 11287671
  71. Viral metagenomics.
    Nat Rev Microbiol. 2005 Jun;3(6):504-10 PMID: 15886693
  72. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  73. Tolerating some redundancy significantly speeds up clustering of large protein databases.
    Bioinformatics. 2002 Jan;18(1):77-82 PMID: 11836214
  74. PP2C phosphatases Ptc2 and Ptc3 are required for DNA checkpoint inactivation after a double-strand break.
    Mol Cell. 2003 Mar;11(3):827-35 PMID: 12667463
  75. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D173-80 PMID: 16381840
  76. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  77. Purification and cloning of Micrococcus luteus ultraviolet endonuclease, an N-glycosylase/abasic lyase that proceeds via an imino enzyme-DNA intermediate.
    J Biol Chem. 1995 Oct 6;270(40):23475-84 PMID: 7559510
  78. Novel subunit-subunit interactions in the structure of glutamine synthetase.
    Nature. 1986 Sep 25-Oct 1;323(6086):304-9 PMID: 2876389
  79. Enzymatic photoreactivation: 50 years and counting.
    Mutat Res. 2000 Jun 30;451(1-2):25-37 PMID: 10915863
  80. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  81. STRING: known and predicted protein-protein associations, integrated and transferred across organisms.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D433-7 PMID: 15608232
  82. Pfam: multiple sequence alignments and HMM-profiles of protein domains.
    Nucleic Acids Res. 1998 Jan 1;26(1):320-2 PMID: 9399864
  83. The ProDom database of protein domain families: more emphasis on 3D.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D212-5 PMID: 15608179
  84. Protein sequence databases.
    Curr Opin Chem Biol. 2004 Feb;8(1):76-80 PMID: 15036160
  85. Ptc1, a type 2C Ser/Thr phosphatase, inactivates the HOG pathway by dephosphorylating the mitogen-activated protein kinase Hog1.
    Mol Cell Biol. 2001 Jan;21(1):51-60 PMID: 11113180
  86. Bacillus subtilis glutamine synthetase mutants pleiotropically altered in glucose catabolite repression.
    J Bacteriol. 1984 Feb;157(2):612-21 PMID: 6141156
  87. Did evolution leap to create the protein universe?
    Curr Opin Struct Biol. 2002 Jun;12(3):409-16 PMID: 12127462
  88. Analysis of the virus population present in equine faeces indicates the presence of hundreds of uncharacterized virus genomes.
    Virus Genes. 2005 Mar;30(2):151-6 PMID: 15744573
  89. Comparison of sequence profiles. Strategies for structural predictions using sequence information.
    Protein Sci. 2000 Feb;9(2):232-41 PMID: 10716175
  90. Evolution of protein structures and functions.
    Curr Opin Struct Biol. 2002 Jun;12(3):400-8 PMID: 12127461
  91. Genome streamlining in a cosmopolitan oceanic bacterium.
    Science. 2005 Aug 19;309(5738):1242-5 PMID: 16109880
  92. CATH--a hierarchic classification of protein domain structures.
    Structure. 1997 Aug 15;5(8):1093-108 PMID: 9309224
  93. Exhaustive enumeration of protein domain families.
    J Mol Biol. 2003 May 2;328(3):749-67 PMID: 12706730
  94. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  95. The Pfam protein families database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41 PMID: 14681378
  96. The ProDom database of protein domain families.
    Nucleic Acids Res. 1998 Jan 1;26(1):323-6 PMID: 9399865
  97. Comparison of the sequences of Turbo and Sulculus indoleamine dioxygenase-like myoglobin genes.
    Gene. 2003 Apr 10;308:89-94 PMID: 12711393
  98. PAML: a program package for phylogenetic analysis by maximum likelihood.
    Comput Appl Biosci. 1997 Oct;13(5):555-6 PMID: 9367129
  99. Update on the pfam5000 strategy for selection of structural genomics targets.
    Conf Proc IEEE Eng Med Biol Soc. 2005;1:751-5 PMID: 17282292
  100. Scaling law in sizes of protein sequence families: from super-families to orphan genes.
    Proteins. 2003 Jun 1;51(4):569-76 PMID: 12784216
  101. Reverse methanogenesis: testing the hypothesis with environmental genomics.
    Science. 2004 Sep 3;305(5689):1457-62 PMID: 15353801
  102. Evolution of function in protein superfamilies, from a structural perspective.
    J Mol Biol. 2001 Apr 6;307(4):1113-43 PMID: 11286560
  103. Bacterial genomes as new gene homes: the genealogy of ORFans in E. coli.
    Genome Res. 2004 Jun;14(6):1036-42 PMID: 15173110
  104. Bacillus subtilis glutamine synthetase. Purification and physical characterization.
    J Biol Chem. 1970 Oct 25;245(20):5195-205 PMID: 4990297
  105. Transfer of photosynthesis genes to and from Prochlorococcus viruses.
    Proc Natl Acad Sci U S A. 2004 Jul 27;101(30):11013-8 PMID: 15256601
  106. Evolution of the glutamine synthetase gene, one of the oldest existing and functioning genes.
    Proc Natl Acad Sci U S A. 1993 Apr 1;90(7):3009-13 PMID: 8096645
  107. Environmental genome shotgun sequencing of the Sargasso Sea.
    Science. 2004 Apr 2;304(5667):66-74 PMID: 15001713
  108. The Protein Data Bank and structural genomics.
    Nucleic Acids Res. 2003 Jan 1;31(1):489-91 PMID: 12520059
  109. Structural and functional diversity of the microbial kinome.
    PLoS Biol. 2007 Mar;5(3):e17 PMID: 17355172
  110. Implications of structural genomics target selection strategies: Pfam5000, whole genome, and random approaches.
    Proteins. 2005 Jan 1;58(1):166-79 PMID: 15521074
  111. JEvTrace: refinement and variations of the evolutionary trace in JAVA.
    Genome Biol. 2002;3(12):RESEARCH0077 PMID: 12537566
  112. ProtoNet: hierarchical classification of the protein space.
    Nucleic Acids Res. 2003 Jan 1;31(1):348-52 PMID: 12520020
  113. The TIGR gene indices: reconstruction and representation of expressed gene sequences.
    Nucleic Acids Res. 2000 Jan 1;28(1):141-5 PMID: 10592205
  114. Who's your neighbor? New computational approaches for functional genomics.
    Nat Biotechnol. 2000 Jun;18(6):609-13 PMID: 10835597
  115. Comparative metagenomics of microbial communities.
    Science. 2005 Apr 22;308(5721):554-7 PMID: 15845853
  116. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  117. TIGRFAMs: a protein family resource for the functional identification of proteins.
    Nucleic Acids Res. 2001 Jan 1;29(1):41-3 PMID: 11125044
  118. TREE-PUZZLE: maximum likelihood phylogenetic analysis using quartets and parallel computing.
    Bioinformatics. 2002 Mar;18(3):502-4 PMID: 11934758
  119. Genome sequence of Oceanobacillus iheyensis isolated from the Iheya Ridge and its unexpected adaptive capabilities to extreme environments.
    Nucleic Acids Res. 2002 Sep 15;30(18):3927-35 PMID: 12235376
  120. RegulonDB (version 4.0): transcriptional regulation, operon organization and growth conditions in Escherichia coli K-12.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D303-6 PMID: 14681419
  121. Structure-function relationships of glutamine synthetases.
    Biochim Biophys Acta. 2000 Mar 7;1477(1-2):122-45 PMID: 10708854
  122. Crystal structure of T4 endonuclease V. An excision repair enzyme for a pyrimidine dimer.
    Ann N Y Acad Sci. 1994 Jul 29;726:198-207 PMID: 8092676
  123. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  124. Structural genomics: an approach to the protein folding problem.
    Proc Natl Acad Sci U S A. 2001 Nov 20;98(24):13488-9 PMID: 11717420
  125. PknB kinase activity is regulated by phosphorylation in two Thr residues and dephosphorylation by PstP, the cognate phospho-Ser/Thr phosphatase, in Mycobacterium tuberculosis.
    Mol Microbiol. 2003 Sep;49(6):1493-508 PMID: 12950916
  126. QuickJoin--fast neighbour-joining tree reconstruction.
    Bioinformatics. 2004 Nov 22;20(17):3261-2 PMID: 15201185
  127. Murine plasmacytoid dendritic cells initiate the immunosuppressive pathway of tryptophan catabolism in response to CD200 receptor engagement.
    J Immunol. 2004 Sep 15;173(6):3748-54 PMID: 15356121
  128. A tour of structural genomics.
    Nat Rev Genet. 2001 Oct;2(10):801-9 PMID: 11584296
  129. Genomic islands and the ecology and evolution of Prochlorococcus.
    Science. 2006 Mar 24;311(5768):1768-70 PMID: 16556843
  130. Close linkage of genes encoding glutamine synthetases I and II in Frankia alni CpI1.
    J Bacteriol. 1993 Jun;175(11):3679-84 PMID: 8099074
  131. The COG database: a tool for genome-scale analysis of protein functions and evolution.
    Nucleic Acids Res. 2000 Jan 1;28(1):33-6 PMID: 10592175
  132. The PASTA domain: a beta-lactam-binding domain.
    Trends Biochem Sci. 2002 Sep;27(9):438 PMID: 12217513
  133. Identification of a PD-(D/E)XK-like domain with a novel configuration of the endonuclease active site in the methyl-directed restriction enzyme Mrr and its homologs.
    Gene. 2001 Apr 18;267(2):183-91 PMID: 11313145
  134. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  135. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  136. Metagenomic analyses of an uncultured viral community from human feces.
    J Bacteriol. 2003 Oct;185(20):6220-3 PMID: 14526037
  137. Structural genomics: a pipeline for providing structures for the biologist.
    Protein Sci. 2002 Apr;11(4):723-38 PMID: 11910018
  138. The K(A)/K(S) ratio test for assessing the protein-coding potential of genomic regions: an empirical and simulation study.
    Genome Res. 2002 Jan;12(1):198-202 PMID: 11779845
  139. Origins of highly mosaic mycobacteriophage genomes.
    Cell. 2003 Apr 18;113(2):171-82 PMID: 12705866
  140. UniProt: the Universal Protein knowledgebase.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D115-9 PMID: 14681372
  141. A functional link between RuBisCO-like protein of Bacillus and photosynthetic RuBisCO.
    Science. 2003 Oct 10;302(5643):286-90 PMID: 14551435
Article Info
Journal
PLoS biology
Abbr.
PLoS Biol
ISSN
1545-7885
Published
2007-03-00
Pages
e16
Language
English
Region
United States
NLM ID
101183755
PMCID
PMC1821046
Subset
IM
Grants
NHGRI NIH HHS · T32 HG000047 · United States
NIGMS NIH HHS · R01 GM073109 · United States
NCRR NIH HHS · 5U54 RR020843-02 · United States
NHGRI NIH HHS · 5T32 HG00047 · United States
NHGRI NIH HHS · K22 HG000056 · United States
NHGRI NIH HHS · K22 HG00056 · United States
NIGMS NIH HHS · P20 GM068136 · United States
NCRR NIH HHS · U54 RR020843 · United States
NCI NIH HHS · P30 CA014195 · United States
Corrections
CommentIn
CommentIn
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com