Home LiteratureArticle Details
PMID: 15937195 Published · epublish English Evaluation Study Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, P.H.S.

Prediction of solvent accessibility and sites of deleterious mutations from protein sequence.

Nucleic acids research ·Vol. 33 ·No. 10 ·2005-00-00 ·Pages 3193-9

Chen H, Zhou HX

Abstract

Residues that form the hydrophobic core of a protein are critical for its stability. A number of approaches have been developed to classify residues as buried or exposed. In order to optimize the classification, we have refined a suite of five methods over a large dataset and proposed a metamethod based on an ensemble average of the individual methods, leading to a two-state classification accuracy of 80%. Many studies have suggested that hydrophobic core residues are likely sites of deleterious mutations, so we wanted to see to what extent these sites can be predicted from the putative buried residues. Residues that were most confidently classified as buried were proposed as sites of deleterious mutations. This proposition was tested on six proteins for which sites of deleterious mutations have previously been identified by stability measurement or functional assay. Of the total of 130 residues predicted as sites of deleterious mutations, 104 (or 80%) were correct.

MeSH Terms
Amino Acid Sequence Amino Acids/chemistry,classification Hydrophobic and Hydrophilic Interactions Molecular Sequence Data Mutation Proteins/chemistry,genetics Reproducibility of Results Sequence Analysis, Protein/methods Solvents/chemistry
Chemicals
Amino Acids Proteins Solvents
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Chen Huiling
Department of Physics, Drexel University Philadelphia, PA 19104, USA.
Zhou Huan-Xiang
References (35)
35 references, click to expand
  1. Human non-synonymous SNPs: server and survey.
    Nucleic Acids Res. 2002 Sep 1;30(17):3894-900 PMID: 12202775
  2. Evaluation of structural and evolutionary contributions to deleterious mutation prediction.
    J Mol Biol. 2002 Sep 27;322(4):891-901 PMID: 12270722
  3. SIFT: Predicting amino acid changes that affect protein function.
    Nucleic Acids Res. 2003 Jul 1;31(13):3812-4 PMID: 12824425
  4. Accurate prediction of solvent accessibility using neural networks-based regression.
    Proteins. 2004 Sep 1;56(4):753-67 PMID: 15281128
  5. Environment and exposure to solvent of protein atoms. Lysozyme and insulin.
    J Mol Biol. 1973 Sep 15;79(2):351-71 PMID: 4760134
  6. Mutations in lambda repressor's amino-terminal domain: implications for protein stability and DNA binding.
    Proc Natl Acad Sci U S A. 1983 May;80(9):2676-80 PMID: 6221342
  7. Dictionary of protein secondary structure: pattern recognition of hydrogen-bonded and geometrical features.
    Biopolymers. 1983 Dec;22(12):2577-637 PMID: 6667333
  8. Complete mutagenesis of the HIV-1 protease.
    Nature. 1989 Aug 3;340(6232):397-400 PMID: 2666861
  9. A symbolic-numeric approach to find patterns in genomes. Application to the translation initiation sites of E. coli.
    Biochimie. 1999 Nov;81(11):1065-72 PMID: 10575363
  10. Decision tree-based formation of consensus protein secondary structure prediction.
    Bioinformatics. 1999 Dec;15(12):1039-46 PMID: 10745994
  11. New method for accurate prediction of solvent accessibility from protein sequence.
    Proteins. 2001 Jan 1;42(1):1-5 PMID: 11093255
  12. Fold recognition and accurate query-template alignment by a combination of PSI-BLAST and threading.
    Proteins. 2001 Jan 1;42(1):23-37 PMID: 11093258
  13. Support vector machine classification and validation of cancer tissue samples using microarray expression data.
    Bioinformatics. 2000 Oct;16(10):906-14 PMID: 11120680
  14. Prediction of protein surface accessibility with information theory.
    Proteins. 2001 Mar 1;42(4):452-9 PMID: 11170200
  15. SNPs, protein structure, and disease.
    Hum Mutat. 2001 Apr;17(4):263-70 PMID: 11295823
  16. Multi-class protein fold recognition using support vector machines and neural networks.
    Bioinformatics. 2001 Apr;17(4):349-58 PMID: 11301304
  17. A novel method of protein secondary structure prediction with high segment overlap measure: support vector machine approach.
    J Mol Biol. 2001 Apr 27;308(2):397-407 PMID: 11327775
  18. Prediction of protein interaction sites from sequence profile and residue neighbor list.
    Proteins. 2001 Aug 15;44(3):336-43 PMID: 11455607
  19. Origins of structure in globular proteins.
    Proc Natl Acad Sci U S A. 1990 Aug;87(16):6388-92 PMID: 2385597
  20. Contributions of the large hydrophobic amino acids to the stability of staphylococcal nuclease.
    Biochemistry. 1990 Sep 4;29(35):8033-41 PMID: 2261461
  21. Systematic mutation of bacteriophage T4 lysozyme.
    J Mol Biol. 1991 Nov 5;222(1):67-88 PMID: 1942069
  22. Conservation and prediction of solvent accessibility in protein families.
    Proteins. 1994 Nov;20(3):216-26 PMID: 7892171
  23. Relationship between in vivo activity and in vitro measures of function and stability of a protein.
    Biochemistry. 1995 Sep 19;34(37):11970-8 PMID: 7547934
  24. Predicting solvent accessibility: higher accuracy using Bayesian statistics and optimized residue substitution classes.
    Proteins. 1996 May;25(1):38-47 PMID: 8727318
  25. Mapping the protein universe.
    Science. 1996 Aug 2;273(5275):595-603 PMID: 8662544
  26. Genetic studies of the Lac repressor. XV: 4000 single amino acid substitutions and analysis of the resulting phenotypes on the basis of the protein structure.
    J Mol Biol. 1996 Aug 30;261(4):509-23 PMID: 8794873
  27. A procedure for the prediction of temperature-sensitive mutants of a globular protein based solely on the amino acid sequence.
    Proc Natl Acad Sci U S A. 1996 Nov 26;93(24):13908-13 PMID: 8943034
  28. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  29. Adaptation of protein surfaces to subcellular location.
    J Mol Biol. 1998 Feb 20;276(2):517-25 PMID: 9512720
  30. Prediction of protein hydration sites from sequence by modular neural networks.
    Protein Eng. 1998 Jan;11(1):11-9 PMID: 9579655
  31. A decision tree system for finding genes in DNA.
    J Comput Biol. 1998 Winter;5(4):667-80 PMID: 10072083
  32. Protein secondary structure prediction based on position-specific scoring matrices.
    J Mol Biol. 1999 Sep 17;292(2):195-202 PMID: 10493868
  33. Prediction of interface residues in protein-protein complexes by a consensus neural network method: test against NMR data.
    Proteins. 2005 Oct 1;61(1):21-35 PMID: 16080151
  34. Prediction of coordination number and relative solvent accessibility in proteins.
    Proteins. 2002 May 1;47(2):142-53 PMID: 11933061
  35. Prediction of protein solvent accessibility using support vector machines.
    Proteins. 2002 Aug 15;48(3):566-70 PMID: 12112679
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2005-00-00
Epub
2005-00-03
Pages
3193-9
Language
English
Region
England
NLM ID
0411011
PMCID
PMC1142490
Subset
IM
Grants
NIGMS NIH HHS · R01 GM058187 · United States
NIGMS NIH HHS · GM58187 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com