Home LiteratureArticle Details
PMID: 15766383 Published · epublish English Journal Article

Speeding disease gene discovery by sequence based candidate prioritization.

BMC bioinformatics ·Vol. 6 ·2005-03-14 ·Pages 55

Adie EA, Adams RR, Evans KL, Porteous DJ, Pickard BS

Abstract

Regions of interest identified through genetic linkage studies regularly exceed 30 centimorgans in size and can contain hundreds of genes. Traditionally this number is reduced by matching functional annotation to knowledge of the disease or phenotype in question. However, here we show that disease genes share patterns of sequence-based features that can provide a good basis for automatic prioritization of candidates by machine learning. We examined a variety of sequence-based features and found that for many of them there are significant differences between the sets of genes known to be involved in human hereditary disease and those not known to be involved in disease. We have created an automatic classifier called PROSPECTR based on those features using the alternating decision tree algorithm which ranks genes in the order of likelihood of involvement in disease. On average, PROSPECTR enriches lists for disease genes two-fold 77% of the time, five-fold 37% of the time and twenty-fold 11% of the time. PROSPECTR is a simple and effective way to identify genes involved in Mendelian and oligogenic disorders. It performs markedly better than the single existing sequence-based classifier on novel data. PROSPECTR could save investigators looking at large regions of interest time and effort by prioritizing positional candidate genes for mutation detection and case-control association studies.

MeSH Terms
Algorithms Automation Computational Biology/methods Conserved Sequence Databases, Genetic Databases, Nucleic Acid Databases, Protein Decision Trees Gene Expression Profiling Genetic Diseases, Inborn/genetics Genetic Linkage Genetic Predisposition to Disease Genetic Testing Genome Genome, Human Humans Linkage Disequilibrium Models, Genetic Models, Statistical Phenotype Polymorphism, Genetic ROC Curve Reproducibility of Results Research Design Sequence Analysis, DNA Software Time Factors
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Adie Euan A
Medical Genetics Section, Department of Medical Sciences, The University of Edinburgh, Edinburgh, UK. euan.adie@ed.ac.uk
Adams Richard R
Evans Kathryn L
Porteous David J
Pickard Ben S
References (26)
26 references, click to expand
  1. The InterPro Database, 2003 brings increased coverage and new features.
    Nucleic Acids Res. 2003 Jan 1;31(1):315-8 PMID: 12520011
  2. Associations between human disease genes and overlapping gene groups and multiple amino acid runs.
    Proc Natl Acad Sci U S A. 2002 Dec 24;99(26):17008-13 PMID: 12473749
  3. Human Gene Mutation Database (HGMD): 2003 update.
    Hum Mutat. 2003 Jun;21(6):577-81 PMID: 12754702
  4. New methods for finding disease-susceptibility genes: impact and potential.
    Genome Biol. 2003;4(10):119 PMID: 14519189
  5. Human disease genes: patterns and predictions.
    Gene. 2003 Oct 30;318:169-75 PMID: 14585509
  6. POCUS: mining genomic sequence annotation to predict disease genes.
    Genome Biol. 2003;4(11):R75 PMID: 14611661
  7. Gene length and proximity to neighbors affect genome-wide expression levels.
    Genome Res. 2003 Dec;13(12):2602-8 PMID: 14613975
  8. Elevated rates of protein secretion, evolution, and disease among tissue-specific genes.
    Genome Res. 2004 Jan;14(1):54-61 PMID: 14707169
  9. The genetic association database.
    Nat Genet. 2004 May;36(5):431-2 PMID: 15118671
  10. Genome information resources - developments at Ensembl.
    Trends Genet. 2004 Jun;20(6):268-72 PMID: 15145580
  11. Genome-wide identification of genes likely to be involved in human genetic disease.
    Nucleic Acids Res. 2004;32(10):3108-14 PMID: 15181176
  12. Overview of commonly used bioinformatics methods and their applications.
    Ann N Y Acad Sci. 2004 May;1020:10-21 PMID: 15208179
  13. Evolutionary conservation and selection of human disease gene orthologs in the rat and mouse genomes.
    Genome Biol. 2004;5(7):R47 PMID: 15239832
  14. Data mining in bioinformatics using Weka.
    Bioinformatics. 2004 Oct 12;20(15):2479-81 PMID: 15073010
  15. CpG islands in vertebrate genomes.
    J Mol Biol. 1987 Jul 20;196(2):261-82 PMID: 3656447
  16. Classification-algorithm evaluation: five performance measures based on confusion matrices.
    J Clin Monit. 1995 May;11(3):189-206 PMID: 7623060
  17. Translational efficiency is regulated by the length of the 3' untranslated region.
    Mol Cell Biol. 1996 Jan;16(1):146-56 PMID: 8524291
  18. 'Going wrong with confidence': misleading sequence analyses of CiaB and clpX.
    Mol Microbiol. 1999 Oct;34(1):195 PMID: 10540297
  19. Intrinsic errors in genome annotation.
    Trends Genet. 2001 Aug;17(8):429-31 PMID: 11485799
  20. Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders.
    Nucleic Acids Res. 2002 Jan 1;30(1):52-5 PMID: 11752252
  21. Large-scale analysis of the human and mouse transcriptomes.
    Proc Natl Acad Sci U S A. 2002 Apr 2;99(7):4465-70 PMID: 11904358
  22. Association of genes to genetically inherited diseases using data mining.
    Nat Genet. 2002 Jul;31(3):316-9 PMID: 12006977
  23. A similarity-based method for genome-wide prediction of disease-relevant human genes.
    Bioinformatics. 2002;18 Suppl 2:S110-5 PMID: 12385992
  24. Modeling the percolation of annotation errors in a database of protein sequences.
    Bioinformatics. 2002 Dec;18(12):1641-9 PMID: 12490449
  25. Finding genes that underlie complex traits.
    Science. 2002 Dec 20;298(5602):2345-9 PMID: 12493905
  26. A new web-based data mining tool for the identification of candidate genes for human genetic disorders.
    Eur J Hum Genet. 2003 Jan;11(1):57-63 PMID: 12529706
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2005-03-14
Epub
2005-00-14
Pages
55
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1274252
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com