Abstract
Regions of interest identified through genetic linkage studies regularly exceed 30 centimorgans in size and can contain hundreds of genes. Traditionally this number is reduced by matching functional annotation to knowledge of the disease or phenotype in question. However, here we show that disease genes share patterns of sequence-based features that can provide a good basis for automatic prioritization of candidates by machine learning. We examined a variety of sequence-based features and found that for many of them there are significant differences between the sets of genes known to be involved in human hereditary disease and those not known to be involved in disease. We have created an automatic classifier called PROSPECTR based on those features using the alternating decision tree algorithm which ranks genes in the order of likelihood of involvement in disease. On average, PROSPECTR enriches lists for disease genes two-fold 77% of the time, five-fold 37% of the time and twenty-fold 11% of the time. PROSPECTR is a simple and effective way to identify genes involved in Mendelian and oligogenic disorders. It performs markedly better than the single existing sequence-based classifier on novel data. PROSPECTR could save investigators looking at large regions of interest time and effort by prioritizing positional candidate genes for mutation detection and case-control association studies.
MeSH Terms
Algorithms
Automation
Computational Biology/methods
Conserved Sequence
Databases, Genetic
Databases, Nucleic Acid
Databases, Protein
Decision Trees
Gene Expression Profiling
Genetic Diseases, Inborn/genetics
Genetic Linkage
Genetic Predisposition to Disease
Genetic Testing
Genome
Genome, Human
Humans
Linkage Disequilibrium
Models, Genetic
Models, Statistical
Phenotype
Polymorphism, Genetic
ROC Curve
Reproducibility of Results
Research Design
Sequence Analysis, DNA
Software
Time Factors
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Adie Euan A
Medical Genetics Section, Department of Medical Sciences, The University of Edinburgh, Edinburgh, UK. euan.adie@ed.ac.uk
Adams Richard R
Evans Kathryn L
Porteous David J
Pickard Ben S
References (26)
26 references, click to expand
-
The InterPro Database, 2003 brings increased coverage and new features.
Nucleic Acids Res. 2003 Jan 1;31(1):315-8
PMID: 12520011
-
Associations between human disease genes and overlapping gene groups and multiple amino acid runs.
Proc Natl Acad Sci U S A. 2002 Dec 24;99(26):17008-13
PMID: 12473749
-
Human Gene Mutation Database (HGMD): 2003 update.
Hum Mutat. 2003 Jun;21(6):577-81
PMID: 12754702
-
New methods for finding disease-susceptibility genes: impact and potential.
Genome Biol. 2003;4(10):119
PMID: 14519189
-
Human disease genes: patterns and predictions.
Gene. 2003 Oct 30;318:169-75
PMID: 14585509
-
POCUS: mining genomic sequence annotation to predict disease genes.
Genome Biol. 2003;4(11):R75
PMID: 14611661
-
Gene length and proximity to neighbors affect genome-wide expression levels.
Genome Res. 2003 Dec;13(12):2602-8
PMID: 14613975
-
Elevated rates of protein secretion, evolution, and disease among tissue-specific genes.
Genome Res. 2004 Jan;14(1):54-61
PMID: 14707169
-
The genetic association database.
Nat Genet. 2004 May;36(5):431-2
PMID: 15118671
-
Genome information resources - developments at Ensembl.
Trends Genet. 2004 Jun;20(6):268-72
PMID: 15145580
-
Genome-wide identification of genes likely to be involved in human genetic disease.
Nucleic Acids Res. 2004;32(10):3108-14
PMID: 15181176
-
Overview of commonly used bioinformatics methods and their applications.
Ann N Y Acad Sci. 2004 May;1020:10-21
PMID: 15208179
-
Evolutionary conservation and selection of human disease gene orthologs in the rat and mouse genomes.
Genome Biol. 2004;5(7):R47
PMID: 15239832
-
Data mining in bioinformatics using Weka.
Bioinformatics. 2004 Oct 12;20(15):2479-81
PMID: 15073010
-
CpG islands in vertebrate genomes.
J Mol Biol. 1987 Jul 20;196(2):261-82
PMID: 3656447
-
Classification-algorithm evaluation: five performance measures based on confusion matrices.
J Clin Monit. 1995 May;11(3):189-206
PMID: 7623060
-
Translational efficiency is regulated by the length of the 3' untranslated region.
Mol Cell Biol. 1996 Jan;16(1):146-56
PMID: 8524291
-
'Going wrong with confidence': misleading sequence analyses of CiaB and clpX.
Mol Microbiol. 1999 Oct;34(1):195
PMID: 10540297
-
Intrinsic errors in genome annotation.
Trends Genet. 2001 Aug;17(8):429-31
PMID: 11485799
-
Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders.
Nucleic Acids Res. 2002 Jan 1;30(1):52-5
PMID: 11752252
-
Large-scale analysis of the human and mouse transcriptomes.
Proc Natl Acad Sci U S A. 2002 Apr 2;99(7):4465-70
PMID: 11904358
-
Association of genes to genetically inherited diseases using data mining.
Nat Genet. 2002 Jul;31(3):316-9
PMID: 12006977
-
A similarity-based method for genome-wide prediction of disease-relevant human genes.
Bioinformatics. 2002;18 Suppl 2:S110-5
PMID: 12385992
-
Modeling the percolation of annotation errors in a database of protein sequences.
Bioinformatics. 2002 Dec;18(12):1641-9
PMID: 12490449
-
Finding genes that underlie complex traits.
Science. 2002 Dec 20;298(5602):2345-9
PMID: 12493905
-
A new web-based data mining tool for the identification of candidate genes for human genetic disorders.
Eur J Hum Genet. 2003 Jan;11(1):57-63
PMID: 12529706