Home LiteratureArticle Details
PMID: 16314312 Published · epublish English Comparative Study Evaluation Study Journal Article Research Support, N.I.H., Extramural

Gene identification in novel eukaryotic genomes by self-training algorithm.

Nucleic acids research ·Vol. 33 ·No. 20 ·2005-00-00 ·Pages 6494-506

Lomsadze A, Ter-Hovhannisyan V, Chernoff YO, Borodovsky M

Abstract

Finding new protein-coding genes is one of the most important goals of eukaryotic genome sequencing projects. However, genomic organization of novel eukaryotic genomes is diverse and ab initio gene finding tools tuned up for previously studied species are rarely suitable for efficacious gene hunting in DNA sequences of a new genome. Gene identification methods based on cDNA and expressed sequence tag (EST) mapping to genomic DNA or those using alignments to closely related genomes rely either on existence of abundant cDNA and EST data and/or availability on reference genomes. Conventional statistical ab initio methods require large training sets of validated genes for estimating gene model parameters. In practice, neither one of these types of data may be available in sufficient amount until rather late stages of the novel genome sequencing. Nevertheless, we have shown that gene finding in eukaryotic genomes could be carried out in parallel with statistical models estimation directly from yet anonymous genomic DNA. The suggested method of parallelization of gene prediction with the model parameters estimation follows the path of the iterative Viterbi training. Rounds of genomic sequence labeling into coding and non-coding regions are followed by the rounds of model parameters estimation. Several dynamically changing restrictions on the possible range of model parameters are added to filter out fluctuations in the initial steps of the algorithm that could redirect the iteration process away from the biologically relevant point in parameter space. Tests on well-studied eukaryotic genomes have shown that the new method performs comparably or better than conventional methods where the supervised model training precedes the gene prediction step. Several novel genomes have been analyzed and biologically interesting findings are discussed. Thus, a self-training algorithm that had been assumed feasible only for prokaryotic genomes has now been developed for ab initio eukaryotic gene identification.

MeSH Terms
Algorithms Animals Exons Genes Genome Genomics/methods Markov Chains Phylogeny Reproducibility of Results
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Lomsadze Alexandre
School of Biology, Georgia Institute of Technology, Atlanta, GA 30332-0230, USA.
Ter-Hovhannisyan Vardges
Chernoff Yury O
Borodovsky Mark
References (37)
37 references, click to expand
  1. Gene prediction with a hidden Markov model and a new intron submodel.
    Bioinformatics. 2003 Oct;19 Suppl 2:ii215-25 PMID: 14534192
  2. GeneSeqer@PlantGDB: Gene structure prediction in plant genomes.
    Nucleic Acids Res. 2003 Jul 1;31(13):3597-600 PMID: 12824374
  3. Gene finding in novel genomes.
    BMC Bioinformatics. 2004 May 14;5:59 PMID: 15144565
  4. Computer methods to locate signals in nucleic acid sequences.
    Nucleic Acids Res. 1984 Jan 11;12(1 Pt 2):505-19 PMID: 6364039
  5. Identification of protein coding regions by database similarity search.
    Nat Genet. 1993 Mar;3(3):266-72 PMID: 8485583
  6. A weight array method for splicing signal analysis.
    Comput Appl Biosci. 1993 Oct;9(5):499-509 PMID: 8293321
  7. Evaluation of gene structure prediction programs.
    Genomics. 1996 Jun 15;34(3):353-67 PMID: 8786136
  8. Gene recognition via spliced sequence alignment.
    Proc Natl Acad Sci U S A. 1996 Aug 20;93(17):9061-6 PMID: 8799154
  9. Gene structure prediction using information on homologous protein sequence.
    Comput Appl Biosci. 1996 Jun;12(3):161-70 PMID: 8872383
  10. Prediction of complete gene structures in human genomic DNA.
    J Mol Biol. 1997 Apr 25;268(1):78-94 PMID: 9149143
  11. Two methods for improving performance of an HMM and their application for gene finding.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:179-86 PMID: 9322033
  12. Integrating database homology in a probabilistic gene structure model.
    Pac Symp Biocomput. 1997;:232-44 PMID: 9390295
  13. GeneMark.hmm: new solutions for gene finding.
    Nucleic Acids Res. 1998 Feb 15;26(4):1107-15 PMID: 9461475
  14. Combining diverse evidence for gene recognition in completely sequenced bacterial genomes.
    Nucleic Acids Res. 1998 Jun 15;26(12):2941-7 PMID: 9611239
  15. Finding intron/exon splice junctions using INFO, INterruption Finder and Organizer.
    J Comput Biol. 1998 Summer;5(2):307-21 PMID: 9672834
  16. Self-identification of protein-coding regions in microbial genomes.
    Proc Natl Acad Sci U S A. 1998 Aug 18;95(17):10026-31 PMID: 9707594
  17. Performance-guarantee gene predictions via spliced alignment.
    Genomics. 1998 Aug 1;51(3):332-9 PMID: 9721203
  18. A computer program for aligning a cDNA sequence with a genomic DNA sequence.
    Genome Res. 1998 Sep;8(9):967-74 PMID: 9750195
  19. How to interpret an anonymous bacterial genome: machine learning approach to gene identification.
    Genome Res. 1998 Nov;8(11):1154-71 PMID: 9847079
  20. Heuristic approach to deriving models for gene finding.
    Nucleic Acids Res. 1999 Oct 1;27(19):3911-20 PMID: 10481031
  21. A dictionary-based approach for gene annotation.
    J Comput Biol. 1999 Fall-Winter;6(3-4):419-30 PMID: 10582576
  22. Evaluation of gene prediction software using a genomic data set: application to Arabidopsis thaliana sequences.
    Bioinformatics. 1999 Nov;15(11):887-99 PMID: 10743555
  23. GeneID in Drosophila.
    Genome Res. 2000 Apr;10(4):511-5 PMID: 10779490
  24. Genie--gene finding in Drosophila melanogaster.
    Genome Res. 2000 Apr;10(4):529-38 PMID: 10779493
  25. On the convergence of a clustering algorithm for protein-coding regions in microbial genomes.
    Bioinformatics. 2000 Apr;16(4):367-71 PMID: 10869034
  26. Human and mouse gene structure: comparative analysis and application to exon prediction.
    Genome Res. 2000 Jul;10(7):950-8 PMID: 10899144
  27. Conservation, regulation, synteny, and introns in a large-scale C. briggsae-C. elegans genomic alignment.
    Genome Res. 2000 Aug;10(8):1115-25 PMID: 10958630
  28. Analysis of the genome sequence of the flowering plant Arabidopsis thaliana.
    Nature. 2000 Dec 14;408(6814):796-815 PMID: 11130711
  29. GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.
    Nucleic Acids Res. 2001 Jun 15;29(12):2607-18 PMID: 11410670
  30. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  31. A draft sequence of the rice genome (Oryza sativa L. ssp. indica).
    Science. 2002 Apr 5;296(5565):79-92 PMID: 11935017
  32. Applications of generalized pair hidden Markov models to alignment and gene finding problems.
    J Comput Biol. 2002;9(2):389-99 PMID: 12015888
  33. Exon discovery by genomic sequence alignment.
    Bioinformatics. 2002 Jun;18(6):777-87 PMID: 12075013
  34. Current methods of gene prediction, their strengths and weaknesses.
    Nucleic Acids Res. 2002 Oct 1;30(19):4103-17 PMID: 12364589
  35. Comparative ab initio prediction of gene structures using pair HMMs.
    Bioinformatics. 2002 Oct;18(10):1309-18 PMID: 12376375
  36. The draft genome of Ciona intestinalis: insights into chordate and vertebrate origins.
    Science. 2002 Dec 13;298(5601):2157-67 PMID: 12481130
  37. EasyGene--a prokaryotic gene finder that ranks ORFs by statistical significance.
    BMC Bioinformatics. 2003 Jun 3;4:21 PMID: 12783628
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2005-00-00
Epub
2005-00-28
Pages
6494-506
Language
English
Region
England
NLM ID
0411011
PMCID
PMC1298918
Subset
IM
Grants
NIGMS NIH HHS · R01 GM058763 · United States
NHGRI NIH HHS · R01 HG000783 · United States
NIGMS NIH HHS · GM58763 · United States
NHGRI NIH HHS · HG00783 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com