Home LiteratureArticle Details
PMID: 14627837 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

Improving gene annotation of complete viral genomes.

Nucleic acids research ·Vol. 31 ·No. 23 ·2003-12-01 ·Pages 7041-55

Mills R, Rozanov M, Lomsadze A, Tatusova T, Borodovsky M

Abstract

Gene annotation in viruses often relies upon similarity search methods. These methods possess high specificity but some genes may be missed, either those unique to a particular genome or those highly divergent from known homologs. To identify potentially missing viral genes we have analyzed all complete viral genomes currently available in GenBank with a specialized and augmented version of the gene finding program GeneMarkS. In particular, by implementing genome-specific self-training protocols we have better adjusted the GeneMarkS statistical models to sequences of viral genomes. Hundreds of new genes were identified, some in well studied viral genomes. For example, a new gene predicted in the genome of the Epstein-Barr virus was shown to encode a protein similar to alpha-herpesvirus minor tegument protein UL14 with heat shock functions. Convincing evidence of this similarity was obtained after only 12 PSI-BLAST iterations. In another example, several iterations of PSI-BLAST were required to demonstrate that a gene predicted in the genome of Alcelaphine herpesvirus 1 encodes a BALF1-like protein which is thought to be involved in apoptosis regulation and, potentially, carcinogenesis. New predictions were used to refine annotations of viral genomes in the RefSeq collection curated by the National Center for Biotechnology Information. Importantly, even in those cases where no sequence similarities were detected, GeneMarkS significantly reduced the number of primary targets for experimental characterization by identifying the most probable candidate genes. The new genome annotations were stored in VIOLIN, an interactive database which provides access to similarity search tools for up-to-date analysis of predicted viral proteins.

MeSH Terms
Amino Acid Sequence Computational Biology/methods Databases, Genetic Genes, Viral/genetics Genome, Viral Herpesviridae/genetics Humans Molecular Sequence Data Sensitivity and Specificity Sequence Alignment Software Viral Proteins/chemistry,genetics
Chemicals
Viral Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Mills Ryan
School of Biology, Georgia Institute of Technology, Atlanta, GA 30332-0230, USA.
Rozanov Michael
Lomsadze Alexandre
Tatusova Tatiana
Borodovsky Mark
References (35)
35 references, click to expand
  1. A giant virus in amoebae.
    Science. 2003 Mar 28;299(5615):2033 PMID: 12663918
  2. GenBank.
    Nucleic Acids Res. 2002 Jan 1;30(1):17-20 PMID: 11752243
  3. Comparing the predicted and observed properties of proteins encoded in the genome of Escherichia coli K-12.
    Electrophoresis. 1997 Aug;18(8):1259-313 PMID: 9298646
  4. A conserved genetic module that encodes the major virion components in both the coliphage T4 and the marine cyanophage S-PM2.
    Proc Natl Acad Sci U S A. 2001 Sep 25;98(20):11411-6 PMID: 11553768
  5. Genome structure of mycobacteriophage D29: implications for phage evolution.
    J Mol Biol. 1998 May 29;279(1):143-64 PMID: 9636706
  6. Sequence analysis of the complete genome of an iridovirus isolated from the tiger frog.
    Virology. 2002 Jan 20;292(2):185-97 PMID: 11878922
  7. Viral Genome DataBase: storing and analyzing genes and proteins from complete viral genomes.
    Bioinformatics. 2000 May;16(5):484-5 PMID: 10871272
  8. The Genome sequence of the SARS-associated coronavirus.
    Science. 2003 May 30;300(5624):1399-404 PMID: 12730501
  9. VIDA: a virus database system for the organization of animal virus genome open reading frames.
    Nucleic Acids Res. 2001 Jan 1;29(1):133-6 PMID: 11125070
  10. The estimation of statistical parameters for local alignment score distributions.
    Nucleic Acids Res. 2001 Jan 15;29(2):351-61 PMID: 11139604
  11. The complete DNA sequence and analysis of the large virulence plasmid of Escherichia coli O157:H7.
    Nucleic Acids Res. 1998 Sep 15;26(18):4196-204 PMID: 9722640
  12. Improving the accuracy of PSI-BLAST protein database searches with composition-based statistics and other refinements.
    Nucleic Acids Res. 2001 Jul 15;29(14):2994-3005 PMID: 11452024
  13. Detecting subtle sequence signals: a Gibbs sampling strategy for multiple alignment.
    Science. 1993 Oct 8;262(5131):208-14 PMID: 8211139
  14. Sequence analysis of the genome of porcine lymphotropic herpesvirus 1 and gene expression during posttransplant lymphoproliferative disease of pigs.
    Virology. 2002 Mar 15;294(2):383-93 PMID: 12009880
  15. The Human Papillomavirus Database.
    J Biomed Sci. 1995 Apr;2(2):90-104 PMID: 11725046
  16. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2001 Jan 1;29(1):11-6 PMID: 11125038
  17. Epstein-Barr virus BALF1 is a BCL-2-like antagonist of the herpesvirus antiapoptotic BCL-2 proteins.
    J Virol. 2002 Mar;76(5):2469-79 PMID: 11836425
  18. The nucleotide sequence of Shiga toxin (Stx) 2e-encoding phage phiP27 is not related to other Stx phage genomes, but the modular genetic structure is conserved.
    Infect Immun. 2002 Apr;70(4):1896-908 PMID: 11895953
  19. Nucleotide sequence of bacteriophage lambda DNA.
    J Mol Biol. 1982 Dec 25;162(4):729-73 PMID: 6221115
  20. Herpes simplex virus type 2 UL14 gene product has heat shock protein (HSP)-like functions.
    J Cell Sci. 2002 Jun 15;115(Pt 12):2517-27 PMID: 12045222
  21. Heuristic approach to deriving models for gene finding.
    Nucleic Acids Res. 1999 Oct 1;27(19):3911-20 PMID: 10481031
  22. Complete DNA sequence and analysis of the large virulence plasmid of Shigella flexneri.
    Infect Immun. 2001 May;69(5):3271-85 PMID: 11292750
  23. Complete genome sequence of the shrimp white spot bacilliform virus.
    J Virol. 2001 Dec;75(23):11811-20 PMID: 11689662
  24. Complete DNA sequence of the rat cytomegalovirus genome.
    J Virol. 2000 Aug;74(16):7656-65 PMID: 10906222
  25. GeneMark.hmm: new solutions for gene finding.
    Nucleic Acids Res. 1998 Feb 15;26(4):1107-15 PMID: 9461475
  26. GeneMarkS: a self-training method for prediction of gene starts in microbial genomes. Implications for finding sequence motifs in regulatory regions.
    Nucleic Acids Res. 2001 Jun 15;29(12):2607-18 PMID: 11410670
  27. The genome of bacteriophage phiKZ of Pseudomonas aeruginosa.
    J Mol Biol. 2002 Mar 15;317(1):1-19 PMID: 11916376
  28. Complete nucleotide sequence of the mycoplasma virus P1 genome.
    Plasmid. 2001 Mar;45(2):122-6 PMID: 11322826
  29. Genome sequence of bovine herpesvirus 4, a bovine Rhadinovirus, and identification of an origin of DNA replication.
    J Virol. 2001 Feb;75(3):1186-94 PMID: 11152491
  30. DNA sequence and comparison of virulence plasmids from Rhodococcus equi ATCC 33701 and 103.
    Infect Immun. 2000 Dec;68(12):6840-7 PMID: 11083803
  31. Sequence logos: a new way to display consensus sequences.
    Nucleic Acids Res. 1990 Oct 25;18(20):6097-100 PMID: 2172928
  32. Complete genomes in WWW Entrez: data representation and analysis.
    Bioinformatics. 1999 Jul-Aug;15(7-8):536-43 PMID: 10487861
  33. Intrinsic and extrinsic approaches for detecting genes in a bacterial genome.
    Nucleic Acids Res. 1994 Nov 11;22(22):4756-67 PMID: 7984428
  34. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  35. Multiple sequence alignment with hierarchical clustering.
    Nucleic Acids Res. 1988 Nov 25;16(22):10881-90 PMID: 2849754
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2003-12-01
Pages
7041-55
Language
English
Region
England
NLM ID
0411011
PMCID
PMC290248
Subset
IM
Grants
NHGRI NIH HHS · R01 HG000783 · United States
NHGRI NIH HHS · HG00783 · United States
Databases
GENBANK
AF083975
RefSeq
NC_000899, NC_001824
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com