Abstract
We have developed a rice (Oryza sativa) genome annotation database (Osa1) that provides structural and functional annotation for this emerging model species. Using the sequence of O. sativa subsp. japonica cv Nipponbare from the International Rice Genome Sequencing Project, pseudomolecules, or virtual contigs, of the 12 rice chromosomes were constructed. Our most recent release, version 3, represents our third build of the pseudomolecules and is composed of 98% finished sequence. Genes were identified using a series of computational methods developed for Arabidopsis (Arabidopsis thaliana) that were modified for use with the rice genome. In release 3 of our annotation, we identified 57,915 genes, of which 14,196 are related to transposable elements. Of these 43,719 non-transposable element-related genes, 18,545 (42.4%) were annotated with a putative function, 5,777 (13.2%) were annotated as encoding an expressed protein with no known function, and the remaining 19,397 (44.4%) were annotated as encoding a hypothetical protein. Multiple splice forms (5,873) were detected for 2,538 genes, resulting in a total of 61,250 gene models in the rice genome. We incorporated experimental evidence into 18,252 gene models to improve the quality of the structural annotation. A series of functional data types has been annotated for the rice genome that includes alignment with genetic markers, assignment of gene ontologies, identification of flanking sequence tags, alignment with homologs from related species, and syntenic mapping with other cereal species. All structural and functional annotation data are available through interactive search and display windows as well as through download of flat files. To integrate the data with other genome projects, the annotation data are available through a Distributed Annotation System and a Genome Browser. All data can be obtained through the project Web pages at http://rice.tigr.org.
MeSH Terms
Computational Biology
Databases, Genetic
Genes, Plant
Genome, Plant
Oryza/genetics
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Yuan Qiaoping
The Institute for Genomic Research, Rockville, Maryland 20850, USA.
Ouyang Shu
Wang Aihui
Zhu Wei
Maiti Rama
Lin Haining
Hamilton John
Haas Brian
Sultana Razvan
Cheung Foo
Wortman Jennifer
Buell C Robin
References (36)
36 references, click to expand
-
The TIGR rice genome annotation resource: annotating the rice genome and creating resources for plant biologists.
Nucleic Acids Res. 2003 Jan 1;31(1):229-33
PMID: 12519988
-
The InterPro Database, 2003 brings increased coverage and new features.
Nucleic Acids Res. 2003 Jan 1;31(1):315-8
PMID: 12520011
-
The generic genome browser: a building block for a model organism system database.
Genome Res. 2002 Oct;12(10):1599-610
PMID: 12368253
-
Annotation of the Arabidopsis genome.
Plant Physiol. 2003 Jun;132(2):461-8
PMID: 12805579
-
Predicting transmembrane protein topology with a hidden Markov model: application to complete genomes.
J Mol Biol. 2001 Jan 19;305(3):567-80
PMID: 11152613
-
High throughput T-DNA insertion mutagenesis in rice: a first step towards in silico reverse genetics.
Plant J. 2004 Aug;39(3):450-64
PMID: 15255873
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
Pack-MULE transposable elements mediate gene evolution in plants.
Nature. 2004 Sep 30;431(7008):569-73
PMID: 15457261
-
Early and multiple Ac transpositions in rice suitable for efficient insertional mutagenesis.
Plant Mol Biol. 2001 May;46(2):215-27
PMID: 11442061
-
A draft sequence of the rice genome (Oryza sativa L. ssp. japonica).
Science. 2002 Apr 5;296(5565):92-100
PMID: 11935018
-
Saturated molecular map of the rice genome based on an interspecific backcross population.
Genetics. 1994 Dec;138(4):1251-74
PMID: 7896104
-
A chromosome bin map of 16,000 expressed sequence tag loci and distribution of genes among the three genomes of polyploid wheat.
Genetics. 2004 Oct;168(2):701-12
PMID: 15514046
-
tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence.
Nucleic Acids Res. 1997 Mar 1;25(5):955-64
PMID: 9023104
-
Comparative genetics in the grasses.
Proc Natl Acad Sci U S A. 1998 Mar 3;95(5):1971-4
PMID: 9482816
-
A draft sequence of the rice genome (Oryza sativa L. ssp. indica).
Science. 2002 Apr 5;296(5565):79-92
PMID: 11935017
-
Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies.
Nucleic Acids Res. 2003 Oct 1;31(19):5654-66
PMID: 14500829
-
Collection, mapping, and annotation of over 28,000 cDNA clones from japonica rice.
Science. 2003 Jul 18;301(5631):376-9
PMID: 12869764
-
Target site specificity of the Tos17 retrotransposon shows a preference for insertion within genes and against insertion in retrotransposon-rich regions of the genome.
Plant Cell. 2003 Aug;15(8):1771-80
PMID: 12897251
-
A high-density rice genetic linkage map with 2275 markers using a single F2 population.
Genetics. 1998 Jan;148(1):479-94
PMID: 9475757
-
InterProScan--an integration platform for the signature-recognition methods in InterPro.
Bioinformatics. 2001 Sep;17(9):847-8
PMID: 11590104
-
The use of the Monsanto draft rice genome sequence in research.
Plant Physiol. 2001 Mar;125(3):1164-5
PMID: 11244095
-
The distributed annotation system.
BMC Bioinformatics. 2001;2:7
PMID: 11667947
-
A tool for analyzing and annotating genomic sequences.
Genomics. 1997 Nov 15;46(1):37-45
PMID: 9403056
-
Prediction of complete gene structures in human genomic DNA.
J Mol Biol. 1997 Apr 25;268(1):78-94
PMID: 9149143
-
Identification of prokaryotic and eukaryotic signal peptides and prediction of their cleavage sites.
Protein Eng. 1997 Jan;10(1):1-6
PMID: 9051728
-
Transposable element annotation of the rice genome.
Bioinformatics. 2004 Jan 22;20(2):155-60
PMID: 14734305
-
GeneSplicer: a new computational method for splice site prediction.
Nucleic Acids Res. 2001 Mar 1;29(5):1185-90
PMID: 11222768
-
A comprehensive rice transcript map containing 6591 expressed sequence tag sites.
Plant Cell. 2002 Mar;14(3):525-35
PMID: 11910001
-
GeneMark.hmm: new solutions for gene finding.
Nucleic Acids Res. 1998 Feb 15;26(4):1107-15
PMID: 9461475
-
International Rice Genome Sequencing Project: the effort to completely sequence the rice genome.
Curr Opin Plant Biol. 2000 Apr;3(2):138-41
PMID: 10712951
-
The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.
Nucleic Acids Res. 2001 Jan 1;29(1):159-64
PMID: 11125077
-
Ab initio gene finding in Drosophila genomic DNA.
Genome Res. 2000 Apr;10(4):516-22
PMID: 10779491
-
The Pfam protein families database.
Nucleic Acids Res. 2002 Jan 1;30(1):276-80
PMID: 11752314
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Rapid, large-scale generation of Ds transposant lines and analysis of the Ds insertion sites in rice.
Plant J. 2004 Jul;39(2):252-63
PMID: 15225289
-
The TIGR Plant Repeat Databases: a collective resource for the identification of repetitive sequences in plants.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D360-3
PMID: 14681434