Abstract
The TIGR Plant Transcript Assemblies (TA) database (http://plantta.tigr.org) uses expressed sequences collected from the NCBI GenBank Nucleotide database for the construction of transcript assemblies. The sequences collected include expressed sequence tags (ESTs) and full-length and partial cDNAs, but exclude computationally predicted gene sequences. The TA database includes all plant species for which more than 1000 EST or cDNA sequences are publicly available. The EST and cDNA sequences are first clustered based on an all-versus-all pairwise sequence comparison, followed by the generation of consensus sequences (TAs) from individual clusters. The clustering and assembly procedures use the TGICL tool, Megablast and the CAP3 assembler. The UniProt Reference Clusters (UniRef100) protein database is used as the reference database for the functional annotation of the assemblies. The transcription orientation of each TA is determined based on the orientation of the alignment with the best protein hit. The TA sequences and annotation are available via web interfaces and FTP downloads. Assemblies can be retrieved by a text-based keyword search or a sequence-based BLAST search. The current version of the TA database is Release 2 (July 17, 2006) and includes a total of 215 plant species.
MeSH Terms
DNA, Complementary/chemistry
Databases, Nucleic Acid
Databases, Protein
Expressed Sequence Tags/chemistry
Internet
Plant Proteins/genetics
RNA, Messenger/chemistry
RNA, Plant/chemistry
User-Computer Interface
Chemicals
DNA, Complementary
Plant Proteins
RNA, Messenger
RNA, Plant
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Childs Kevin L
The Institute for Genomic Research, 9712 Medical Center Drive, Rockville, MD 20850, USA.
Hamilton John P
Zhu Wei
Ly Eugene
Cheung Foo
Wu Hank
Rabinowicz Pablo D
Town Chris D
Buell C Robin
Chan Agnes P
References (13)
13 references, click to expand
-
InterPro: an integrated documentation resource for protein families, domains and functional sites.
Brief Bioinform. 2002 Sep;3(3):225-35
PMID: 12230031
-
Database resources of the National Center for Biotechnology.
Nucleic Acids Res. 2003 Jan 1;31(1):28-33
PMID: 12519941
-
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets.
Bioinformatics. 2003 Mar 22;19(5):651-2
PMID: 12651724
-
PlantGDB, plant genome database and analysis tools.
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D354-9
PMID: 14681433
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
CAP3: A DNA sequence assembly program.
Genome Res. 1999 Sep;9(9):868-77
PMID: 10508846
-
The Universal Protein Resource (UniProt).
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D154-9
PMID: 15608167
-
The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D71-4
PMID: 15608288
-
Comparative plant genomics resources at PlantGDB.
Plant Physiol. 2005 Oct;139(2):610-8
PMID: 16219921
-
The Gene Ontology (GO) project in 2006.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D322-6
PMID: 16381878
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
A greedy algorithm for aligning DNA sequences.
J Comput Biol. 2000 Feb-Apr;7(1-2):203-14
PMID: 10890397
-
The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.
Nucleic Acids Res. 2001 Jan 1;29(1):159-64
PMID: 11125077