Home LiteratureArticle Details
PMID: 8628656 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Cleaning the GenBank Arabidopsis thaliana data set.

Nucleic acids research ·Vol. 24 ·No. 2 ·1996-01-15 ·Pages 316-20

Korning PG, Hebsgaard SM, Rouze P, Brunak S

Abstract

Data driven computational biology relies on the large quantities of genomic data stored in international sequence data banks. However, the possibilities are drastically impaired if the stored data is unreliable. During a project aiming to predict splice sites in the dicot Arabidopsis thaliana, we extracted a data set from the A.thaliana entries in GenBank. A number of simple 'sanity' checks, based on the nature of the data, revealed an alarmingly high error rate. More than 15% of the most important entries extracted did contain erroneous information. In addition, a number of entries had directly conflicting assignments of exons and introns, not stemming from alternative splicing. In a few cases the errors are due to mere typographical misprints, which may be corrected by comparison to the original papers, but errors caused by wrong assignments of splice sites from experimental data are the most common. It is proposed that the level of error correction should be increased and that gene structure sanity checks should be incorporated--also at the submitter level--to avoid or reduce the problem in the future. A non-redundant and error corrected subset of the data for A.thaliana is made available through anonymous FTP.

MeSH Terms
Algorithms Arabidopsis/genetics Base Sequence DNA, Plant/genetics Databases, Factual Genome, Plant Introns Molecular Sequence Data Neural Networks, Computer RNA Splicing/genetics
Chemicals
DNA, Plant
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Korning P G
Center for Biological Sequence Analysis, Technical University of Denmark, Lyngby, Denmark.
Hebsgaard S M
Rouze P
Brunak S
References (9)
9 references, click to expand
  1. Neural network detects errors in the assignment of mRNA splice sites.
    Nucleic Acids Res. 1990 Aug 25;18(16):4797-801 PMID: 2395643
  2. Cleaning up gene databases.
    Nature. 1990 Jan 11;343(6254):123 PMID: 2296305
  3. Arabidopsis phosphoribosylanthranilate isomerase: molecular genetic analysis of triplicate tryptophan pathway genes.
    Plant Cell. 1995 Apr;7(4):447-61 PMID: 7773017
  4. Nucleotide and protein sequences of a cytoplasmic ribosomal protein S15a gene from Arabidopsis thaliana.
    Plant Physiol. 1994 Sep;106(1):401-2 PMID: 7972526
  5. The unusual 5' splicing border GC is used in myrosinase genes of the Brassicaceae.
    Plant Mol Biol. 1995 Oct;29(1):167-71 PMID: 7579162
  6. The minimum functional length of pre-mRNA introns in monocots and dicots.
    Plant Mol Biol. 1990 May;14(5):727-33 PMID: 2102851
  7. Prediction of human mRNA donor and acceptor sites from the DNA sequence.
    J Mol Biol. 1991 Jul 5;220(1):49-65 PMID: 2067018
  8. Database of homology-derived protein structures and the structural meaning of sequence alignment.
    Proteins. 1991;9(1):56-68 PMID: 2017436
  9. Selection of representative protein data sets.
    Protein Sci. 1992 Mar;1(3):409-17 PMID: 1304348
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
0305-1048
Published
1996-01-15
Pages
316-20
Language
English
Region
England
NLM ID
0411011
PMCID
PMC145627
Subset
IM
Databases
GENBANK
D13044, L22568, L24119, L27461, M12196, M37247, M84343, M85253, M93023, U08315, U09339, U11033, U18969, X13708, X51474, X51799, X55970, X62281, X68146, X69376, X70990, X73839, X74515, X74734, X81799, X81800, Z12614, Z22958, Z31589, Z31715
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com