Home LiteratureArticle Details
PMID: 11934736 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

Separation of nearly identical repeats in shotgun assemblies using defined nucleotide positions, DNPs.

Bioinformatics (Oxford, England) ·Vol. 18 ·No. 3 ·2002-03-00 ·Pages 379-88

Tammi MT, Arner E, Britton T, Andersson B

Abstract

An increasingly important problem in genome sequencing is the failure of the commonly used shotgun assembly programs to correctly assemble repetitive sequences. The assembly of non-repetitive regions or regions containing repeats considerably shorter than the average read length is in practice easy to solve, while longer repeats have been a difficult problem. We here present a statistical method to separate arbitrarily long, almost identical repeats, which makes it possible to correctly assemble complex repetitive sequence regions. The differences between repeat units may be as low as 1% and the sequencing error may be up to ten times higher. The method is based on the realization that a comparison of only a part of all overlapping sequences at a time in a data set does not generate enough information for a conclusive analysis. Our method uses optimal multi-alignments consisting of all the overlaps of each read. This makes it possible to determine defined nucleotide positions, DNPs, which constitute the differences between the repeat units. Differences between repeats are distinguished from sequencing errors using statistical methods, where the probabilities of obtaining certain combinations of candidate DNPs are calculated using the information from the multi-alignments. The use of DNPs and combinations of DNPs will allow for optimal and rapid assemblies of repeated regions. This method can solve repeats that differ in only two positions in a read length, which is the theoretical limit for repeat separation. We predict that this method will be highly useful in shotgun sequencing in the future.

MeSH Terms
Algorithms Base Sequence Cluster Analysis Computational Biology/methods Computer Simulation Deoxyribonucleoproteins/genetics Feasibility Studies Models, Genetic Models, Statistical Molecular Sequence Data Repetitive Sequences, Nucleic Acid/genetics Sensitivity and Specificity Sequence Alignment/methods,statistics & numerical data Sequence Analysis, DNA/methods,statistics & numerical data
Chemicals
Deoxyribonucleoproteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Tammi Martti T
Department of Genetics and Pathology, Rudbeck Laboratory, Uppsala University, Uppsala, Sweden.
Arner Erik
Britton Tom
Andersson Björn
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2002-03-00
Pages
379-88
Language
English
Region
England
NLM ID
9808944
Subset
IM
Grants
NIAID NIH HHS · 5 U01 AI 45061-02 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com