Home LiteratureArticle Details
PMID: 19561018 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads.

Bioinformatics (Oxford, England) ·Vol. 25 ·No. 21 ·2009-11-01 ·Pages 2865-71

Ye K, Schulz MH, Long Q, Apweiler R, Ning Z

Abstract

There is a strong demand in the genomic community to develop effective algorithms to reliably identify genomic variants. Indel detection using next-gen data is difficult and identification of long structural variations is extremely challenging. We present Pindel, a pattern growth approach, to detect breakpoints of large deletions and medium-sized insertions from paired-end short reads. We use both simulated reads and real data to demonstrate the efficiency of the computer program and accuracy of the results. The binary code and a short user manual can be freely downloaded from http://www.ebi.ac.uk/ approximately kye/pindel/. k.ye@lumc.nl; zn1@sanger.ac.uk.

MeSH Terms
Algorithms Chromosome Breakpoints Computational Biology/methods DNA Breaks Genome INDEL Mutation Sequence Analysis, DNA Software
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Ye Kai
EMBL Outstation European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, UK. k.ye@lumc.nl
Schulz Marcel H
Long Quan
Apweiler Rolf
Ning Zemin
References (13)
13 references, click to expand
  1. SSAHA: a fast search method for large DNA databases.
    Genome Res. 2001 Oct;11(10):1725-9 PMID: 11591649
  2. Short read fragment assembly of bacterial genomes.
    Genome Res. 2008 Feb;18(2):324-30 PMID: 18083777
  3. The complete genome of an individual by massively parallel DNA sequencing.
    Nature. 2008 Apr 17;452(7189):872-6 PMID: 18421352
  4. An initial map of insertion and deletion (INDEL) variation in the human genome.
    Genome Res. 2006 Sep;16(9):1182-90 PMID: 16902084
  5. The diploid genome sequence of an individual human.
    PLoS Biol. 2007 Sep 4;5(10):e254 PMID: 17803354
  6. Large-scale copy number polymorphism in the human genome.
    Science. 2004 Jul 23;305(5683):525-8 PMID: 15273396
  7. Velvet: algorithms for de novo short read assembly using de Bruijn graphs.
    Genome Res. 2008 May;18(5):821-9 PMID: 18349386
  8. Mapping and sequencing of structural variation from eight human genomes.
    Nature. 2008 May 1;453(7191):56-64 PMID: 18451855
  9. An efficient, versatile and scalable pattern growth approach to mine frequent patterns in unaligned protein sequences.
    Bioinformatics. 2007 Mar 15;23(6):687-93 PMID: 17237070
  10. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  11. The generalised k-Truncated Suffix Tree for time-and space-efficient searches in multiple DNA or protein sequences.
    Int J Bioinform Res Appl. 2008;4(1):81-95 PMID: 18283030
  12. Detection of large-scale variation in the human genome.
    Nat Genet. 2004 Sep;36(9):949-51 PMID: 15286789
  13. Natural genetic variation caused by transposable elements in humans.
    Genetics. 2004 Oct;168(2):933-51 PMID: 15514065
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2009-11-01
Epub
2009-00-26
Pages
2865-71
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2781750
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com