Home LiteratureArticle Details
PMID: 21092284 Published · epublish English Journal Article Research Support, N.I.H., Extramural

Pash 3.0: A versatile software package for read mapping and integrative analysis of genomic and epigenomic variation using massively parallel DNA sequencing.

BMC bioinformatics ·Vol. 11 ·2010-11-23 ·Pages 572

Coarfa C, Yu F, Miller CA, Chen Z, Harris RA, Milosavljevic A

Abstract

Massively parallel sequencing readouts of epigenomic assays are enabling integrative genome-wide analyses of genomic and epigenomic variation. Pash 3.0 performs sequence comparison and read mapping and can be employed as a module within diverse configurable analysis pipelines, including ChIP-Seq and methylome mapping by whole-genome bisulfite sequencing. Pash 3.0 generally matches the accuracy and speed of niche programs for fast mapping of short reads, and exceeds their performance on longer reads generated by a new generation of massively parallel sequencing technologies. By exploiting longer read lengths, Pash 3.0 maps reads onto the large fraction of genomic DNA that contains repetitive elements and polymorphic sites, including indel polymorphisms. We demonstrate the versatility of Pash 3.0 by analyzing the interaction between CpG methylation, CpG SNPs, and imprinting based on publicly available whole-genome shotgun bisulfite sequencing data. Pash 3.0 makes use of gapped k-mer alignment, a non-seed based comparison method, which is implemented using multi-positional hash tables. This allows Pash 3.0 to run on diverse hardware platforms, including individual computers with standard RAM capacity, multi-core hardware architectures and large clusters.

MeSH Terms
Base Sequence DNA/chemistry Databases, Genetic Epigenomics/methods Genetic Variation Genome Polymorphism, Single Nucleotide Sequence Analysis, DNA/methods Software
Chemicals
DNA
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Coarfa Cristian
Department of Molecular and Human Genetics, Baylor College of Medicine, One Baylor Plaza Houston, TX 77030, USA. coarfa@bcm.edu
Yu Fuli
Miller Christopher A
Chen Zuozhou
Harris R Alan
Milosavljevic Aleksandar
References (35)
35 references, click to expand
  1. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  2. Using quality scores and longer reads improves accuracy of Solexa read mapping.
    BMC Bioinformatics. 2008 Feb 28;9:128 PMID: 18307793
  3. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
    Nature. 2007 Aug 2;448(7153):553-60 PMID: 17603471
  4. Mapping allele-specific DNA methylation: a new tool for maximizing information from GWAS.
    Am J Hum Genet. 2010 Feb 12;86(2):109-12 PMID: 20159108
  5. Whole-genome re-sequencing.
    Curr Opin Genet Dev. 2006 Dec;16(6):545-52 PMID: 17055251
  6. BRAT: bisulfite-treated reads analysis tool.
    Bioinformatics. 2010 Feb 15;26(4):572-3 PMID: 20031974
  7. Putting epigenome comparison into practice.
    Nat Biotechnol. 2010 Oct;28(10):1053-6 PMID: 20944597
  8. Circular binary segmentation for the analysis of array-based DNA copy number data.
    Biostatistics. 2004 Oct;5(4):557-72 PMID: 15475419
  9. Updates to the RMAP short-read mapping software.
    Bioinformatics. 2009 Nov 1;25(21):2841-2 PMID: 19736251
  10. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
  11. Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
    Genome Res. 1998 Mar;8(3):175-85 PMID: 9521921
  12. SSAHA: a fast search method for large DNA databases.
    Genome Res. 2001 Oct;11(10):1725-9 PMID: 11591649
  13. Pash: efficient genome-scale sequence anchoring by Positional Hashing.
    Genome Res. 2004 Apr;14(4):672-8 PMID: 15060009
  14. Good spaced seeds for homology search.
    Bioinformatics. 2004 May 1;20(7):1053-9 PMID: 14764573
  15. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  16. BSMAP: whole genome bisulfite sequence MAPping program.
    BMC Bioinformatics. 2009 Jul 27;10:232 PMID: 19635165
  17. Allelic skewing of DNA methylation is widespread across the genome.
    Am J Hum Genet. 2010 Feb 12;86(2):196-212 PMID: 20159110
  18. mrsFAST: a cache-oblivious algorithm for short-read mapping.
    Nat Methods. 2010 Aug;7(8):576-7 PMID: 20676076
  19. Origins and functional impact of copy number variation in the human genome.
    Nature. 2010 Apr 1;464(7289):704-12 PMID: 19812545
  20. dbSNP: the NCBI database of genetic variation.
    Nucleic Acids Res. 2001 Jan 1;29(1):308-11 PMID: 11125122
  21. Fast and accurate long-read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2010 Mar 1;26(5):589-95 PMID: 20080505
  22. A survey of sequence alignment algorithms for next-generation sequencing.
    Brief Bioinform. 2010 Sep;11(5):473-83 PMID: 20460430
  23. Quantitative comparison of genome-wide DNA methylation mapping technologies.
    Nat Biotechnol. 2010 Oct;28(10):1106-14 PMID: 20852634
  24. Comparison of sequencing-based methods to profile DNA methylation and identification of monoallelic epigenetic modifications.
    Nat Biotechnol. 2010 Oct;28(10):1097-105 PMID: 20852635
  25. Human DNA methylomes at base resolution show widespread epigenomic differences.
    Nature. 2009 Nov 19;462(7271):315-22 PMID: 19829295
  26. The many faces of sequence alignment.
    Brief Bioinform. 2005 Mar;6(1):6-22 PMID: 15826353
  27. Pash 2.0: scaleable sequence anchoring for next-generation sequencing technologies.
    Pac Symp Biocomput. 2008;:102-13 PMID: 18229679
  28. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
  29. Massive parallel bisulfite sequencing of CG-rich DNA fragments reveals that methylation of many X-chromosomal CpG islands in female blood DNA is incomplete.
    Hum Mol Genet. 2009 Apr 15;18(8):1439-48 PMID: 19223391
  30. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  31. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  32. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  33. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  34. Using the FASTA program to search protein and DNA sequence databases.
    Methods Mol Biol. 1994;24:307-31 PMID: 8205202
  35. Detection of large-scale variation in the human genome.
    Nat Genet. 2004 Sep;36(9):949-51 PMID: 15286789
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-11-23
Epub
2010-00-23
Pages
572
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3001746
Subset
IM
Grants
NHGRI NIH HHS · 5R01HG004009 · United States
NIDA NIH HHS · 5U01DA025956 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com