Home LiteratureArticle Details
PMID: 15520295 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

A novel method for multiple alignment of sequences with repeated and shuffled elements.

Genome research ·Vol. 14 ·No. 11 ·2004-11-00 ·Pages 2336-46

Raphael B, Zhi D, Tang H, Pevzner P

Abstract

We describe ABA (A-Bruijn alignment), a new method for multiple alignment of biological sequences. The major difference between ABA and existing multiple alignment methods is that ABA represents an alignment as a directed graph, possibly containing cycles. This representation provides more flexibility than does a traditional alignment matrix or the recently introduced partial order alignment (POA) graph by allowing a larger class of evolutionary relationships between the aligned sequences. Our graph representation is particularly well-suited to the alignment of protein sequences with shuffled and/or repeated domain structure, and allows one to construct multiple alignments of proteins containing (1) domains that are not present in all proteins, (2) domains that are present in different orders in different proteins, and (3) domains that are present in multiple copies in some proteins. In addition, ABA is useful in the alignment of genomic sequences that contain duplications and inversions. We provide several examples illustrating the applications of ABA.

MeSH Terms
Algorithms Databases, Genetic Sequence Alignment/methods Sequence Analysis, Protein/methods Sequence Homology, Amino Acid Software
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Raphael Benjamin
Department of Computer Science and Engineering, University of California, San Diego, La Jolla, California 92093-0114, USA. braphael@ucsd.edu
Zhi Degui
Tang Haixu
Pevzner Pavel
References (39)
39 references, click to expand
  1. 1-Tuple DNA sequencing: computer analysis.
    J Biomol Struct Dyn. 1989 Aug;7(1):63-73 PMID: 2684223
  2. An Eulerian path approach to global multiple alignment for DNA sequences.
    J Comput Biol. 2003;10(6):803-19 PMID: 14980012
  3. A workbench for multiple alignment construction and analysis.
    Proteins. 1991;9(3):180-90 PMID: 2006136
  4. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  5. The multiplicity of domains in proteins.
    Annu Rev Biochem. 1995;64:287-314 PMID: 7574483
  6. A new algorithm for DNA sequence assembly.
    J Comput Biol. 1995 Summer;2(2):291-306 PMID: 7497130
  7. Approximate matching of network expressions with spacers.
    J Comput Biol. 1996 Spring;3(1):33-51 PMID: 8697238
  8. On the complexity of multiple sequence alignment.
    J Comput Biol. 1994 Winter;1(4):337-48 PMID: 8790475
  9. Multiple DNA and protein sequence alignment based on segment-to-segment comparison.
    Proc Natl Acad Sci U S A. 1996 Oct 29;93(22):12098-103 PMID: 8901539
  10. Extracting protein alignment models from the sequence database.
    Nucleic Acids Res. 1997 May 1;25(9):1665-77 PMID: 9108146
  11. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  12. DIALIGN: finding local similarities by multiple sequence alignment.
    Bioinformatics. 1998;14(3):290-4 PMID: 9614273
  13. Splicing graphs and EST assembly problem.
    Bioinformatics. 2002;18 Suppl 1:S181-8 PMID: 12169546
  14. A computational method for resequencing long DNA targets by universal oligonucleotide arrays.
    Proc Natl Acad Sci U S A. 2002 Nov 26;99(24):15492-6 PMID: 12429861
  15. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  16. Sources of systematic error in functional annotation of genomes: domain rearrangement, non-orthologous gene displacement and operon disruption.
    In Silico Biol. 1998;1(1):55-67 PMID: 11471243
  17. Scale-free behavior in protein domain networks.
    Mol Biol Evol. 2001 Sep;18(9):1694-702 PMID: 11504849
  18. An Eulerian path approach to DNA fragment assembly.
    Proc Natl Acad Sci U S A. 2001 Aug 14;98(17):9748-53 PMID: 11504945
  19. Multiple sequence alignment using partial order graphs.
    Bioinformatics. 2002 Mar;18(3):452-64 PMID: 11934745
  20. Recent progress in multiple sequence alignment: a survey.
    Pharmacogenomics. 2002 Jan;3(1):131-44 PMID: 11966409
  21. Large scale sequencing by hybridization.
    J Comput Biol. 2002;9(2):413-28 PMID: 12015890
  22. The human genome browser at UCSC.
    Genome Res. 2002 Jun;12(6):996-1006 PMID: 12045153
  23. Comparative analysis of protein domain organization.
    Genome Res. 2004 Mar;14(3):343-53 PMID: 14993202
  24. Aligning multiple genomic sequences with the threaded blockset aligner.
    Genome Res. 2004 Apr;14(4):708-15 PMID: 15060014
  25. Combining partial order alignment and progressive multiple sequence alignment increases alignment speed and scalability to very large alignment problems.
    Bioinformatics. 2004 Jul 10;20(10):1546-56 PMID: 14962922
  26. Mauve: multiple alignment of conserved genomic sequence with rearrangements.
    Genome Res. 2004 Jul;14(7):1394-403 PMID: 15231754
  27. Progressive sequence alignment as a prerequisite to correct phylogenetic trees.
    J Mol Evol. 1987;25(4):351-60 PMID: 3118049
  28. CLUSTAL: a package for performing multiple sequence alignment on a microcomputer.
    Gene. 1988 Dec 15;73(1):237-44 PMID: 3243435
  29. A tool for multiple sequence alignment.
    Proc Natl Acad Sci U S A. 1989 Jun;86(12):4412-5 PMID: 2734293
  30. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  31. CDD: a curated Entrez database of conserved domain alignments.
    Nucleic Acids Res. 2003 Jan 1;31(1):383-7 PMID: 12520028
  32. Human-mouse alignments with BLASTZ.
    Genome Res. 2003 Jan;13(1):103-7 PMID: 12529312
  33. PCMA: fast and accurate multiple sequence alignment based on profile consistency.
    Bioinformatics. 2003 Feb 12;19(3):427-8 PMID: 12584134
  34. MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.
    Nucleic Acids Res. 2003 Jul 1;31(13):3518-24 PMID: 12824357
  35. Estimating the repeat structure and length of DNA sequences using L-tuples.
    Genome Res. 2003 Aug;13(8):1916-22 PMID: 12902383
  36. Comparative analyses of multi-species sequences from targeted genomic regions.
    Nature. 2003 Aug 14;424(6950):788-93 PMID: 12917688
  37. Divide-and-conquer multiple alignment with segment-based constraints.
    Bioinformatics. 2003 Oct;19 Suppl 2:ii189-95 PMID: 14534189
  38. The Pfam protein families database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41 PMID: 14681378
  39. Motif recognition and alignment for many sequences by comparison of dot-matrices.
    J Mol Biol. 1991 Mar 5;218(1):33-43 PMID: 1900535
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2004-11-00
Pages
2336-46
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC525693
Subset
IM
Grants
NHGRI NIH HHS · R01 HG002366 · United States
NHGRI NIH HHS · 1 R01 HG02366 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com