Home LiteratureArticle Details
PMID: 15060014 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S.

Aligning multiple genomic sequences with the threaded blockset aligner.

Genome research ·Vol. 14 ·No. 4 ·2004-04-00 ·Pages 708-15

Blanchette M, Kent WJ, Riemer C, Elnitski L, Smit AF, Roskin KM, Baertsch R, Rosenbloom K, Clawson H, Green ED, Haussler D, Miller W

Abstract

We define a "threaded blockset," which is a novel generalization of the classic notion of a multiple alignment. A new computer program called TBA (for "threaded blockset aligner") builds a threaded blockset under the assumption that all matching segments occur in the same order and orientation in the given sequences; inversions and duplications are not addressed. TBA is designed to be appropriate for aligning many, but by no means all, megabase-sized regions of multiple mammalian genomes. The output of TBA can be projected onto any genome chosen as a reference, thus guaranteeing that different projections present consistent predictions of which genomic positions are orthologous. This capability is illustrated using a new visualization tool to view TBA-generated alignments of vertebrate Hox clusters from both the mammalian and fish perspectives. Experimental evaluation of alignment quality, using a program that simulates evolutionary change in genomic sequences, indicates that TBA is more accurate than earlier programs. To perform the dynamic-programming alignment step, TBA runs a stand-alone program called MULTIZ, which can be used to align highly rearranged or incompletely sequenced genomes. We describe our use of MULTIZ to produce the whole-genome multiple alignments at the Santa Cruz Genome Browser.

MeSH Terms
Animals Base Sequence Cats Cattle Computational Biology/methods,standards,trends Computer Simulation Dogs Evaluation Studies as Topic Evolution, Molecular Genes, Homeobox/genetics Genes, fos/genetics Genome Genome, Human Humans Mice Molecular Sequence Data Multigene Family/genetics Rats Ribosomal Proteins/genetics Sequence Alignment/methods,standards,trends Software/trends
Chemicals
Ribosomal Proteins ribosomal protein L34
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Blanchette Mathieu
Howard Hughes Medical Institute, University of California at Santa Cruz, Santa Cruz, California 95064, USA.
Kent W James
Riemer Cathy
Elnitski Laura
Smit Arian F A
Roskin Krishna M
Baertsch Robert
Rosenbloom Kate
Clawson Hiram
Green Eric D
Haussler David
Miller Webb
References (25)
25 references, click to expand
  1. Approximate matching of network expressions with spacers.
    J Comput Biol. 1996 Spring;3(1):33-51 PMID: 8697238
  2. MAVID multiple alignment server.
    Nucleic Acids Res. 2003 Jul 1;31(13):3525-6 PMID: 12824358
  3. Comparative analyses of multi-species sequences from targeted genomic regions.
    Nature. 2003 Aug 14;424(6950):788-93 PMID: 12917688
  4. Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes.
    Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9 PMID: 14500911
  5. Identification and characterization of multi-species conserved sequences.
    Genome Res. 2003 Dec;13(12):2507-18 PMID: 14656959
  6. Phylogenetic estimation of context-dependent substitution rates by maximum likelihood.
    Mol Biol Evol. 2004 Mar;21(3):468-88 PMID: 14660683
  7. Genome sequence of the Brown Norway rat yields insights into mammalian evolution.
    Nature. 2004 Apr 1;428(6982):493-521 PMID: 15057822
  8. Benchmarking tools for the alignment of functional noncoding DNA.
    BMC Bioinformatics. 2004 Jan 21;5:6 PMID: 14736341
  9. A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given.
    Mol Biol Evol. 1989 Nov;6(6):649-68 PMID: 2488477
  10. Globin gene server: a prototype E-mail database server featuring extensive multiple alignments and data compilation for electronic genetic analysis.
    Genomics. 1994 May 15;21(2):344-53 PMID: 8088828
  11. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  12. Fast and sensitive alignment of large genomic sequences.
    Proc IEEE Comput Soc Bioinform Conf. 2002;1:138-47 PMID: 15838131
  13. PipMaker--a web server for aligning two genomic DNA sequences.
    Genome Res. 2000 Apr;10(4):577-86 PMID: 10779500
  14. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  15. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  16. Multiple sequence alignment using partial order graphs.
    Bioinformatics. 2002 Mar;18(3):452-64 PMID: 11934745
  17. The human genome browser at UCSC.
    Genome Res. 2002 Jun;12(6):996-1006 PMID: 12045153
  18. Whole-genome shotgun assembly and analysis of the genome of Fugu rubripes.
    Science. 2002 Aug 23;297(5585):1301-10 PMID: 12142439
  19. Initial sequencing and comparative analysis of the mouse genome.
    Nature. 2002 Dec 5;420(6915):520-62 PMID: 12466850
  20. Human-mouse alignments with BLASTZ.
    Genome Res. 2003 Jan;13(1):103-7 PMID: 12529312
  21. LAGAN and Multi-LAGAN: efficient tools for large-scale multiple alignment of genomic DNA.
    Genome Res. 2003 Apr;13(4):721-31 PMID: 12654723
  22. A vision for the future of genomics research.
    Nature. 2003 Apr 24;422(6934):835-47 PMID: 12695777
  23. Evolutionary conservation of regulatory elements in vertebrate Hox gene clusters.
    Genome Res. 2003 Jun;13(6A):1111-22 PMID: 12799348
  24. MultiPipMaker and supporting tools: Alignments and analysis of multiple genomic DNA sequences.
    Nucleic Acids Res. 2003 Jul 1;31(13):3518-24 PMID: 12824357
  25. DIALIGN 2: improvement of the segment-to-segment approach to multiple sequence alignment.
    Bioinformatics. 1999 Mar;15(3):211-8 PMID: 10222408
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2004-04-00
Pages
708-15
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC383317
Subset
IM
Grants
NHGRI NIH HHS · HG-02238 · United States
NHGRI NIH HHS · P41 HG002371 · United States
NHGRI NIH HHS · 1P41HG02371 · United States
NHGRI NIH HHS · F32 HG002325 · United States
NHGRI NIH HHS · HG02325 · United States
NHGRI NIH HHS · R01 HG002238 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com