Home LiteratureArticle Details
PMID: 11827946 Published · ppublish English Comparative Study Journal Article Research Support, U.S. Gov't, P.H.S.

Molecular fossils in the human genome: identification and analysis of the pseudogenes in chromosomes 21 and 22.

Genome research ·Vol. 12 ·No. 2 ·2002-02-00 ·Pages 272-80

Harrison PM, Hegyi H, Balasubramanian S, Luscombe NM, Bertone P, Echols N, Johnson T, Gerstein M

Abstract

We have developed an initial approach for annotating and surveying pseudogenes in the human genome. We search human genomic DNA for regions that are similar to known protein sequences and contain obvious disablements (i.e., mid-sequence stop codons or frameshifts), while ensuring minimal overlap with annotations of known genes. Pseudogenes can be divided into "processed" and "nonprocessed"; the former are reverse transcribed from mRNA (and therefore have no intron structure), whereas the latter presumably arise from genomic duplications. We annotate putative processed pseudogenes based on whether there is a continuous span of homology that is >70% of the length of the closest matching human protein (i.e., with introns removed), or whether there is evidence of polyadenylation. We have applied our approach to chromosomes 21 and 22, the first parts of the human genome completely sequenced, finding 190 new pseudogene annotations beyond the 264 reported by the sequencing centers. In total, on chromosomes 21 and 22, there are 189 processed pseudogenes, 195 nonprocessed pseudogenes, and, additionally, 70 pseudogenic immunoglobulin gene segments. (Detailed assignments are available at http://bioinfo.mbb.yale.edu/genome/pseudogene or http://genecensus.org/pseudogene.) By extrapolation, we predict that there could be up to approximately 20,000 pseudogenes in the whole human genome, with a little more than half of them processed. We have determined the main populations and clusters of pseudogenes on chromosomes 21 and 22. There are notable excesses of pseudogenes relative to genes near the centromeres of both chromosomes, indicating the existence of pseudogenic "hot-spots" in the genome. We have looked at the distribution of InterPro families and Gene Ontology (GO) functional categories in our pseudogenes. Overall, the families in both processed and nonprocessed pseudogene populations occur according to a similar power-law distribution as that found for the occurrence of gene families, with a few big families and many small ones. The processed population is, in particular, enriched in highly expressed ribosomal-protein sequences (approximately 20%), which appear fairly evenly distributed across the chromosomes. We compared processed pseudogenes of different evolutionary ages, observing a high degree of similarity between "ancient" and "modern" subpopulations. This may be attributable to the consistently high expression of ribosomal proteins over evolutionary time. Finally, we find that chromosome 22 pseudogene population is dominated by immunoglobulin segments, which have a greater rate of disablement per amino acid than the other pseudogene populations and are also substantially more diverged.

MeSH Terms
Chromosome Mapping/methods Chromosomes, Human, Pair 21/genetics Chromosomes, Human, Pair 22/genetics Evolution, Molecular Fossils Genes, Immunoglobulin Genes, Overlapping Genome, Human Humans Multigene Family Pseudogenes RNA Processing, Post-Transcriptional/genetics Sequence Analysis, DNA/statistics & numerical data
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Harrison Paul M
Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, Connecticut 06520-8114, USA.
Hegyi Hedi
Balasubramanian Suganthi
Luscombe Nicholas M
Bertone Paul
Echols Nathaniel
Johnson Ted
Gerstein Mark
References (32)
32 references, click to expand
  1. Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.
    Nucleic Acids Res. 2001 Feb 1;29(3):818-30 PMID: 11160906
  2. The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
    J Mol Biol. 1999 Apr 23;288(1):147-64 PMID: 10329133
  3. Mining the draft human genome.
    Nature. 2001 Feb 15;409(6822):827-8 PMID: 11236999
  4. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  5. The complete human olfactory subgenome.
    Genome Res. 2001 May;11(5):685-702 PMID: 11337468
  6. Computational inference of homologous gene structures in the human genome.
    Genome Res. 2001 May;11(5):803-16 PMID: 11337476
  7. A complete map of the human ribosomal protein genes: assignment of 80 genes to the cytogenetic map and implications for human disorders.
    Genomics. 2001 Mar 15;72(3):223-30 PMID: 11401437
  8. A draft annotation and overview of the human genome.
    Genome Biol. 2001;2(7):RESEARCH0025 PMID: 11516338
  9. Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.
    J Mol Biol. 2001 Nov 2;313(4):673-81 PMID: 11697896
  10. Processed pseudogenes: characteristics and evolution.
    Annu Rev Genet. 1985;19:253-72 PMID: 3909943
  11. Analysis of compositionally biased regions in sequence databases.
    Methods Enzymol. 1996;266:554-71 PMID: 8743706
  12. Comparative analysis of 1196 orthologous mouse and human full-length mRNA and protein sequences.
    Genome Res. 1996 Sep;6(9):846-57 PMID: 8889551
  13. Evolution by the birth-and-death process in multigene families of the vertebrate immune system.
    Proc Natl Acad Sci U S A. 1997 Jul 22;94(15):7799-806 PMID: 9223266
  14. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  15. Comparison of DNA sequences with protein sequences.
    Genomics. 1997 Nov 15;46(1):24-36 PMID: 9403055
  16. A structural census of genomes: comparing bacterial, eukaryotic, and archaeal genomes in terms of protein structure.
    J Mol Biol. 1997 Dec 12;274(4):562-76 PMID: 9417935
  17. The DNA sequence of human chromosome 22.
    Nature. 1999 Dec 2;402(6761):489-95 PMID: 10591208
  18. GenBank.
    Nucleic Acids Res. 2000 Jan 1;28(1):15-8 PMID: 10592170
  19. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  20. The GeneQuiz web server: protein functional analysis through the Web.
    Trends Biochem Sci. 2000 Jan;25(1):33-5 PMID: 10637611
  21. Vertebrate pseudogenes.
    FEBS Lett. 2000 Feb 25;468(2-3):109-14 PMID: 10692568
  22. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  23. The DNA sequence of human chromosome 21.
    Nature. 2000 May 18;405(6784):311-9 PMID: 10830953
  24. Analysis of expressed sequence tags indicates 35,000 human genes.
    Nat Genet. 2000 Jun;25(2):232-4 PMID: 10835644
  25. Estimate of human gene number provided by genome-wide analysis using Tetraodon nigroviridis DNA sequence.
    Nat Genet. 2000 Jun;25(2):235-8 PMID: 10835645
  26. The sequence of the human genome.
    Science. 2001 Feb 16;291(5507):1304-51 PMID: 11181995
  27. Evolution of genome size: new approaches to an old problem.
    Trends Genet. 2001 Jan;17(1):23-8 PMID: 11163918
  28. MaskerAid: a performance enhancement to RepeatMasker.
    Bioinformatics. 2000 Nov;16(11):1040-1 PMID: 11159316
  29. InterPro--an integrated documentation resource for protein families, domains and functional sites.
    Bioinformatics. 2000 Dec;16(12):1145-50 PMID: 11159333
  30. The genome sequence of Rickettsia prowazekii and the origin of mitochondria.
    Nature. 1998 Nov 12;396(6707):133-40 PMID: 9823893
  31. Patterns of protein-fold usage in eight microbial genomes: a comprehensive structural census.
    Proteins. 1998 Dec 1;33(4):518-34 PMID: 9849936
  32. Massive gene decay in the leprosy bacillus.
    Nature. 2001 Feb 22;409(6823):1007-11 PMID: 11234002
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2002-02-00
Pages
272-80
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC155275
Subset
IM
Grants
NIGMS NIH HHS · P50 GM062413 · United States
NIGMS NIH HHS · P50 GM62413-01 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com