Abstract
We have developed an initial approach for annotating and surveying pseudogenes in the human genome. We search human genomic DNA for regions that are similar to known protein sequences and contain obvious disablements (i.e., mid-sequence stop codons or frameshifts), while ensuring minimal overlap with annotations of known genes. Pseudogenes can be divided into "processed" and "nonprocessed"; the former are reverse transcribed from mRNA (and therefore have no intron structure), whereas the latter presumably arise from genomic duplications. We annotate putative processed pseudogenes based on whether there is a continuous span of homology that is >70% of the length of the closest matching human protein (i.e., with introns removed), or whether there is evidence of polyadenylation. We have applied our approach to chromosomes 21 and 22, the first parts of the human genome completely sequenced, finding 190 new pseudogene annotations beyond the 264 reported by the sequencing centers. In total, on chromosomes 21 and 22, there are 189 processed pseudogenes, 195 nonprocessed pseudogenes, and, additionally, 70 pseudogenic immunoglobulin gene segments. (Detailed assignments are available at http://bioinfo.mbb.yale.edu/genome/pseudogene or http://genecensus.org/pseudogene.) By extrapolation, we predict that there could be up to approximately 20,000 pseudogenes in the whole human genome, with a little more than half of them processed. We have determined the main populations and clusters of pseudogenes on chromosomes 21 and 22. There are notable excesses of pseudogenes relative to genes near the centromeres of both chromosomes, indicating the existence of pseudogenic "hot-spots" in the genome. We have looked at the distribution of InterPro families and Gene Ontology (GO) functional categories in our pseudogenes. Overall, the families in both processed and nonprocessed pseudogene populations occur according to a similar power-law distribution as that found for the occurrence of gene families, with a few big families and many small ones. The processed population is, in particular, enriched in highly expressed ribosomal-protein sequences (approximately 20%), which appear fairly evenly distributed across the chromosomes. We compared processed pseudogenes of different evolutionary ages, observing a high degree of similarity between "ancient" and "modern" subpopulations. This may be attributable to the consistently high expression of ribosomal proteins over evolutionary time. Finally, we find that chromosome 22 pseudogene population is dominated by immunoglobulin segments, which have a greater rate of disablement per amino acid than the other pseudogene populations and are also substantially more diverged.
MeSH Terms
Chromosome Mapping/methods
Chromosomes, Human, Pair 21/genetics
Chromosomes, Human, Pair 22/genetics
Evolution, Molecular
Fossils
Genes, Immunoglobulin
Genes, Overlapping
Genome, Human
Humans
Multigene Family
Pseudogenes
RNA Processing, Post-Transcriptional/genetics
Sequence Analysis, DNA/statistics & numerical data
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Harrison Paul M
Department of Molecular Biophysics and Biochemistry, Yale University, New Haven, Connecticut 06520-8114, USA.
Hegyi Hedi
Balasubramanian Suganthi
Luscombe Nicholas M
Bertone Paul
Echols Nathaniel
Johnson Ted
Gerstein Mark
References (32)
32 references, click to expand
-
Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.
Nucleic Acids Res. 2001 Feb 1;29(3):818-30
PMID: 11160906
-
The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
J Mol Biol. 1999 Apr 23;288(1):147-64
PMID: 10329133
-
Mining the draft human genome.
Nature. 2001 Feb 15;409(6822):827-8
PMID: 11236999
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011
-
The complete human olfactory subgenome.
Genome Res. 2001 May;11(5):685-702
PMID: 11337468
-
Computational inference of homologous gene structures in the human genome.
Genome Res. 2001 May;11(5):803-16
PMID: 11337476
-
A complete map of the human ribosomal protein genes: assignment of 80 genes to the cytogenetic map and implications for human disorders.
Genomics. 2001 Mar 15;72(3):223-30
PMID: 11401437
-
A draft annotation and overview of the human genome.
Genome Biol. 2001;2(7):RESEARCH0025
PMID: 11516338
-
Protein family and fold occurrence in genomes: power-law behaviour and evolutionary model.
J Mol Biol. 2001 Nov 2;313(4):673-81
PMID: 11697896
-
Processed pseudogenes: characteristics and evolution.
Annu Rev Genet. 1985;19:253-72
PMID: 3909943
-
Analysis of compositionally biased regions in sequence databases.
Methods Enzymol. 1996;266:554-71
PMID: 8743706
-
Comparative analysis of 1196 orthologous mouse and human full-length mRNA and protein sequences.
Genome Res. 1996 Sep;6(9):846-57
PMID: 8889551
-
Evolution by the birth-and-death process in multigene families of the vertebrate immune system.
Proc Natl Acad Sci U S A. 1997 Jul 22;94(15):7799-806
PMID: 9223266
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
Comparison of DNA sequences with protein sequences.
Genomics. 1997 Nov 15;46(1):24-36
PMID: 9403055
-
A structural census of genomes: comparing bacterial, eukaryotic, and archaeal genomes in terms of protein structure.
J Mol Biol. 1997 Dec 12;274(4):562-76
PMID: 9417935
-
The DNA sequence of human chromosome 22.
Nature. 1999 Dec 2;402(6761):489-95
PMID: 10591208
-
GenBank.
Nucleic Acids Res. 2000 Jan 1;28(1):15-8
PMID: 10592170
-
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):45-8
PMID: 10592178
-
The GeneQuiz web server: protein functional analysis through the Web.
Trends Biochem Sci. 2000 Jan;25(1):33-5
PMID: 10637611
-
Vertebrate pseudogenes.
FEBS Lett. 2000 Feb 25;468(2-3):109-14
PMID: 10692568
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
The DNA sequence of human chromosome 21.
Nature. 2000 May 18;405(6784):311-9
PMID: 10830953
-
Analysis of expressed sequence tags indicates 35,000 human genes.
Nat Genet. 2000 Jun;25(2):232-4
PMID: 10835644
-
Estimate of human gene number provided by genome-wide analysis using Tetraodon nigroviridis DNA sequence.
Nat Genet. 2000 Jun;25(2):235-8
PMID: 10835645
-
The sequence of the human genome.
Science. 2001 Feb 16;291(5507):1304-51
PMID: 11181995
-
Evolution of genome size: new approaches to an old problem.
Trends Genet. 2001 Jan;17(1):23-8
PMID: 11163918
-
MaskerAid: a performance enhancement to RepeatMasker.
Bioinformatics. 2000 Nov;16(11):1040-1
PMID: 11159316
-
InterPro--an integrated documentation resource for protein families, domains and functional sites.
Bioinformatics. 2000 Dec;16(12):1145-50
PMID: 11159333
-
The genome sequence of Rickettsia prowazekii and the origin of mitochondria.
Nature. 1998 Nov 12;396(6707):133-40
PMID: 9823893
-
Patterns of protein-fold usage in eight microbial genomes: a comprehensive structural census.
Proteins. 1998 Dec 1;33(4):518-34
PMID: 9849936
-
Massive gene decay in the leprosy bacillus.
Nature. 2001 Feb 22;409(6823):1007-11
PMID: 11234002