Home LiteratureArticle Details
PMID: 11160906 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Digging for dead genes: an analysis of the characteristics of the pseudogene population in the Caenorhabditis elegans genome.

Nucleic acids research ·Vol. 29 ·No. 3 ·2001-02-01 ·Pages 818-30

Harrison PM, Echols N, Gerstein MB

Abstract

Pseudogenes are non-functioning copies of genes in genomic DNA, which may either result from reverse transcription from an mRNA transcript (processed pseudogenes) or from gene duplication and subsequent disablement (non-processed pseudogenes). As pseudogenes are apparently 'dead', they usually have a variety of obvious disablements (e.g., insertions, deletions, frameshifts and truncations) relative to their functioning homologs. We have derived an initial estimate of the size, distribution and characteristics of the pseudogene population in the Caenorhabditis elegans genome, performing a survey in 'molecular archaeology'. Corresponding to the 18 576 annotated proteins in the worm (i.e., in Wormpep18), we have found an estimated total of 2168 pseudogenes, about one for every eight genes. Few of these appear to be processed. Details of our pseudogene assignments are available from http://bioinfo.mbb.yale.edu/genome/worm/pseudogene. The population of pseudogenes differs significantly from that of genes in a number of respects: (i) pseudogenes are distributed unevenly across the genome relative to genes, with a disproportionate number on chromosome IV; (ii) the density of pseudogenes is higher on the arms of the chromosomes; (iii) the amino acid composition of pseudogenes is midway between that of genes and (translations of) random intergenic DNA, with enrichment of Phe, Ile, Leu and Lys, and depletion of Asp, Ala, Glu and Gly relative to the worm proteome; and (iv) the most common protein folds and families differ somewhat between genes and pseudogenes-whereas the most common fold found in the worm proteome is the immunoglobulin fold and the most common 'pseudofold' is the C-type lectin. In addition, the size of a gene family bears little overall relationship to the size of its corresponding pseudogene complement, indicating a highly dynamic genome. There are in fact a number of families associated with large populations of pseudogenes. For example, one family of seven-transmembrane receptors (represented by gene B0334.7) has one pseudogene for every four genes, and another uncharacterized family (represented by gene B0403.1) is approximately two-thirds pseudogenic. Furthermore, over a hundred apparent pseudogenic fragments do not have any obvious homologs in the worm.

MeSH Terms
Amino Acid Sequence Animals Caenorhabditis elegans/genetics Chromosomes/genetics DNA, Helminth/genetics Genes, Helminth/genetics Genome Molecular Sequence Data Pseudogenes/genetics
Chemicals
DNA, Helminth
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Harrison P M
Department of Molecular Biophysics and Biochemistry, Yale University, 260 Whitney Avenue, PO Box 208114, New Haven, CT 06511-8114, USA.
Echols N
Gerstein M B
References (44)
44 references, click to expand
  1. Two large families of chemoreceptor genes in the nematodes Caenorhabditis elegans and Caenorhabditis briggsae reveal extensive gene duplication, diversification, movement, and intron loss.
    Genome Res. 1998 May;8(5):449-63 PMID: 9582190
  2. Processed pseudogenes: characteristics and evolution.
    Annu Rev Genet. 1985;19:253-72 PMID: 3909943
  3. Patterns of protein-fold usage in eight microbial genomes: a comprehensive structural census.
    Proteins. 1998 Dec 1;33(4):518-34 PMID: 9849936
  4. Genome sequence of the nematode C. elegans: a platform for investigating biology.
    Science. 1998 Dec 11;282(5396):2012-8 PMID: 9851916
  5. Comparison of the complete protein sets of worm and yeast: orthology and divergence.
    Science. 1998 Dec 11;282(5396):2022-8 PMID: 9851918
  6. How representative are the known structures of the proteins in a complete genome? A comprehensive structural census.
    Fold Des. 1998;3(6):497-512 PMID: 9889159
  7. Cloning, mRNA localization and evolutionary conservation of a human 5-HT7 receptor pseudogene.
    Gene. 1999 Feb 4;227(1):63-9 PMID: 9931439
  8. Transcriptional analysis of the PTEN/MMAC1 pseudogene, psiPTEN.
    Oncogene. 1999 Mar 4;18(9):1765-9 PMID: 10208437
  9. Neuronal expression of neural nitric oxide synthase (nNOS) protein is suppressed by an antisense RNA transcribed from an NOS pseudogene.
    J Neurosci. 1999 Sep 15;19(18):7711-20 PMID: 10479675
  10. Genomic analysis of Caenorhabditis elegans reveals ancient families of retroviral-like elements.
    Genome Res. 1999 Oct;9(10):924-35 PMID: 10523521
  11. Detection of a putative HLA-A*31012 processed (intronless) pseudogene in a laryngeal squamous cell carcinoma.
    Genes Chromosomes Cancer. 2000 Jan;27(1):26-34 PMID: 10564583
  12. The DNA sequence of human chromosome 22.
    Nature. 1999 Dec 2;402(6761):489-95 PMID: 10591208
  13. Nonviral retroposons: genes, pseudogenes, and transposable elements generated by the reverse flow of genetic information.
    Annu Rev Biochem. 1986;55:631-61 PMID: 2427017
  14. An amber mutation of prion protein in Gerstmann-Sträussler syndrome with mutant PrP plaques.
    Biochem Biophys Res Commun. 1993 Apr 30;192(2):525-31 PMID: 8097911
  15. Selection of representative protein data sets.
    Protein Sci. 1992 Mar;1(3):409-17 PMID: 1304348
  16. Unusual molecular evolution of an Adh pseudogene in Drosophila.
    Mol Biol Evol. 1994 May;11(3):443-58 PMID: 8015438
  17. Structure, expression and duplication of genes which encode phosphoglyceromutase of Drosophila melanogaster.
    Genetics. 1994 Oct;138(2):352-63 PMID: 7828819
  18. Evolution of immunoglobulin VH pseudogenes in chickens.
    Mol Biol Evol. 1995 Jan;12(1):94-102 PMID: 7877500
  19. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  20. The size distribution of insertions and deletions in human and rodent pseudogenes suggests the logarithmic gap penalty for sequence alignment.
    J Mol Evol. 1995 Apr;40(4):464-73 PMID: 7769622
  21. Rte-1, a retrotransposon-like element in Caenorhabditis elegans.
    FEBS Lett. 1996 Feb 12;380(1-2):1-7 PMID: 8603714
  22. Analysis of compositionally biased regions in sequence databases.
    Methods Enzymol. 1996;266:554-71 PMID: 8743706
  23. Computer analyses reveal a hobo-like element in the nematode Caenorhabditis elegans, which presents a conserved transposase domain common with the Tc1-Mariner transposon family.
    Gene. 1996 Oct 3;174(2):265-71 PMID: 8890745
  24. High intrinsic rate of DNA loss in Drosophila.
    Nature. 1996 Nov 28;384(6607):346-9 PMID: 8934517
  25. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  26. A structural census of the current population of protein sequences.
    Proc Natl Acad Sci U S A. 1997 Oct 28;94(22):11911-6 PMID: 9342336
  27. Comparison of DNA sequences with protein sequences.
    Genomics. 1997 Nov 15;46(1):24-36 PMID: 9403055
  28. A structural census of genomes: comparing bacterial, eukaryotic, and archaeal genomes in terms of protein structure.
    J Mol Biol. 1997 Dec 12;274(4):562-76 PMID: 9417935
  29. Patterns and rates of indel evolution in processed pseudogenes from humans and murids.
    Gene. 1997 Dec 31;205(1-2):191-202 PMID: 9461394
  30. ProtoMap: automatic classification of protein sequences and hierarchy of protein families.
    Nucleic Acids Res. 2000 Jan 1;28(1):49-55 PMID: 10592179
  31. The intronerator: exploring introns and alternative splicing in Caenorhabditis elegans.
    Nucleic Acids Res. 2000 Jan 1;28(1):91-3 PMID: 10592190
  32. Evidence for DNA loss as a determinant of genome size.
    Science. 2000 Feb 11;287(5455):1060-2 PMID: 10669421
  33. The large srh family of chemoreceptor genes in Caenorhabditis nematodes reveals processes of genome evolution involving large duplications and deletions and intron gains and losses.
    Genome Res. 2000 Feb;10(2):192-203 PMID: 10673277
  34. Evidence for a high frequency of simultaneous double-nucleotide substitutions.
    Science. 2000 Feb 18;287(5456):1283-6 PMID: 10678838
  35. Analysis of the yeast transcriptome with structural and functional categories: characterizing highly expressed proteins.
    Nucleic Acids Res. 2000 Mar 15;28(6):1481-8 PMID: 10684945
  36. Vertebrate pseudogenes.
    FEBS Lett. 2000 Feb 25;468(2-3):109-14 PMID: 10692568
  37. Nature and structure of human genes that generate retropseudogenes.
    Genome Res. 2000 May;10(5):672-8 PMID: 10810090
  38. Analysis of expressed sequence tags indicates 35,000 human genes.
    Nat Genet. 2000 Jun;25(2):232-4 PMID: 10835644
  39. Gene index analysis of the human genome estimates approximately 120,000 genes.
    Nat Genet. 2000 Jun;25(2):239-40 PMID: 10835646
  40. Protein folds in the worm genome.
    Pac Symp Biocomput. 2000;:30-41 PMID: 10902154
  41. A global profile of germline gene expression in C. elegans.
    Mol Cell. 2000 Sep;6(3):605-16 PMID: 11030340
  42. Patterns of nucleotide substitution in pseudogenes and functional genes.
    J Mol Evol. 1982;18(5):360-9 PMID: 7120431
  43. Nonrandomness of point mutation as reflected in nucleotide substitutions in pseudogenes and its evolutionary implications.
    J Mol Evol. 1984;21(1):58-71 PMID: 6442359
  44. Biased usages of arginines and lysines in proteins are correlated with local-scale fluctuations of the G + C content of DNA sequences.
    J Mol Evol. 1998 Oct;47(4):385-93 PMID: 9767684
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2001-02-01
Pages
818-30
Language
English
Region
England
NLM ID
0411011
PMCID
PMC30377
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: product@genelibs.com