Abstract
RNA-sequencing (RNA-seq) allows quantitative measurement of expression levels of genes and their transcripts. In this study, we sequenced complementary DNA fragments of cultured human B-cells and obtained 879 million 50-bp reads comprising 44 Gb of sequence. The results allowed us to study the gene expression profile of B-cells and to determine experimental parameters for sequencing-based expression studies. We identified 20,766 genes and 67,453 of their alternatively spliced transcripts. More than 90% of the genes with multiple exons are alternatively spliced; for most genes, one isoform is predominantly expressed. We found that while chromosomes differ in gene density, the percentage of transcribed genes in each chromosome is less variable. In addition, genes involved in related biological processes are expressed at more similar levels than genes with different functions. Besides characterizing gene expression, we also used the data to investigate the effect of sequencing depth on gene expression measurements. While 100 million reads are sufficient to detect most expressed genes and transcripts, about 500 million reads are needed to measure accurately their expression levels. We provide examples in which deep sequencing is needed to determine the relative abundance of genes and their isoforms. With data from 20 individuals and about 40 million sequence reads per sample, we uncovered only 21 alternatively spliced, multi-exon genes that are not in databases; this result suggests that at this sequence coverage, we can detect most of the known genes. Results from this project are available on the UCSC Genome Browser to allow readers to study the expression and structure of genes in human B-cells.
MeSH Terms
B-Lymphocytes/metabolism
DNA, Complementary/genetics
Gene Expression Profiling/methods
High-Throughput Nucleotide Sequencing/methods
Humans
Protein Isoforms/metabolism
Proteins/metabolism
Sequence Analysis, RNA/methods
Chemicals
DNA, Complementary
Protein Isoforms
Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Toung Jonathan M
Genomics and Computational Biology Program, University of Pennsylvania, Philadelphia, Pennsylvania 19104, USA.
Morley Michael
Li Mingyao
Cheung Vivian G
References (27)
27 references, click to expand
-
Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing.
Nat Genet. 2008 Dec;40(12):1413-5
PMID: 18978789
-
Gene expression analysis by massively parallel signature sequencing (MPSS) on microbead arrays.
Nat Biotechnol. 2000 Jun;18(6):630-4
PMID: 10835600
-
GENCODE: producing a reference annotation for ENCODE.
Genome Biol. 2006;7 Suppl 1:S4.1-9
PMID: 16925838
-
Genetic analysis of genome-wide variation in human gene expression.
Nature. 2004 Aug 12;430(7001):743-7
PMID: 15269782
-
Noisy splicing drives mRNA isoform diversity in human cells.
PLoS Genet. 2010 Dec 09;6(12):e1001236
PMID: 21151575
-
Highly integrated single-base resolution maps of the epigenome in Arabidopsis.
Cell. 2008 May 2;133(3):523-36
PMID: 18423832
-
Purification to homogeneity and partial characterization of cytotoxic lymphocyte maturation factor from human B-lymphoblastoid cells.
Proc Natl Acad Sci U S A. 1990 Sep;87(17):6808-12
PMID: 2204066
-
Stem cell transcriptome profiling via massive-scale mRNA sequencing.
Nat Methods. 2008 Jul;5(7):613-9
PMID: 18516046
-
Accurate whole human genome sequencing using reversible terminator chemistry.
Nature. 2008 Nov 6;456(7218):53-9
PMID: 18987734
-
A genome-wide association study of global gene expression.
Nat Genet. 2007 Oct;39(10):1202-7
PMID: 17873877
-
Mapping and quantifying mammalian transcriptomes by RNA-Seq.
Nat Methods. 2008 Jul;5(7):621-8
PMID: 18516045
-
Multiplexed biochemical assays with biological chips.
Nature. 1993 Aug 5;364(6437):555-6
PMID: 7687751
-
CTLA-4 is a second receptor for the B cell activation antigen B7.
J Exp Med. 1991 Sep 1;174(3):561-9
PMID: 1714933
-
Segregation of MHC class II molecules from MHC class I molecules in the Golgi complex for transport to lysosomal compartments.
Nature. 1991 Feb 21;349(6311):669-76
PMID: 1847504
-
Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
Genome Biol. 2009;10(3):R25
PMID: 19261174
-
Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation.
Nat Biotechnol. 2010 May;28(5):511-5
PMID: 20436464
-
Centre d'etude du polymorphisme humain (CEPH): collaborative genetic mapping of the human genome.
Genomics. 1990 Mar;6(3):575-7
PMID: 2184120
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Dynamic repertoire of a eukaryotic transcriptome surveyed at single-nucleotide resolution.
Nature. 2008 Jun 26;453(7199):1239-43
PMID: 18488015
-
Transcriptome genetics using second generation sequencing in a Caucasian population.
Nature. 2010 Apr 1;464(7289):773-7
PMID: 20220756
-
Alternative isoform regulation in human tissue transcriptomes.
Nature. 2008 Nov 27;456(7221):470-6
PMID: 18978772
-
ENCODE whole-genome data in the UCSC Genome Browser.
Nucleic Acids Res. 2010 Jan;38(Database issue):D620-5
PMID: 19920125
-
Heritability and linkage analysis of sensitivity to cisplatin-induced cytotoxicity.
Cancer Res. 2004 Jun 15;64(12):4353-6
PMID: 15205351
-
Use of a cDNA microarray to analyse gene expression patterns in human cancer.
Nat Genet. 1996 Dec;14(4):457-60
PMID: 8944026
-
Serial analysis of gene expression.
Science. 1995 Oct 20;270(5235):484-7
PMID: 7570003
-
The transcriptional landscape of the yeast genome defined by RNA sequencing.
Science. 2008 Jun 6;320(5881):1344-9
PMID: 18451266
-
TopHat: discovering splice junctions with RNA-Seq.
Bioinformatics. 2009 May 1;25(9):1105-11
PMID: 19289445