Improvements in DNA sequencing technology mean that genome-wide DNA sequencing-based analyses can now be carried out using tiny amounts of starting mate rial, and genome-wide DNA sequencing in single human cells and single cells of model organisms is a rapidly advancing field of research. The term single-cell genomics is now widely used to cover broad DNA sequencing-based analyses in isolated single cells; it is not limited to analyses of genomic DNA, many of the studies carried out being devoted to analyzing transcriptomes and, to a lesser extent, epigenomes (mostly covering DNA methylation and histone modification states and chromatin conformation). That is, single-cell genomics describes large-scale DNA sequencing-based assays that follow changes in genomic DNA, chromatin, or RNA transcripts, often at a genome-wide level (Figure 1). Other types of single-cell analysis are also being carried out that use different methods and/or operate on a smaller scale, such as tracking transcripts or proteins expressed by one or often a small number of genes of interest (using fluorescence hybridization with antisense RNA probes or by using specific antibodies).

Fig1. Major facets of single-cell genomics.
Given the limiting amounts of starting material (which pushes existing technology to the limits), why should so much effort be expended to apply DNA sequencing methods to study individual human cells? The answer is that traditional cellular analyses have notable limitations because they are typically carried out on cell populations (such as bulk tissue samples and cell culture preparations). Because of cell-to-cell variation in phenotype, in gene expression, and even in genomic DNA, the resulting data are necessarily aggregate values: unless single cells are analyzed, all the original intercellular variation is hidden, and important rare cells can be overlooked.
An era has just begun where systems biology will be incrementally applied to understanding the workings of single cells. Cell-to-cell variation has important consequences in health and disease, and we consider first how single-cell analyses will provide exciting new inroads into both basic biology and medical research. We finish by describing some of the technical aspects that have allowed this revolutionary perspective. This is a very fast-moving area, and interested readers are recommended to consult recent reviews.
Understanding cell-to-cell variation: the myriad applications of single-cell genomics in biology
The Human Genome Project paved the way for a biological equivalent of chemistry’s periodic table: a periodic table of genes. Admittedly, that table is not a universal one (it varies from organism to organism, but there is a great deal of similarity in the genes of closely related organisms, such as different mammals and vertebrates). There is, however, another fundamental characteristic of living things: cells. What single-cell analyses now offer is the prospect of ultimately establishing a definitive catalog of the cells of multicellular organisms. In addition to the question of defining cell identity, we also briefly outline below how single-cell analyses can offer important insights into other aspects of cell-to-cell variation. Important applications are being found in diverse areas, including neuroscience, development, and also immunology (where, for example, little has previously been known about heterogeneous transcriptional responses in immune cells after activation).
Natural DNA variation in single cells
Though largely stable, the genomic DNA content of our cells does vary as a result of somatic DNA changes. This variation includes programmed changes in the DNA of various normal cell types (and also cancer cells, as described below). In addition, incremental random somatic mutations occur in all cells, so that each cell in our bodies has a unique genome.
Some topical areas of interest include studies of germ-cell DNA (charting aneuploidy, which occurs quite frequently in the gametes of normal humans) and recombination (by comparing the genomes of diploid cells with those of individual gametes from the same individual, crossovers can be mapped across the genome). Certain types of somatic cells are of interest because of a naturally high frequency of large structural changes in their DNA, notably neurons where large deletions are especially frequent, and large-scale duplications are also common. The question of to what extent the changes are related to neuron diversification is being explored. The study of natural DNA variation in human cells also permits, for the first time, detailed tracing of human cell lineages, as described in the next subsection.
Cell lineage tracing
The only complete metazoan cell lineage tree—a cell fate map beginning from the fertilized egg for the nematode C. elegans—was reported by John Sulston and colleagues, describing postembryonic lineages in 1976 and embryonic lineages in 1983. This tour de force was carried out by time-lapse microscopy, aided by the roundworm’s optical transparency, and by the modest size (several hundred cells) and invariant nature of the cell lineage. (Interested readers can find the worm lineages by typing the query “lineage” at http://www.wormatlas.org/ and in the original papers at PMID 838129 and PMID 6684600.) Partial lineage tracing has been carried out in additional model organisms. Often the studies seek to identify progeny of a cell of interest using clonal marking. To do that, the cell of interest is experimentally marked in some way (initially dyes or radiolabeled markers were used; more recently, genetic markers have been introduced into the cell), and descendants of that cell are followed by assaying for the marker.
Experimental clonal marking is not applicable for lineage tracing in humans, but somatic mutation represents a natural way of genetically marking cells: at each cell division, starting from the zygote, new somatic mutations are introduced that are transmit ted to progeny (Figure 2). Sequencing of hypervariable and other highly mutable sites across the genomes of single cells now allows lineage analyses in human cells (see Table 1 for some examples). This new dimension will allow progenitor cells to be defined for multiple cell types where we have limited existing knowledge (see the example of identifying novel candidate stem cells in Table 1).

Fig2. Cell lineage tracing using somatic mutations. Each cell in a multicellular organism has a unique genome because, starting from the zygote, many somatic mutations arise de novo (shown by red arrows) in a cell’s genome prior to each cell division (usually when the DNA replicates), and the two daughter cells inherit different sets of somatic mutations. The figure shows first-, second-, and third generation descendants of an ancestor cell (AC), and the colored boxes positioned on the lines connecting parent cell to daughter cells indicate new sets of mutations (not present in AC) that have occurred before the first (1), second (2), and third (3) generations. The vertically arranged boxes to the left or right of each descendant indicate acquired somatic mutations not present in AC but accumulated immediately prior to the first, second, and third generations. Thus, for example, the first two descendants of AC have acquired sets of somatic mutations (collectively labeled 1) not present in AC; the black and white numerals indicate that the somatic mutation sets acquired by the two cells are different. The descendants in the second and third generations have acquired additional somatic mutations that were subsequently generated. Although the spectrum of somatic mutations will differ from cell to cell, all eight cells in the third generation will inherit the new somatic mutations acquired by AC (0). By carrying out analyses at hypervariable DNA regions across the genomes of single cells, somatic mutational differences between cells can be identified and analyzed to construct lineage trees.

Table1. EXAMPLES OF RESEARCH AREAS WHERE SINGLE-CELL GENOME-WIDE DNA SEQUENCING IS ILLUMINATING UNDERSTANDING OF HUMAN AND MAMMALIAN BIOLOGY
Cell identity
Perhaps the most exciting application of single-cell genomics is to clarify our understanding of cell identity. We are likely to have somewhere in the region of 20 to 100 trillion cells, but the division into cell types has been based largely on just anatomical and morphological grounds. About 200 human cell types have been distinguished on this basis, but that number is widely regarded as a gross underestimate (some cell types, notably neurons and T cells, are known to be highly heterogeneous). Not only do we lack a complete catalog of our cells, but there is also limited knowledge of different cell states. (We even lack precise definitions for the intuitive terms “cell type” and “cell state.”)
Studying the transcriptomes and epigenomes of single cells holds the promise of fine-scale classification of cells, and identification of novel molecular markers associated with different novel cell subtypes. Rare, functionally important cells can also be identified that are simply not detectable by standard methods designed to study cell populations. While we currently appreciate the more obvious transient cell states, such as when cells transition through different stages of the cell cycle, deep molecular profiling of single cells should also identify more subtle cell states.
Table 1 gives some early examples of single-cell analyses that have begun to expand the range of cell types, plus identified previously unappreciated, functionally important rare cells. The greatly enhanced ability to profile single cells at the molecular level has recently prompted proposals for an international Human Cell Atlas project (see www. humancellatlas.org and PMID 29206104). The aim is to develop a comprehensive catalog of all human cells based on both their stable properties and transient features, and on cell positions and cell lineages. As well as providing invaluable markers, molecular sig natures, and tools for basic research, a Human Cell Atlas should have important clinical applications, as described below.
The prospects of advancing medical research using single-cell genomics
With the exception of pre-implantation diagnosis (where DNA-based diagnosis using single blastomeres from the early in vitro fertilization [IVF] embryo has long been available in many countries), single-cell analyses have not been part of medical practice. But the new single cell genomics technologies offer huge scope for advancing medical research and translational opportunities. One important area is cancer research, where single-cell genomics has been applied since 2011; we outline some applications in the subsection below.
Single-cell genomics offers important new dimensions in various general areas of medical research. One is the cellular basis of disease. Up until recently, cellular analyses of the pathogenic state have been achieved by analyzing mixtures of cells from dis ease tissues or other heterogeneous cell populations. Inevitably, the picture obtained is clouded by heterogeneity because tissues are made up of different cell types. Cell-to-cell differences in the case of the primary disease cells have not traditionally been explored, and there is no great understanding of the roles of neighboring secondary cells in initiating, promoting, or restraining disease. Analyses of single cells from dissociated tissue or even in situ analysis should provide a clearer picture, giving a full description of the expression profiles of individual cells, the dynamic states within each cell type, and the proportion of, and spatial relationships of, the various cell types. Single-cell analyses across many patients can then show how the picture varies during the course of the dis ease and the responses to treatment.
Other important benefits that might be expected to accrue from single-cell expression profiling include more precisely defined disease markers, expression signatures that identify stages of the disease process and of recovery after treatment, and cell therapy (more precise expression profiling might lead to more accurately defined desired cells to be used in therapy).
Cancer research
Cancers are defined by natural selection acting at the level of the cell to promote abnormal proliferation. During the development of cancerous changes, cancer cells acquire extraordinary genetic and epigenetic changes, and give rise to subpopulations with different cell properties. As a result, tumors are heterogeneous and clonal evolution is important in development of different properties. Rare cancer stem cells may be responsible for regrowth of tumors after cancer treatment. Other initially rare cells in a tumor may give rise to subpopulations of cells that are important in metastasis and in developing resistance to drug treatment. Identifying cell-to-cell variation within (and between) tumors is therefore an especially relevant application of single-cell genomics. We will consider this aspect in Chapter 19.
The technology of DNA-based sequencing assays in single cells
Single-cell genomics involves a succession of methodologies. First, there must be a way of isolating single cells efficiently. Secondly, DNA must be isolated in some way to provide a substrate that can be assayed. It may be done directly—by isolating whole genomic DNA or targeted subsets of genomic DNA from single cells—or indirectly (in transcriptome analysis, RNA is isolated from single cells and converted into cDNA). Because the amount of DNA or RNA isolated from single cells is so tiny, the DNA/cDNA must be amplified to give sufficient material for sequencing. Finally, the end game: high-throughput sequencing and data analysis. We consider some component steps in the subsections below.
Isolating single cells
Different methods have been used to isolate single cells. Manual methods have been used in the past (such as cell capture from tissues by laser microdissection, and manual micro manipulation) but automated methods predominate in single-cell genomics. Some popular methods depend on using labeled antibodies to bind to cell surface proteins characteristic of a cell type of interest; others rely on fluid flow in microchambers—examples are listed below.
• Flow cytometry. Cells are labeled using fluorescently labeled antibodies, then sorted by the degree of fluorescence they exhibit; because of spectral overlap, however, resolution is limited.
• Mass cytometry. A cross between flow cytometry and mass spectrometry, the cells are labeled using antibodies conjugated with heavy metal ions; it can dis criminate significantly more simultaneous signals than flow cytometry.
• Microfluidic cell sorting. Cells are suspended in fluid that flows through channels in prefabricated microchambers (where separation can occur according to inherent physical properties of the cells). Recently developed methods rely on mixing an aqueous solution containing cells with oil to produce an emulsion with very fine droplets that contain single cells (for an example, see the Drop-Seq method for whole-transcriptome analysis, as described below).
Isolating DNA fragments and epigenome analyses
The starting DNA may be fragmented genomic DNA (for genomic analyses) or fragmented RNA that has been converted by reverse transcription to give DNA fragments (for transcriptome analysis). For analysis of the epigenome, different properties—such as the genome-wide locations of methylated cytosines, specific histone variants, and bound transcription factors, and the conformation of chromatin—can each be tracked, ultimately, by a DNA sequencing-based assay. Example methods are described briefly below, and will be detailed in Chapters 9 and 10.
• In the ChIP-Seq method, antibodies specific for DNA-binding proteins of interest, such as specific histone variants and individual types of transcription factor, are used to define DNA binding sites within chromatin of the proteins of interest. That is, they can map all locations across the genome where such a protein of interest is bound. The assay depends on first adding agents that will cause chemical cross linking of all proteins bound to DNA within cells (proteins bound to chromatin by noncovalent bonds then become covalently bound to the DNA).
• Methyl-Seq involves treating DNA fragments with sodium bisulfite. Nonmethylated cytosines are chemically converted to give uracils, while 5-methylcytosine and hydroxymethylcytosine are unaffected, allowing map ping of methylated cytosines.
• HiC-Seq and other DNA sequencing-based methods for analyzing chromatin conformation will be described in Section 10.1.
Whole-genome amplification
The total genomic DNA of a human cell is typically less than 10 pg, and current technology requires amplification of the DNA fragments of interest. Three types of method are used for whole-genome amplification, as listed below each with advantages and disadvantages. One problem is amplification bias: certain sequences in the starting genome do not amplify very well compared to others. As a result, genome coverage may be comparatively poor (a significant proportion of the genome region is not represented in the amplified DNA fragments), and sometimes specific alleles may be preferentially amplified or not amplified at all (allele dropout).
• PCR-based amplification methods. Adaptor-ligation PCR has been used, but more widely used methods use degenerate oligonucleotide primers. That is, instead of having single primers, sets of primers with closely related sequences are used where for each degenerate nucleotide position, some members of the primer set have an A, others have a C, a G, or a T. In degenerate oligonucleotide primer PCR (DOP-PCR) about six or so nucleotide positions, located some distance away from the 3′ end of each primer, are designed to be degenerate. The method is disadvantaged by relatively poor genome coverage, but amplification is otherwise uniform and it is useful for studying copy number variation.
• Multiple displacement amplification (MDA). An isothermal amplification method, it uses random hexanucleotide primers and the φ29 polymerase, which binds very strongly to single-stranded DNA and is highly-processive (it remains bound to DNA over long distances, synthesizing DNA as it goes; other polymerases drop off the DNA more readily). As a result of its highly-processive properties, the φ29 polymerase readily displaces other newly synthesized strands previously formed from primers binding downstream, resulting in branched structures (Figure 3A) and exponential amplification. With MDA, genome coverage is high, often 80–90% of the genome, but amplification is nonuniform and so it is not well suited to assaying copy number variation.
• Hybrid methods. Multiple annealing and looping-based amplification cycles (MALBAC) begins with a quasilinear MDA-like amplification designed to minimize amplification bias (special primers are used that enable looping of the amplicons to pre vent them from being further amplified in subsequent MALBAC cycles). After a number of looping cycles have taken place, PCR amplification is carried out (Figure 3B). Genome coverage is high and amplification is uniform, but the DNA polymerase used in MALBAC has a relatively high error rate compared to the φ29 polymerase.

Fig3. Multiple displacement amplification (MDA) and multiple annealing and looping-based amplification cycles (MALBAC). (A) Multiple displacement amplification. Random primers are used for isothermal amplification with the φ29 DNA polymerase, which has a strong displacement activity and generates new DNA strands that are many kilobases in length. (B) MALBAC. Random primers with a fixed sequence are used in a temperature cycle in which only the original genomic DNA and semiamplicons are linearly amplified, and full amplicons are protected from further amplification by the formation of DNA loops owing to the complementarity of the fixed sequences at the 3′ and 5′ ends. The DNA loops are PCR-amplified at the final stage. Here, m is the number of temperature cycles (m = 0 ∼ 10) and n is the number of primers bound; (m + 1) × n is the number of semiamplicons present at the mth cycle, and m × n2 is the number of full amplicons generated in the mth cycle. (A and B, from Huang L et al. [2015] Annu Rev Genomics Hum Genet 16:79–102; PMID 26077818. With permission from Annual Reviews. Permission conveyed through Copyright Clearance Center, Inc.)
Whole-transcriptome analysis
Until very recently, single-cell transcriptome analysis was hampered because existing methods were not suited to analyzing very large numbers of whole transcriptomes, and required the transcriptomes to be sequenced one after another. In 2015, however, new methods were reported that both sequenced the transcriptomes in parallel and were scalable, allowing highly-parallel whole-transcriptome sequencing.
The new methods involve using primers attached to microbeads and tiny reaction chambers. In one case, a miniature dish with tiny wells is used that contains at most single cells—the volume of the well is just 20 picoliters and the dosage of cells is deliberately kept very low so that most wells do not receive any cells, but those that do can be expected to contain single cells. Other methods rely on microfluidic systems and creating emulsions to trap individual cells in tiny aqueous droplets. An additional feature of the new methods is the use of molecular barcoding systems. As an illustration we describe in Figure 4 the Drop-Seq method that was used in 2015 to carry out parallel sequencing of the transcriptomes of several tens of thousands of cells. It relies on using molecular barcoding to identify both cell of origin and individual RNA molecules.

Fig4. Drop-Seq: highly-parallel single-cell transcriptome sequencing using droplet reaction chambers and random sequence tags (“barcodes”) to indicate cell and molecule of origin for sequence reads. (A) Method overview. After dissociation from a tissue, individual cells are encapsulated in droplets along with microparticles (gray circles) that deliver barcoded primers. Each cell is lysed within a droplet; its mRNAs bind to the primers on its companion microparticle. The mRNAs are reverse-transcribed into cDNAs, generating a set of beads called STAMPs (single-cell transcriptomes attached to microparticles). Barcoded STAMPs are amplified in pools for high-throughput mRNA-Seq to analyze any desired number of individual cells. (B) Each microparticle contains more than 108 individual primers with four types of sequence. The two end sequences are invariant: a “PCR handle’’ sequence (to allow PCR amplification after STAMP formation) and an oligo(dT) sequence (to capture mRNAs). The two central sequences are variable: a 12-nucleotide cell barcode (common to all primers within one microparticle, but differing between microparticles) and an 8-nucleotide unique molecular identifier (UMIs differ from one primer to another within a microparticle, enabling mRNA transcripts to be digitally counted). (C) Cell barcode synthesis. The pool of microparticles is repeatedly split into four equally sized oligonucleotide synthesis reactions, to each of which one of the four DNA bases is added, and which are then pooled together after each cycle, in a total of 12 split-pool cycles. The barcode synthesized on any individual bead reflects that bead’s unique path through the series of synthesis reactions. The result is a pool of microparticles, each possessing one of 412 (16,777,216) possible 12-nucleotide sequences on its entire complement of primers. (D) Molecular barcode synthesis. Following completion of the ‘‘split-and-pool’’ synthesis cycles, all microparticles are together subjected to eight rounds of degenerate synthesis with all four DNA bases available during each cycle, such that each individual primer receives one of 48 (65,536) possible 8-nucleotide sequences (UMIs). (E) In silico reconstruction. Millions of paired-end reads are generated from a Drop-Seq library, representing many thousands of single-cell transcriptomes. The reads are first aligned to a reference genome to identify the gene of origin of the cDNA. Next, reads are organized by their cell barcodes, and individual UMIs are counted for each gene in each cell. The result (shown at bottom extreme right) is a ‘‘digital expression matrix’’ (far right): each column corresponds to a cell, each row corresponds to a gene, and each entry is the integer number of transcripts detected from that gene in that cell. (From Macosko EZ et al. [2015] Cell 161:1202–1214; PMID 26000488. With permission from Elsevier.)