DNA SEQUENCING
DNA sequencing is the process of determining the precise order of nucleotides within a DNA molecule. It includes any method or technology that is used to determine the order of the four bases—adenine, guanine, cytosine, and thymine—in a strand of DNA. The advent of rapid DNA sequencing methods has greatly accelerated biological and medical research and discovery.
Knowledge of DNA sequences has become indispensable for basic biological research, and in numerous applied fields such as diagnostic, biotechnology, forensic biology, and biological systematics. The rapid speed of sequencing attained with modern DNA sequencing technology has been instrumental in the sequencing of complete DNA sequences, or genomes of numerous types and species of life, including the human genome and other complete DNA sequences of many animal, plant, and microbial species.An example of the results of automated chain-termination DNA sequencing.
The first DNA sequences were obtained in the early 1970s by academic researchers using laborious methods based on two-dimensional chromatography. Following the development of fluorescence-based sequencing methods with automated analysis,DNA sequencing has
DNA sequencing may be used to determine the sequence of individual genes, larger genetic regions (i.e. clusters of genes or operons), full chromosomes or entire genomes. Depending on the methods used, sequencing may provide the order of nucleotides in DNA or RNA isolated from cells of animals, plants, bacteria, archaea, or virtually any other source of genetic information. The resulting sequences may be used by researchers in molecular biology or genetics to further scientific progress or may be used by medical personnel to make treatment decisions or aid in genetic counseling.
Though the structure of DNA was established as a double helix in 1953,several decades would pass before fragments of DNA could be reliably analyzed for their sequence in the laboratory. RNA sequencing was one of the earliest forms of nucleotide sequencing. The major landmark of RNA sequencing is the sequence of the first complete gene and the complete genome of Bacteriophage MS2, identified and published by Walter Fiers and his coworkers at the University of Ghent (Ghent, Belgium), in 1972and 1976.
Several notable advancements in DNA sequencing were made during the 1970s. Frederick Sanger developed rapid DNA sequencing methods at the MRC Centre, Cambridge, UK and published a method for "DNA sequencing with chain-terminating inhibitors" in 1977.Walter Gilbert and Allan Maxam at Harvard also developed sequencing methods, including one for "DNA sequencing by chemical degradation".In 1973, Gilbert and Maxam reported the sequence of 24 basepairs using a method known as wandering-spot analysis.Advancements in sequencing were aided by the concurrent development of recombinant DNA technology, allowing DNA samples to be isolated from sources other than viruses.
The first full DNA genome to be sequenced was that of bacteriophage φX174 in 1977.Medical Research Council scientists deciphered the complete DNA sequence of the Epstein-Barr virus in 1984, finding it to be 170 thousand base-pairs long.
Leroy E. Hood's laboratory at the California Institute of Technology and Smith announced the first semi-automated DNA sequencing machine in 1986.[citation needed] This was followed by Applied Biosystems' marketing of the first fully automated sequencing machine, the ABI 370, in 1987. By 1990, the U.S. National Institutes of Health (NIH) had begun large-scale sequencing trials on Mycoplasma capricolum, Escherichia coli, Caenorhabditis elegans, and Saccharomyces cerevisiae at a cost of US$0.75 per base. Meanwhile, sequencing of human cDNA sequences called expressed sequence tags began in Craig Venter's lab, an attempt to capture the coding fraction of the human genome.In 1995, Venter, Hamilton Smith, and colleagues at The Institute for Genomic Research (TIGR) published the first complete genome of a free-living organism, the bacterium Haemophilus influenzae. The circular chromosome contains 1,830,137 bases and its publication in the journal Science marked the first published use of whole-genome shotgun sequencing, eliminating the need for initial mapping efforts. By 2001, shotgun sequencing methods had been used to produce a draft sequence of the human genome.
Several new methods for DNA sequencing were developed in the mid to late 1990s. These techniques comprise the first of the "next-generation" sequencing methods. In 1996, Pål Nyrén and his student Mostafa Ronaghi at the Royal Institute of Technology in Stockholm published their method of pyrosequencing.A year later, Pascal Mayer and Laurent Farinelli submitted patents to the World Intellectual Property Organization describing DNA colony sequencing.Lynx Therapeutics published and marketed "Massively parallel signature sequencing", or MPSS, in 2000. This method incorporated a parallelized, adapter/ligation-mediated, bead-based sequencing technology and served as the first commercially available "next-generation" sequencing method, though no DNA sequencers were sold to independent laboratorie.In 2004, 454 Life Sciences marketed a parallelized version of pyrosequencing.The first version of their machine reduced sequencing costs 6-fold compared to automated Sanger sequencing, and was the second of the new generation of sequencing technologies, after MPSS.
The large quantities of data produced by DNA sequencing have also required development of new methods and programs for sequence analysis. Phil Green and Brent Ewing of the University of Washington described their phred quality score for sequencer data analysis in 1998.
Types of DNA Sequencing
Maxam-Gilbert sequencing
Allan Maxam and Walter Gilbert published a DNA sequencing method in 1977 based on chemical modification of DNA and subsequent cleavage at specific bases.Also known as chemical sequencing, this method allowed purified samples of double-stranded DNA to be used without further cloning. This method's use of radioactive labeling and its technical complexity discouraged extensive use after refinements in the Sanger methods had been made.
Maxam-Gilbert sequencing requires radioactive labeling at one 5' end of the DNA and purification of the DNA fragment to be sequenced. Chemical treatment then generates breaks at a small proportion of one or two of the four nucleotide bases in each of four reactions (G, A+G, C, C+T). The concentration of the modifying chemicals is controlled to introduce on average one modification per DNA molecule. Thus a series of labeled fragments is generated, from the radiolabeled end to the first "cut" site in each molecule. The fragments in the four reactions are electrophoresed side by side in denaturing acrylamide gels for size separation. To visualize the fragments, the gel is exposed to X-ray film for autoradiography, yielding a series of dark bands each corresponding to a radiolabeled DNA fragment, from which the sequence may be inferred.
Procedure
Maxam–Gilbert sequencing requires radioactive labeling at one 5′ end of the DNA fragment to be sequenced (typically by a kinase reaction using gamma-32P ATP) and purification of the DNA. Chemical treatment generates breaks at a small proportion of one or two of the four nucleotide bases in each of four reactions (G, A+G, C, C+T). For example, the purines (A+G) are depurinated using formic acid, the guanines (and to some extent the adenines) are methylated by dimethyl sulfate, and the pyrimidines (C+T) are methylated using hydrazine. The addition of salt (sodium chloride) to the hydrazine reaction inhibits the reaction of thymine for the C-only reaction. The modified DNAs may then be cleaved by hot piperidine at the position of the modified base. The concentration of the modifying chemicals is controlled to introduce on average one modification per DNA molecule. Thus a series of labeled fragments is generated, from the radiolabeled end to the first "cut" site in each molecule.
The fragments in the four reactions are electrophoresed side by side in denaturing acrylamide gels for size separation. To visualize the fragments, the gel is exposed to X-ray film for autoradiography, yielding a series of dark bands each showing the location of identical radiolabeled DNA molecules. From presence and absence of certain fragments the sequence may be inferred.
Chain-termination methods
Sanger sequencing
The chain-termination method developed by Frederick Sanger and coworkers in 1977 soon became the method of choice, owing to its relative ease and reliability.[6][22] When invented, the chain-terminator method used fewer toxic chemicals and lower amounts of radioactivity than the Maxam and Gilbert method. Because of its comparative ease, the Sanger method was soon automated and was the method used in the first generation of DNA sequencers.
Sanger sequencing is the method which prevailed from the 80's until the mid-2000's. Over that period, great advances were made in the technique, such as fluorescent labelling, capillary electrophoresis, and general automation. These developments allowed much more efficient sequencing, leading to lower costs. The Sanger method, in mass production form, is the technology which produced the first human genome in 2001, ushering in the age of genomics. However, later in the decade, radically different approaches reached the market, bringing the cost per genome down from $100 million in 2001 to $10,000 in 2011.
Method
The classical chain-termination method requires a single-stranded DNA template, a DNA primer, a DNA polymerase, normal deoxynucleosidetriphosphates (dNTPs), and modified nucleotides (dideoxyNTPs) that terminate DNA strand elongation. These chain-terminating nucleotides lack a 3'-OH group required for the formation of a phosphodiester bond between two nucleotides, causing DNA polymerase to cease extension of DNA when a ddNTP is incorporated. The ddNTPs may be radioactively or fluorescently labelled for detection in automated sequencing machines.
The DNA sample is divided into four separate sequencing reactions, containing all four of the standard deoxynucleotides (dATP, dGTP, dCTP and dTTP) and the DNA polymerase. To each reaction is added only one of the four dideoxynucleotides (ddATP, ddGTP, ddCTP, or ddTTP), while three other nucleotides are ordinary ones. Putting it in a more sensible order, four separate reactions are needed in this process to test all four ddNTPs. Following rounds of template DNA extension from the bound primer, the resulting DNA fragments are heat denatured and separated by size using gel electrophoresis. This is frequently performed using a denaturing polyacrylamide-urea gel with each of the four reactions run in one of four individual lanes (lanes A, T, G, C). The DNA bands may then be visualized by autoradiography or UV light and the DNA sequence can be directly read off the X-ray film or gel image.
Part of a radioactively labelled sequencing gel
In the image on the right, X-ray film was exposed to the gel, and the dark bands correspond to DNA fragments of different lengths. A dark band in a lane indicates a DNA fragment that is the result of chain termination after incorporation of a dideoxynucleotide (ddATP, ddGTP, ddCTP, or ddTTP). The relative positions of the different bands among the four lanes, from bottom to top, are then used to read the DNA sequence.
DNA fragments are labelled with a radioactive or fluorescent tag on the primer, in the new DNA strand with a labeled dNTP, or with a labeled ddNTP.
Technical variations of chain-termination sequencing include tagging with nucleotides containing radioactive phosphorus for radiolabelling, or using a primer labeled at the 5' end with a fluorescent dye. Dye-primer sequencing facilitates reading in an optical system for faster and more economical analysis and automation. The later development by Leroy Hood and coworkersof fluorescently labeled ddNTPs and primers set the stage for automated, high-throughput DNA sequencing.
Sequence ladder by radioactive sequencing compared to fluorescent peaks
Chain-termination methods have greatly simplified DNA sequencing. For example, chain-termination-based kits are commercially available that contain the reagents needed for sequencing, pre-aliquoted and ready to use. Limitations include non-specific binding of the primer to the DNA, affecting accurate read-out of the DNA sequence, and DNA secondary structures affecting the fidelity of the sequence.
Dye-terminator sequencing
Capillary electrophoresis
Dye-terminator sequencing utilizes labelling of the chain terminator ddNTPs, which permits sequencing in a single reaction, rather than four reactions as in the labelled-primer method. In dye-terminator sequencing, each of the four dideoxynucleotide chain terminators is labelled with fluorescent dyes, each of which emit light at different wavelengths.
Owing to its greater expediency and speed, dye-terminator sequencing is now the mainstay in automated sequencing. Its limitations include dye effects due to differences in the incorporation of the dye-labelled chain terminators into the DNA fragment, resulting in unequal peak heights and shapes in the electronic DNA sequence trace chromatogram after capillary electrophoresis
This problem has been addressed with the use of modified DNA polymerase enzyme systems and dyes that minimize incorporation variability, as well as methods for eliminating "dye blobs". The dye-terminator sequencing method, along with automated high-throughput DNA sequence analyzers, is now being used for the vast majority of sequencing projects.
Automation and sample preparation
Automated DNA-sequencing instruments (DNA sequencers) can sequence up to 384 DNA samples in a single batch (run) in up to 24 runs a day. DNA sequencers carry out capillary electrophoresis for size separation, detection and recording of dye fluorescence, and data output as fluorescent peak trace chromatograms. Sequencing reactions by thermocycling, cleanup and re-suspension in a buffer solution before loading onto the sequencer are performed separately. A number of commercial and non-commercial software packages can trim low-quality DNA traces automatically. These programs score the quality of each peak and remove low-quality base peaks (generally located at the ends of the sequence). The accuracy of such algorithms is below visual examination by a human operator, but sufficient for automated processing of large sequence data sets.
Challenges
Common challenges of DNA sequencing with the Sanger method include poor quality in the first 15-40 bases of the sequence due to primer binding and deteriorating quality of sequencing traces after 700-900 bases. Base calling software such as Phred typically provides an estimate of quality to aid in trimming of low-quality regions of sequences.
In cases where DNA fragments are cloned before sequencing, the resulting sequence may contain parts of the cloning vector. In contrast, PCR-based cloning and next-generation sequencing technologies based on pyrosequencing often avoid using cloning vectors. Recently, one-step Sanger sequencing (combined amplification and sequencing) methods such as Ampliseq and SeqSharp have been developed that allow rapid sequencing of target genes without cloning or prior amplification.[7][8]
Current methods can directly sequence only relatively short (300-1000 nucleotides long) DNA fragments in a single reaction. The main obstacle to sequencing DNA fragments above this size limit is insufficient power of separation for resolving large DNA fragments that differ in length by only one nucleotide.
Advanced methods and de novo sequencing
Genomic DNA is fragmented into random pieces and cloned as a bacterial library. DNA from individual bacterial clones is sequenced and the sequence is assembled by using overlapping DNA regions.
Large-scale sequencing often aims at sequencing very long DNA pieces, such as whole chromosomes, although large-scale sequencing can also be used to generate very large numbers of short sequences, such as found in phage display. For longer targets such as chromosomes, common approaches consist of cutting (with restriction enzymes) or shearing (with mechanical forces) large DNA fragments into shorter DNA fragments. The fragmented DNA may then be cloned into a DNA vector and amplified in a bacterial host such as Escherichia coli. Short DNA fragments purified from individual bacterial colonies are individually sequenced and assembled electronically into one long, contiguous sequence.
The term "de novo sequencing" specifically refers to methods used to determine the sequence of DNA with no previously known sequence. De novo translates from Latin as "from the beginning". Gaps in the assembled sequence may be filled by primer walking. The different strategies have different tradeoffs in speed and accuracy; shotgun methods are often used for sequencing large genomes, but its assembly is complex and difficult, particularly with sequence repeats often causing gaps in genome assembly.
Most sequencing approaches use an in vitro cloning step to amplify individual DNA molecules, because their molecular detection methods are not sensitive enough for single molecule sequencing. Emulsion PCR isolates individual DNA molecules along with primer-coated beads in aqueous droplets within an oil phase. A polymerase chain reaction (PCR) then coats each bead with clonal copies of the DNA molecule followed by immobilization for later sequencing. Emulsion PCR is used in the methods developed by Marguilis et al. (commercialized by 454 Life Sciences), Shendure and Porreca et al. (also known as "Polony sequencing") and SOLiD sequencing, (developed by Agencourt, later Applied Biosystems, now Life Technologies).
Shotgun sequencing
Shotgun sequencing is a sequencing method designed for analysis of DNA sequences longer than 1000 base pairs, up to and including entire chromosomes. This method requires the target DNA to be broken into random fragments. After sequencing individual fragments, the sequences can be reassembled on the basis of their overlapping regions.
In genetics, shotgun sequencing, also known as shotgun cloning, is a method used for sequencing long DNA strands. It is named by analogy with the rapidly expanding, quasi-random firing pattern of a shotgun.
Since the chain termination method of DNA sequencing can only be used for fairly short strands (100 to 1000 basepairs), longer sequences must be subdivided into smaller fragments, and subsequently re-assembled to give the overall sequence. Two principal methods are used for this: chromosome walking, which progresses through the entire strand, piece by piece, and shotgun sequencing, which is a faster but more complex process, and uses random fragments.
In shotgun sequencing,DNA is broken up randomly into numerous small segments, which are sequenced using the chain termination method to obtain reads. Multiple overlapping reads for the target DNA are obtained by performing several rounds of this fragmentation and sequencing. Computer programs then use the overlapping ends of different reads to assemble them into a continuous sequence.
Shotgun sequencing was one of the precursor technologies that was responsible for enabling full genome sequencing.
Bridge PCR
Another method for in vitro clonal amplification is bridge PCR, in which fragments are amplified upon primers attached to a solid surface and form "DNA colonies" or "DNA clusters". This method is used in the Illumina Genome Analyzer sequencers. Single-molecule methods, such as that developed by Stephen Quake's laboratory (later commercialized by Helicos) are an exception: they use bright fluorophores and laser excitation to detect base addition events from individual DNA molecules fixed to a surface, eliminating the need for molecular amplification.
Next-generation methods
The high demand for low-cost sequencing has driven the development of high-throughput sequencing (or next-generation sequencing) technologies that parallelize the sequencing process, producing thousands or millions of sequences concurrently.High-throughput sequencing technologies are intended to lower the cost of DNA sequencing beyond what is possible with standard dye-terminator methods.In ultra-high-throughput sequencing as many as 500,000 sequencing-by-synthesis operations may be run in parallel.Multiple, fragmented sequence reads must be assembled together on the basis of their overlapping areas.
Next-generation (massively parallel, or second-generation) sequencing technologies have largely supplanted first-generation technologies. These newer approaches enable many DNA fragments (sometimes on the order of millions of fragments) to be sequenced at one time and are more cost-efficient and much faster than first-generation technologies. The utility of next-generation technologies was improved significantly by advances in bioinformatics that allowed for increased data storage and facilitated the analysis and manipulation of very large data sets, often in the gigabase range (1 gigabase = 1,000,000,000 base pairs of DNA).
Massively parallel signature sequencing (MPSS)
The first of the next-generation sequencing technologies, massively parallel signature sequencing (or MPSS), was developed in the 1990s at Lynx Therapeutics, a company founded in 1992 by Sydney Brenner and Sam Eletr. MPSS was a bead-based method that used a complex approach of adapter ligation followed by adapter decoding, reading the sequence in increments of four nucleotides. This method made it susceptible to sequence-specific bias or loss of specific sequences. Because the technology was so complex, MPSS was only performed 'in-house' by Lynx Therapeutics and no DNA sequencing machines were sold to independent laboratories. Lynx Therapeutics merged with Solexa (later acquired by Illumina) in 2004, leading to the development of sequencing-by-synthesis, a more simple approach acquired from Manteia Predictive Medicine, which rendered MPSS obsolete. However, the essential properties of the MPSS output were typical of later "next-generation" data types, including hundreds of thousands of short DNA sequences. In the case of MPSS, these were typically used for sequencing cDNA for measurements of gene expression levels.
Life Sciences Technology
A parallelized version of pyrosequencing was developed by 454 Life Sciences, which has since been acquired by Roche Diagnostics. The method amplifies DNA inside water droplets in an oil solution (emulsion PCR), with each droplet containing a single DNA template attached to a single primer-coated bead that then forms a clonal colony. The sequencing machine contains many picoliter-volume wells each containing a single bead and sequencing enzymes. Pyrosequencing uses luciferase to generate light for detection of the individual nucleotides added to the nascent DNA, and the combined data are used to generate sequence read-outs.This technology provides intermediate read length and price per base compared to Sanger sequencing on one end and Solexa and SOLiD on the other.
Illumina dye sequencing
Solexa, now part of Illumina, developed a sequencing method based on reversible dye-terminators technology, and engineered polymerases, that it developed internally.The terminated chemistry was developed internally at Solexa and the concept of the Solexa system was invented by Balasubramanian and Klennerman from Cambridge University's chemistry department. In 2004, Solexa acquired the company Manteia Predictive Medicine in order to gain a massivelly parallel sequencing technology based on "DNA Clusters", which involves the clonal amplification of DNA on a surface. The cluster technology was co-acquired with Lynx Therapeutics of California. Solexa Ltd. later merged with Lynx to form Solexa Inc.
Illumina dye sequencing is a technique used to determine the series of base pairs in DNA, also known as DNA sequencing. It was developed by researchers at Manteia Predictive Medicine (acquired by Solexa in 2004) and Solexa, a company later acquired by Illumina. This sequencing method is based on reversible dye-terminators that enable the identification of single bases as they are introduced into DNA strands. It is often employed to sequence difficult regions, such as homopolymers and repetitive sequences. It can also be used for whole-genome and region sequencing, transcriptome analysis, small RNA discovery, methylation profiling, and genome-wide protein-nucleic acid interaction analysis.
An Illumina MiSeq sequencer
In this method, DNA molecules and primers are first attached on a slide and amplified with polymerase so that local clonal DNA colonies, later coined "DNA clusters", are formed. To determine the sequence, four types of reversible terminator bases (RT-bases) are added and non-incorporated nucleotides are washed away. A camera takes images of the fluorescently labeled nucleotides, then the dye, along with the terminal 3' blocker, is chemically removed from the DNA, allowing for the next cycle to begin. Unlike pyrosequencing, the DNA chains are extended one nucleotide at a time and image acquisition can be performed at a delayed moment, allowing for very large arrays of DNA colonies to be captured by sequential images taken from a single camera.
Decoupling the enzymatic reaction and the image capture allows for optimal throughput and theoretically unlimited sequencing capacity. With an optimal configuration, the ultimately reachable instrument throughput is thus dictated solely by the analog-to-digital conversion rate of the camera, multiplied by the number of cameras and divided by the number of pixels per DNA colony required for visualizing them optimally (approximately 10 pixels/colony). In 2012, with cameras operating at more than 10 MHz A/D conversion rates and available optics, fluidics and enzymatics, throughput can be multiples of 1 million nucleotides/second, corresponding roughly to 1 human genome equivalent at 1x coverage per hour per instrument, and 1 human genome re-sequenced (at approx. 30x) per day per instrument (equipped with a single camera).
DNA NANOBALL
DNA nanoball sequencing is a type of high throughput sequencing technology used to determine the entire genomic sequence of an organism. The company Complete Genomics uses this technology to sequence samples submitted by independent researchers. The method uses rolling circle replication to amplify small fragments of genomic DNA into DNA nanoballs. Unchained sequencing by ligation is then used to determine the nucleotide sequence.This method of DNA sequencing allows large numbers of DNA nanoballs to be sequenced per run and at low reagent costs compared to other next generation sequencing platforms.However, only short sequences of DNA are determined from each DNA nanoball which makes mapping the short reads to a reference genome difficult.[55] This technology has been used for multiple genome sequencing projects and is scheduled to be used for more.
Heliscope single molecule sequencing
HEliscope sequencing is a method of single-molecule sequencing developed by Helicos Biosciences. It uses DNA fragments with added poly-A tail adapters which are attached to the flow cell surface. The next steps involve extension-based sequencing with cyclic washes of the flow cell with fluorescently labeled nucleotides (one nucleotide type at a time, as with the Sanger method). The reads are performed by the Heliscope sequencer. The reads are short, up to 55 bases per run, but recent improvements allow for more accurate reads of stretches of one type of nucleotides.
This sequencing method and equipment were used to sequence the genome of the M13 bacteriophage.
Helicos™ Single Molecule Sequencing (SMS) provides a unique view of genome biology through direct sequencing of cellular nucleic acids in an unbiased manner, providing both accurate quantitation and sequence information. Sample preparation does not require ligation or PCR amplification, avoiding the GC-content and size biases observed in other technologies. DNA is simply sheared, tailed with poly(A), and hybridized to a flow cell surface containing oligo(dT) for sequencing-by-synthesis of billions of molecules in parallel. This process also requires far less material than other technologies. Gene expression measurements can be done using first-strand cDNA-based methods (RNA-Seq) or using a novel approach that allows direct hybridization and sequencing of cellular RNA for the most direct quantitation possible. In this unit, principles and methods for using the Helicos® Genetic Analysis System are discussed.
Single molecule real time (SMRT) sequencing
SMRT sequencing is based on the sequencing by synthesis approach. The DNA is synthesized in zero-mode wave-guides (ZMWs) – small well-like containers with the capturing tools located at the bottom of the well. The sequencing is performed with use of unmodified polymerase (attached to the ZMW bottom) and fluorescently labelled nucleotides flowing freely in the solution. The wells are constructed in a way that only the fluorescence occurring by the bottom of the well is detected. The fluorescent label is detached from the nucleotide at its incorporation into the DNA strand, leaving an unmodified DNA strand. According to Pacific Biosciences, the SMRT technology developer, this methodology allows detection of nucleotide modifications (such as cytosine methylation). This happens through the observation of polymerase kinetics. This approach allows reads of 20,000 nucleotides or more, with average read lengths of 5 kilobases.
Methods in development
DNA sequencing methods currently under development include labeling the DNA polymerase,reading the sequence as a DNA strand transits through nanopores,and microscopy-based techniques, such as atomic force microscopy or transmission electron microscopy that are used to identify the positions of individual nucleotides within long DNA fragments (>5,000 bp) by nucleotide labeling with heavier elements (e.g., halogens) for visual detection and recording.Third generation technologies aim to increase throughput and decrease the time to result and cost by eliminating the need for excessive reagents and harnessing the processivity of DNA polymerase.
Nanopore sequencing
This method is based on the readout of electrical signals occurring at nucleotides passing by alpha-hemolysin pores covalently bound with cyclodextrin. The DNA passing through the nanopore changes its ion current. This change is dependent on the shape, size and length of the DNA sequence. Each type of the nucleotide blocks the ion flow through the pore for a different period of time. The method has a potential of development as it does not require modified nucleotides, however single nucleotide resolution is not yet available.
DNA could be passed through the nanopore for various reasons. For example, electrophoresis might attract the DNA towards the nanopore, and it might eventually pass through it. Or, enzymes attached to the nanopore might guide DNA towards the nanopore. The scale of the nanopore means that the DNA may be forced through the hole as a long string, one base at a time, rather like thread through the eye of a needle. As it does so, each nucleotide on the DNA molecule may obstruct the nanopore to a different, characteristic degree. The amount of current which can pass through the nanopore at any given moment therefore varies depending on whether the nanopore is blocked by an A, a C, a G or a T. The change in the current through the nanopore as the DNA molecule passes through the nanopore represents a direct reading of the DNA sequence. Alternatively, a nanopore might be used to identify individual DNA bases as they pass through the nanopore in the correct order - this approach has been shown by Oxford Nanopore Technologies and Professor Hagan Bayley.
The potential is that a single molecule of DNA can be sequenced directly using a nanopore, without the need for an intervening PCR amplification step or a chemical labelling step or the need for optical instrumentation to identify the chemical label. As of July 2010, information available to the public indicates that nanopore sequencing is still in the development stage, with some laboratory-based data to back up the different components of the sequencing method. Despite these advancements, nanopore sequencing is not currently commercially available, parallelized, routineized, nor cost-effective enough to compete with "next generation sequencing" methods. Nanopore-based DNA analysis techniques are being industrially developed by Oxford Nanopore Technologies (developing direct exonuclease sequencing and strand sequencing using protein nanopores, and solid-state sequencing through internal R&D and collaborations with academic institutions), NabSys (using a library of DNA probes and using nanopores to detect where these probes have hybridized to single stranded DNA) and NobleGen(using nanopores in combination with fluorescent labels). IBM has noted research projects on computer simulations of translocation of a DNA strand through a solid-state nanopore, but not projects on identifying the DNA bases on that strand.
Tunnelling currents DNA sequencing
Another approach uses measurements of the electrical tunnelling currents across single-strand DNA as it moves through a channel. Depending on its electronic structure each base affects the tunnelling current differently, allowing statistical differentiation between different bases.
The use of tunnelling currents has the potential to sequence orders of magnitude faster than ionic current methods and the sequencing of several DNA oligomers and micro-RNA has already been achieved.
It has been proposed that single molecules of DNA could be sequenced by measuring the physical properties of the bases as they pass through a nanopore1, 2. Theoretical calculations suggest that electron tunnelling can identify bases in single-stranded DNA without enzymatic processing3, 4, 5, and it was recently experimentally shown that tunnelling can sense individual nucleotides6 and nucleosides7. Here, we report that tunnelling electrodes functionalized with recognition reagents can identify a single base flanked by other bases in short DNA oligomers. The residence time of a single base in a recognition junction is on the order of a second, but pulling the DNA through the junction with a force of tens of piconewtons would yield reading speeds of tens of bases per second.
Development initiatives
In October 2006, the X Prize Foundation established an initiative to promote the development of full genome sequencing technologies, called the Archon X Prize, intending to award $10 million to "the first Team that can build a device and use it to sequence 100 human genomes within 10 days or less, with an accuracy of no more than one error in every 100,000 bases sequenced, with sequences accurately covering at least 98% of the genome, and at a recurring cost of no more than $10,000 (US) per genome."
Each year the National Human Genome Research Institute, or NHGRI, promotes grants for new research and developments in genomics. 2010 grants and 2011 candidates include continuing work in microfluidic, polony and base-heavy sequencing methodologies.As DNA Analysis has become more prevalent and necessary in modern time, so have methods of sequencing.Genome.gov holds programs and grants for those with the most rewarding possible future methods of DNA Sequencing. Below are a few that have some great potential.
Microfluidic DNA Sequencing
A single molecule detection method leveraging droplet-based microfluidics. This should limit the amount of reagent required to sequence DNA to less than several milliliters, while still retaining the ability to amplify the template that thereby enables us to use relatively inexpensive and robust detection. The method is simple and does not require enzymes.
Into a droplet library of DNA probe strands, each containing a set of fluorescent dyes at discrete concentration levels that are used as fluorescent bar codes, the unknown DNA template strand is injected. If the complementary sequence of the probe strand is found in the template, it can hybrizide to the template. By detecting the hybridization event and at the same time the fluorescent barcode label, it is possible to obtain information about the template sequence. By systematically probing all possible probe strands, i.e. sequences of the four bases (A,C,G,T), the sequence of the template strand can be reconstructed.
Related side projects I am working on include the parallelized encapsulation of samples to generate large droplet libraries and fluorescent barcode labels to identify droplets.
In a custom microscope setup, we focus a 488 nm laser into a flow channel where droplets are detected. When flowing by the laser spot, fluorescent dyes inside the droplets are excited, the fluorescence emission is collected through the same objective of the microscope and guided through a series of dichroic mirrors. Currently we have three detection channels for three different fluorescent species that we can detect simultaneously in each droplet. The same detection principle is used in flow cytometers.
Polony Sequencing
Ultra-high throughput polony genome sequencing, generating raw data to re-sequence the human genome in one week. The goal of the method is to increase the polony sequencing read length using a cyclic ligation strategy that involves enzymatic cleavage, and increase read density by using different clonal amplification strategies.
Polony sequencing is an inexpensive but highly accurate multiplex sequencing technique that can be used to “read” millions of immobilized DNA sequences in parallel. This technique was first developed by Dr. George Church's group at Harvard Medical School. Unlike other sequencing techniques, Polony sequencing technology is an open platform with freely downloadable, open source software and protocols. Also, the hardware of this technique can be easily set up with a commonly available epifluorescence microscopy and a computer-controlled flowcell/fluidics system. Polony sequencing is generally performed on paired-end tags library that each molecule of DNA template is of 135 bp in length with two 17–18 bp paired genomic tags separated and flanked by common sequences. The current read length of this technique is 26 bases per amplicon and 13 bases per tag, leaving a gap of 4–5 bases in each tag.
The Polony sequencing method, developed in the laboratory of George M. Church at Harvard, was among the first next-generation sequencing systems and was used to sequence a full genome in 2005. It combined an in vitro paired-tag library with emulsion PCR, an automated microscope, and ligation-based sequencing chemistry to sequence an E. coli genome at an accuracy of >99.9999% and a cost approximately 1/9 that of Sanger sequencing.The technology was licensed to Agencourt Biosciences, subsequently spun out into Agencourt Personal Genomics, and eventually incorporated into the Applied Biosystems SOLiD platform, which is now owned by Life Technologies.
Polony sequencing is an inexpensive but highly accurate multiplex sequencing technique that can be used to “read” millions of immobilized DNA sequences in parallel. This technique was first developed by Dr. George Church's group at Harvard Medical School. Unlike other sequencing techniques, Polony sequencing technology is an open platform with freely downloadable, open source software and protocols. Also, the hardware of this technique can be easily set up with a commonly available epifluorescence microscopy and a computer-controlled flowcell/fluidics system. Polony sequencing is generally performed on paired-end tags library that each molecule of DNA template is of 135 bp in length with two 17–18 bp paired genomic tags separated and flanked by common sequences. The current read length of this technique is 26 bases per amplicon and 13 bases per tag, leaving a gap of 4–5 bases in each tag.
Millikan Sequencing by Nucleotide
This novel sequencing-by-synthesis approach measures the increased charge as nucleotides are added to DNA templates attached to a tethered bead. Opposing electrical, hydrodynamic and entropic forces will be used to measure the bead displacement, which is a function of the length of DNA attached to the bead. The much lower per-bead copy number required compared to the 454 system should enable amplification options other than emulsion PCR, such as bridge PCR, making initial sample preparation easier and cheaper.
A nucleic acid sequence is a succession of letters that indicate the order of nucleotides within a DNA (using GACT) or RNA (GACU) molecule. By convention, sequences are usually presented from the 5' end to the 3' end. Because nucleic acids are normally linear (unbranched) polymers, specifying the sequence is equivalent to defining the covalent structure of the entire molecule. For this reason, the nucleic acid sequence is also termed the primary structure.
The sequence has capacity to represent information. Biological DNA represents the information which directs the functions of a living thing. In that context, the term genetic sequence is often used. Sequences can be read from the biological raw material through DNA sequencing methods.
Nucleic acids also have a secondary structure and tertiary structure. Primary structure is sometimes mistakenly referred to as primary sequence. Conversely, there is no parallel concept of secondary or tertiary sequence.
Single-Molecule DNA Sequencing with Engineered Nanopores
In nanopore strand sequencing, a single strand of DNA moves through a narrow pore and the bases are identified as they pass a reading head. Here, we focus on the remaining tasks required to put into practice strand sequencing with the ±-hemolysin (±HL) protein nanopore. Nanopore sequencing is a rapid real-time technology; it does not require the time-consuming cyclic addition of reagents. After implementing a chip with 106 pores, we expect nanopore sequencing to achieve a 15-minute genome by 2014 with a very short sample preparation time. In addition, nanopore sequencing will be able to identify modified bases and to sequence RNA directly. The latest goal is to refine base recognition by using ±HL nanopores, engineered by conventional mutagenesis, unnatural amino acid mutagenesis and targeted chemical modification, to produce DNA reading heads fit for real-time sequencing.
Direct Real-time Single Molecule DNA Sequencing
Direct real-time sequencing of single DNA molecules from genomic DNA at the speed and accuracy of the natural DNA polymerases using native nucleotides. Unlike the difficult to engineer man-made nanostructures used in nanopore sequencing to distinguish the 4 base types in close proximity and constant fluctuation, DNA polymerases have precise atomic-resolution 3D structures and can synthesize very long DNA molecules with high fidelity and velocity. The strategy is to engineer sensors onto the surface of the polymerase by protein engineering to monitor the subtle yet distinct conformational changes accompanying the incorporation of each base type.
Single molecule real time sequencing (also known as SMRT) is a parallelized single molecule DNA sequencing by synthesis technology developed by Pacific Biosciences. Single molecule real time sequencing utilizes the zero-mode waveguide (ZMW), developed in the laboratories of Harold G. Craighead and Watt W. Webb at Cornell University. A single DNA polymerase enzyme is affixed at the bottom of a ZMW with a single molecule of DNA as a template. The ZMW is a structure that creates an illuminated observation volume that is small enough to observe only a single nucleotide of DNA (also known as a base) being incorporated by DNA polymerase. Each of the four DNA bases is attached to one of four different fluorescent dyes. When a nucleotide is incorporated by the DNA polymerase, the fluorescent tag is cleaved off and diffuses out of the observation area of the ZMW where its fluorescence is no longer observable. A detector detects the fluorescent signal of the nucleotide incorporation, and the base call is made according to the corresponding fluorescence of the dye. Sequence data generated from single molecule real time sequencing was first published in January 2009 in the journal Science.
Tunnel Junction for Reading All Four Bases with High Discrimination
Distinct tunneling signals can be generated for all four nucleosides (and 5-methyldeoxycytidine) using one pair of tunneling electrodes functionalized with a simple reagent containing a hydrogen-bond donor and a hydrogen bond acceptor. The goals of this proposal are to extend the measurements to nucleotides in aqueous electrolyte, and then to small oligomers.
We have discovered that distinct tunneling signals can be generated for all four nucleosides (and 5-methyldeoxycytidine) using one pair of tunneling electrodes functionalized with a simple reagent containing a hydrogen-bond donor and a hydrogen bond acceptor. The goals of this proposal are to extend the measurements to nucleotides in aqueous electrolyte, and then to small oligomers. We will quantify the fraction of single-molecule reads and determine the factors that control this fraction with the goal of eliminating signals that come from more than one nucleotide in the gap at a time. We will explore the factors that control the width of the distribution of current signals for all four bases (and 5-methyl C) with the goal of improving the discrimination of a single read. We will measure the fraction of successful reads and characterize the time required for the complex (that gives rise to the signal) to form in the tunnel gap. From these measurements, we will identify improvements needed to increase the readout efficiency and also develop criteria for design of a nanopore sequencing system equipped with tunneling electrodes. The reagents developed during the course of this research will be made available to other research groups developing nanopore sequencers that use electron tunneling as the readout. At least seven NIH-supported groups are exploring sequencing methods that propose to use electron tunneling as the readout for a nanopore sequencer, an approach that might greatly reduce the cost of sequencing. We have shown that all four nucleosides and 5-methyl cytidine can be read by functionalized electrodes and we will develop reagents suitable for DNA sequencing in aqueous electrolyte and make these widely available.
Single Molecule Sequencing by Nanopore-induced Photon Emission (SM-SNIPE)
Nanopore induced photon emission (SNIPE), utilizes optical detection rather than the more ubiquitous electrical detection. Dramatically increase the throughput, speed and accuracy of SNIPE. Develop and optimize our proprietary DNA conversion approach, Circular DNA conversion (CDC).
Nanopore-based DNA analysis is an extremely attractive area of research due to the simplicity of the method, and the ability to not only probe individual molecules, but also to detect very small amounts of genomic material. Here, we describe the materials and methods of a novel, nanopore-based, single-molecule DNA sequencing system that utilizes optical detection. We convert target DNA according to a binary code, which is recognized by molecular beacons with two types of fluorophores. Solid-state nanopores are then used to sequentially strip off the beacons, leading to a series of photon bursts that can be detected with a custom-made microscope. We do not use any enzymes in the readout stage; thus, our system is not limited by the highly variable processivity, lifetime, and inaccuracy of individual enzymes that can hinder throughput and reliability. Furthermore, because our system uses purely optical readout, we can take advantage of high-end, wide-field imaging devices to record from multiple nanopores simultaneously. This allows an extremely straightforward parallelization of our system to nanopore arrays.
Modeling Macromolecular Transport for Sequencing Technologies
In Nanopore-based electrophoretic experiments, translocation of single molecules of DNA is monitored as they pass through protein channels and solid-state nanopores under an external electric field. This proposed method deals with a fundamental understanding of the behavior of DNA in nanopore environments under the influence of electrical and hydrodynamic forces.
Base-selective Heavy Atom Labels for Electron Microscopy-based DNA Sequencing
Since efficient electron scattering to a detector is highly dependent on atomic number (Z), it is possible to label single stranded DNA (ssDNA) with heavy atoms. To test the limits of this trend, this method proposes a multipronged approach to selectively prepared metal-DNA base pair complexes, focusing on the selective labeling of DNA bases and the development of an appropriate assay to evaluate our success.
Transmission electron microscopy DNA sequencing is an emerging third-generation, single-molecule sequencing technology that uses transmission electron microscopy techniques. DNA is visible under the electron microscope; however, it must be labeled with heavy atoms so that the DNA bases can be clearly visualized. In addition, specialized imaging techniques and aberration corrected optics are beneficial for obtaining the resolution required to image the labeled DNA molecule. Transmission electron microscopy DNA sequencing advantageously may provide extremely long read lengths, but it is not yet commercially available.
The Need for DNA Sequencing
The process of DNA sequencing translates the DNA of a specific organism into a format that is decipherable by researchers and scientists. DNA sequencing has given a massive boost to numerous fields such as forensic biology, biotechnology and more. By mapping the basic sequence of nucleotides, DNA sequencing has allowed scientists to better understand genes and their role in the creation of the human body.
Forensic biology uses DNA sequences to identify the organism which it is unique to. Although identifying an individual is less accurate currently, but as the processes evolves further, direct comparisons of large DNA segments, and maybe even genomes, will be more practical and viable and will allow precise identification of an individual. Scientists will be able to isolate the genes responsible for genetic diseases like Cystic Fibrosis, Alzheimer’s disease, myotonic dystrophy, etc., which are caused by the inability of genes to work properly.
Agriculture has been helped immensely by DNA sequencing. It has allowed scientists to make plants more resistant to insects and pests, by understanding their genes. Likewise, the same technique has been utilized to increase the productivity and quality of the milk, as well as the meat, produced by livestock.
Applications of DNA sequencing technologies
Knowledge of the sequence of a DNA segment has many uses. First, it can be used to find genes, segments of DNA that code for a specific protein or phenotype. If a region of DNA has been sequenced, it can be screened for characteristic features of genes. For example, open reading frames (ORFs)—long sequences that begin with a start codon (three adjacent nucleotides; the sequence of a codon dictates amino acid production) and are uninterrupted by stop codons (except for one at their termination)—suggest a protein-coding region. Also, human genes are generally adjacent to so-called CpG islands—clusters of cytosine and guanine, two of the nucleotides that make up DNA. If a gene with a known phenotype (such as a disease gene in humans) is known to be in the chromosomal region sequenced, then unassigned genes in the region will become candidates for that function. Second, homologous DNA sequences of different organisms can be compared in order to plot evolutionary relationships both within and between species. Third, a gene sequence can be screened for functional regions. In order to determine the function of a gene, various domains can be identified that are common to proteins of similar function. For example, certain amino acid sequences within a gene are always found in proteins that span a cell membrane; such amino acid stretches are called transmembrane domains. If a transmembrane domain is found in a gene of unknown function, it suggests that the encoded protein is located in the cellular membrane. Other domains characterize DNA-binding proteins. Several public databases of DNA sequences are available for analysis by any interested individual.
The applications of next-generation sequencing technologies are vast, owing to their relatively low cost and large-scale high-throughput capacity. Using these technologies, scientists have been able to rapidly sequence entire genomes (whole genome sequencing) of organisms, to discover genes involved in disease, and to better understand genomic structure and diversity among species generally.
The Future
DNA sequencing, in time, may allow us to manipulate the process of evolution. Diseases will vanish, and the human race will be stronger, smarter and better than ever before.