Purpose and scientific scope
DNA sequencing is the experimental process of determining the order of nucleotides in a DNA molecule. Modern sequencing is not only a laboratory technique; it is an integrated workflow that combines sample preparation, chemical detection, instrument engineering, and bioinformatic analysis. Next-generation sequencing (NGS) usually refers to massively parallel short-read technologies, although the broader field now also includes long-read single-molecule platforms.
From Sanger sequencing to massively parallel sequencing
Sanger sequencing, introduced in the 1970s, uses DNA polymerase, a primer, normal deoxynucleotides, and chain-terminating dideoxynucleotides. The incorporation of a dideoxynucleotide stops extension, generating fragments of different lengths. Modern capillary electrophoresis separates these fragments and detects fluorescent labels to reconstruct the sequence. Sanger sequencing remains highly useful for validating plasmids, PCR products, and selected genomic loci, but it is low-throughput compared with NGS.
NGS changed the scale of sequencing by reading millions to billions of DNA fragments in parallel. Instead of producing one sequence read at a time, NGS creates a library of fragments and reads them simultaneously. The result is a large digital dataset that must be processed computationally before it becomes biologically meaningful.
Core NGS workflow
A typical NGS experiment starts with nucleic-acid extraction. DNA sequencing workflows fragment genomic DNA and ligate adapters to create a sequencing library. RNA-seq workflows usually convert RNA into complementary DNA (cDNA) before library construction. Library quality, fragment-size distribution, concentration, and contamination are critical because poor library preparation cannot be fully corrected by later bioinformatics.
Many short-read platforms amplify library molecules before sequencing. Illumina systems use bridge amplification on a flow cell to generate clusters of identical DNA fragments. Other historical platforms used emulsion PCR on beads. Long-read technologies such as PacBio and Oxford Nanopore can read single molecules without clonal amplification, although library preparation remains essential.
After sequencing, the instrument produces raw signal data that are converted into base calls and quality scores. The most common first output is FASTQ. Downstream analysis may include quality control, adapter trimming, alignment to a reference genome, de novo assembly, variant calling, quantification, taxonomic profiling, or expression analysis depending on the study design.
Main sequencing chemistries and signals
Applications
NGS is used for whole-genome sequencing, targeted gene panels, exome sequencing, RNA-seq, metagenomics, pathogen surveillance, cancer genomics, population genetics, epigenomic assays, and ancient-DNA studies. In clinical and public-health settings, sequencing can help identify pathogens, track outbreaks, detect antimicrobial-resistance genes, characterize inherited variants, and support precision oncology. Interpretation must always consider sample quality, coverage depth, reference bias, contamination, and validation requirements.
Limitations and quality-control principles
NGS is powerful but not error-free. Short-read sequencing can struggle with repetitive regions, large structural variants, highly homologous gene families, and complex genomic rearrangements. Long-read sequencing improves structural resolution but requires careful error correction, coverage planning, and computational resources. For scientific reliability, every study should report library strategy, read length, depth or coverage, reference genome version, quality filters, software versions, and validation method where appropriate.
Key scientific takeaway
NGS should be understood as a complete experimental and computational system. The biological conclusion is only as strong as the sample quality, library design, sequencing chemistry, coverage, bioinformatic pipeline, and validation strategy.
Main sequencing approaches
| Approach | Main signal/principle | Strength | Important limitation |
| Sanger sequencing | Chain termination and electrophoretic separation | High accuracy for short targets | Low throughput and expensive per base |
| Illumina SBS | Fluorescent reversible terminators imaged cycle by cycle | High throughput and high short-read accuracy | Short reads limit repeat and structural-variant resolution |
| Ion Torrent | pH changes from hydrogen-ion release during synthesis | No optical imaging and relatively simple detection | Homopolymer length errors can be problematic |
| PacBio SMRT | Single-molecule real-time fluorescent detection | Long reads and useful consensus accuracy | Higher input and analysis requirements |
| Oxford Nanopore | Ionic-current disruption as molecules pass through nanopores | Very long reads, portability, direct DNA/RNA signals | Raw-read error profile and data handling require care |
References
- Sanger, F., Nicklen, S., & Coulson, A. R. (1977). DNA sequencing with chain-terminating inhibitors. Proceedings of the National Academy of Sciences, 74(12), 5463-5467.
- Mardis, E. R. (2008). Next-generation DNA sequencing methods. Annual Review of Genomics and Human Genetics, 9, 387-402.
- Goodwin, S., McPherson, J. D., & McCombie, W. R. (2016). Coming of age: Ten years of next-generation sequencing technologies. Nature Reviews Genetics, 17, 333-351.
- Sequencing by synthesis technology overview. https://www.illumina.com/science/technology/next-generation-sequencing/sequencing-technology.html
- Logsdon, G. A., Vollger, M. R., & Eichler, E. E. (2020). Long-read human genome sequencing and its applications. Nature Reviews Genetics, 21, 597-614.