HomeTechniquesNext Generation Sequencing(NGS): Principles, Workflow & Applications

Next Generation Sequencing(NGS): Principles, Workflow & Applications

- Advertisement -

Next generation sequencing (NGS) is a group of advanced DNA and RNA sequencing technologies that allow researchers to analyze millions of nucleic acid fragments simultaneously. Unlike traditional Sanger sequencing, which generally analyzes one DNA fragment at a time, NGS uses massively parallel sequencing to generate large amounts of genetic information in a single experiment.

The development of NGS has transformed genomics and molecular biology. Researchers can now sequence entire genomes, analyze protein-coding regions, investigate gene expression, identify genetic variants, and study cancer-associated mutations with much greater throughput than was possible with conventional sequencing methods.

NGS is now widely used in cancer research, genetic disease diagnosis, infectious disease research, transcriptomics, pharmacogenomics, precision medicine, and biomarker discovery. Depending on the experimental objective, researchers can sequence a small group of selected genes or investigate nearly all of the genetic material in a biological sample.

Although NGS platforms differ in their chemistry and sequencing mechanisms, most workflows follow several fundamental steps, including nucleic acid extraction, library preparation, sequencing, signal detection, and bioinformatic analysis.

This guide explains what next generation sequencing is, how it works, the major sequencing technologies, its applications, and its main advantages and limitations.

What Is Next Generation Sequencing?

Next generation sequencing is a high-throughput sequencing approach that determines the nucleotide sequence of large numbers of DNA molecules simultaneously.

The term “next generation” was introduced to distinguish these technologies from first-generation sequencing methods, particularly Sanger sequencing.

Sanger sequencing remains an important technique and is highly useful when researchers need to sequence a relatively small number of DNA fragments. However, sequencing large numbers of genes or entire genomes using Sanger sequencing would require substantial time and resources.

NGS addresses this limitation by using massively parallel sequencing.

Instead of sequencing one DNA fragment at a time, an NGS instrument can process millions of DNA fragments during a single sequencing run. Each fragment produces sequencing information that can subsequently be analyzed computationally.

From Sanger sequencing to NGS

Sanger sequencing relies on chain-termination chemistry to determine the sequence of a DNA fragment. It is characterized by relatively long and accurate reads but has limited throughput compared with modern NGS technologies.

NGS introduced a fundamentally different approach based on parallelization.

A DNA sample is first fragmented into many smaller pieces. These fragments are then prepared as a sequencing library and simultaneously processed by the sequencing instrument.

The resulting sequences are called reads.

Bioinformatics software subsequently processes these reads and may align them to a reference genome, assemble them into longer sequences, or compare them with other samples.

What can be sequenced using NGS?

NGS can be applied to different types of sequencing projects.

Whole-genome sequencing (WGS) analyzes genomic DNA across essentially the entire genome.

Whole-exome sequencing (WES) focuses on the exons, which represent the protein-coding portions of genes.

Targeted sequencing analyzes selected genes or genomic regions. This approach is particularly useful when researchers are interested in a predefined group of genes, such as cancer-associated genes.

NGS can also be used for RNA sequencing (RNA-seq). In this case, RNA molecules are converted into a sequencing-compatible library, allowing researchers to investigate gene expression and transcriptome composition.

The appropriate approach depends on the biological question, sample type, required sequencing depth, and available resources.

How Does Next Generation Sequencing Work?

Although different sequencing platforms use different chemistries, a typical next generation sequencing workflow can be divided into several major stages.

1. Sample preparation and nucleic acid extraction

The process begins with a biological sample.

Depending on the experiment, the starting material may include blood, tissue, cultured cells, microorganisms, tumor samples, or other biological specimens.

DNA or RNA is extracted from the sample using an appropriate extraction method.

Read our full guide on DNA extraction.

Read our full guide on RNA extraction.

The quality and quantity of the extracted nucleic acid are important because poor-quality starting material can affect subsequent library preparation and sequencing performance.

For DNA sequencing, researchers generally evaluate factors such as DNA concentration, purity, integrity, and fragment size.

For RNA sequencing, RNA integrity is particularly important because RNA can degrade relatively easily.

2. DNA fragmentation

For many NGS workflows, genomic DNA is fragmented into smaller pieces.

Fragmentation can be performed using physical or enzymatic methods.

The desired fragment size depends on the sequencing platform and experimental design.

After fragmentation, the DNA consists of many individual fragments representing different regions of the original genome.

Not every NGS workflow requires fragmentation in the same way. Some long-read sequencing approaches are designed to preserve much longer DNA molecules.

3. Library preparation

One of the most important stages of NGS is library preparation.

An NGS library is a collection of DNA fragments that have been modified so they can be recognized and processed by the sequencing platform.

During library preparation, specialized sequences called adapters are attached to the ends of DNA fragments.

These adapters serve several purposes.

They can allow DNA fragments to bind to the sequencing surface, provide sequences required for amplification or sequencing, and contain sample-specific indexes.

Indexes are particularly useful when several samples are sequenced together in the same run. This process is known as multiplexing.

After sequencing, bioinformatics software can use the index sequences to determine which reads originated from which sample.

4. Library amplification

Depending on the sequencing technology and workflow, the prepared library may undergo amplification.

Amplification increases the amount of sequencing material available for detection.

For example, in some short-read sequencing systems, individual DNA fragments are amplified into clusters on a sequencing surface.

However, amplification is not universal across all modern sequencing technologies. Some single-molecule sequencing approaches can sequence individual DNA molecules directly.

5. Sequencing

The prepared library is introduced into the sequencing instrument.

The exact sequencing mechanism depends on the platform.

In sequencing-by-synthesis approaches, the instrument detects the incorporation of nucleotides as a complementary DNA strand is synthesized.

Other platforms use different mechanisms.

For example, semiconductor sequencing detects changes associated with nucleotide incorporation, while nanopore sequencing measures changes in electrical current as nucleic acid molecules pass through nanopores.

The result is a large collection of sequence reads.

6. Signal detection and base calling

The sequencing instrument detects a physical or chemical signal associated with nucleotide incorporation or movement.

These signals are converted into nucleotide sequences through a process known as base calling.

The output contains sequence information as strings of nucleotides such as:

A, T, C, and G

for DNA.

In RNA-related workflows, the original RNA is often converted into complementary DNA before sequencing.

The quality of each base call can also be estimated. Sequencing quality scores provide information about the confidence associated with individual nucleotide calls.

7. Bioinformatics analysis

Sequencing generates enormous quantities of data, so computational analysis is an essential part of NGS.

Raw sequencing reads can undergo several processing steps, including:

  • Quality control
  • Adapter trimming
  • Removal of low-quality sequences
  • Alignment to a reference genome
  • Read assembly
  • Variant calling
  • Gene expression analysis
  • Annotation
  • Statistical analysis

For example, in a DNA sequencing experiment designed to identify mutations, sequencing reads may first be aligned to a reference genome.

Software can then identify differences between the patient’s sequence and the reference sequence.

These differences may include single-nucleotide variants, insertions, deletions, and other genomic alterations.

Sequencing depth and coverage

Two important concepts in NGS are sequencing depth and coverage.

Sequencing depth refers broadly to how many sequencing reads support a particular position or region.

Higher depth can increase the ability to detect variants, particularly variants present at low frequencies.

Coverage describes how much of the target genomic region has been successfully sequenced.

For example, a sequencing experiment may have high average depth but still contain genomic regions with insufficient coverage.

Therefore, evaluating both sequencing depth and coverage is important when interpreting NGS results.

Major Next Generation Sequencing Technologies and Platforms

NGS is not a single technology. Several sequencing platforms use different approaches to determine nucleotide sequences.

Illumina sequencing

Illumina sequencing is one of the most widely used short-read sequencing approaches.

It primarily uses a sequencing-by-synthesis mechanism.

DNA fragments are attached to a sequencing surface and amplified into clusters. During sequencing, nucleotides are incorporated into newly synthesized DNA strands, and the instrument detects signals associated with nucleotide incorporation.

Illumina platforms are widely used for:

  • Whole-genome sequencing
  • Whole-exome sequencing
  • Targeted sequencing
  • RNA sequencing
  • Cancer genomic profiling
  • Microbial genomics

Their combination of high throughput and high accuracy has made short-read sequencing particularly useful for many research and clinical applications.

Ion Torrent sequencing

Ion Torrent technology uses a different approach known as semiconductor sequencing.

Instead of detecting fluorescent signals, the system detects changes associated with the release of hydrogen ions during nucleotide incorporation.

This allows DNA sequence information to be converted into an electrical signal.

Ion Torrent systems have been used for applications such as targeted sequencing, cancer panels, microbial sequencing, and other genomic studies.

Oxford Nanopore sequencing

Oxford Nanopore technology uses nanopores to analyze individual DNA or RNA molecules.

As a nucleic acid molecule passes through a nanopore, it causes changes in electrical current. These changes can be interpreted to determine the nucleotide sequence.

A major characteristic of nanopore sequencing is its ability to generate very long reads.

Long reads can be particularly useful for studying:

  • Structural variants
  • Repetitive genomic regions
  • Large genomic rearrangements
  • Complex genomic regions
  • Full-length transcripts

Nanopore sequencing can also provide real-time sequencing data, making it useful for applications where rapid analysis is important.

PacBio sequencing

PacBio sequencing is another important long-read technology.

It uses single-molecule real-time sequencing, allowing individual DNA molecules to be observed during synthesis.

Long-read sequencing can help researchers resolve genomic regions that are difficult to analyze using short reads.

PacBio technologies have applications in genome assembly, structural variant detection, transcriptome analysis, and characterization of complex genomic regions.

Short-read vs. long-read sequencing

One important distinction between sequencing technologies is read length.

Short-read technologies generate relatively short sequence fragments but can provide high accuracy and high throughput.

Long-read technologies generate substantially longer sequences, which can make it easier to resolve repetitive sequences, structural variations, and complex genomic arrangements.

Neither approach is universally suitable for every experiment.

The choice depends on the research question, genome complexity, required accuracy, available sequencing depth, turnaround time, and budget.

Applications of Next Generation Sequencing

The ability to analyze large quantities of genetic information has made NGS an important tool across biomedical research and clinical genomics.

NGS in cancer research

Cancer research is one of the major applications of next generation sequencing.

Cancer cells can accumulate genetic alterations that influence tumor development, progression, and response to treatment.

NGS allows researchers to investigate these alterations at high throughput.

For example, targeted cancer sequencing panels can analyze genes commonly associated with specific cancer types.

Researchers can identify mutations in oncogenes and tumor suppressor genes and investigate other genomic alterations.

NGS can also be used to study tumor heterogeneity, in which different populations of cancer cells within the same tumor may carry different genetic alterations.

This information can contribute to the identification of potential biomarkers and therapeutic targets.

Whole-genome sequencing

Whole-genome sequencing provides a comprehensive view of genomic DNA.

It can be used to investigate coding and non-coding regions of the genome and can identify a broad range of genetic variants.

WGS is particularly valuable in research because it provides information beyond predefined gene panels.

However, the large amount of data generated also increases computational and interpretive requirements.

Whole-exome sequencing

Whole-exome sequencing focuses primarily on protein-coding regions.

Because the exome represents only a fraction of the entire genome, WES can provide a more focused alternative to WGS when the research objective is to identify coding variants.

WES has been widely used in genetic disease research and cancer genomics.

RNA sequencing

RNA sequencing (RNA-seq) uses NGS to investigate the transcriptome.

Researchers can use RNA-seq to determine which genes are expressed in a biological sample and compare expression patterns between conditions.

Applications include:

  • Gene expression profiling
  • Transcript discovery
  • Alternative splicing analysis
  • Cancer transcriptomics
  • Biomarker discovery
  • Investigation of cellular responses

RNA-seq therefore provides information about gene activity rather than simply identifying the DNA sequence.

Genetic disease research

NGS has become an important tool for investigating inherited genetic disorders.

Instead of analyzing individual genes one at a time, researchers can use targeted panels, WES, or WGS to investigate multiple genes simultaneously.

This can be particularly useful for genetically heterogeneous diseases, where similar clinical features can result from variants in different genes.

Infectious disease research

NGS can also be used to sequence microbial genomes and investigate infectious diseases.

Researchers can characterize bacterial, viral, fungal, or other microbial genomes and examine genetic variation between isolates.

Sequencing can help investigate pathogen evolution, transmission patterns, antimicrobial resistance, and outbreak dynamics.

Pharmacogenomics and precision medicine

Individuals can respond differently to medications partly because of genetic variation.

NGS can contribute to pharmacogenomic research by identifying genetic variants associated with drug metabolism, drug response, or adverse reactions.

In precision medicine, genomic information can be combined with clinical information to support more individualized approaches to disease management.

Biomarker discovery

NGS can generate large datasets that researchers can use to search for molecular biomarkers.

Potential biomarkers may include genetic variants, gene expression patterns, or other molecular signatures associated with disease characteristics.

In cancer research, for example, NGS can help identify molecular features associated with tumor subtype, disease progression, or treatment response.

Advantages, Limitations, and Challenges of Next Generation Sequencing

NGS provides capabilities that were difficult or impossible to achieve with traditional sequencing methods. However, it also presents several technical and analytical challenges.

Advantages of NGS

One of the main advantages is high throughput.

Thousands to millions of DNA fragments can be analyzed simultaneously, allowing researchers to investigate large genomic regions efficiently.

Another advantage is scalability.

Researchers can select an approach that matches their experimental needs, ranging from targeted sequencing of a limited number of genes to whole-genome sequencing.

NGS can also identify multiple types of genomic alterations within the same experiment.

Depending on the platform and experimental design, researchers may investigate single-nucleotide variants, insertions, deletions, structural variants, or other genomic features.

NGS is also highly valuable for multiplexing.

Multiple samples can be prepared with different molecular indexes and sequenced together, reducing the need to perform separate sequencing runs for every sample.

Finally, NGS has opened new possibilities in areas such as cancer genomics, transcriptomics, microbial genomics, and precision medicine.

Large amounts of data

One of the main challenges associated with NGS is the enormous quantity of data generated.

A sequencing experiment can produce millions or even billions of reads.

Processing these data requires appropriate computational resources, storage capacity, and bioinformatics pipelines.

Researchers therefore need to consider bioinformatics during experimental planning rather than treating data analysis as an optional final step.

Library preparation can affect results

NGS data quality depends partly on the quality of library preparation.

Poor DNA or RNA quality, inefficient adapter ligation, amplification bias, contamination, or inappropriate fragment sizes can affect sequencing results.

For this reason, quality control should be performed throughout the workflow.

Sequencing errors

Every sequencing technology has characteristic sources of error.

The type and frequency of errors can differ between platforms.

Understanding these error profiles is particularly important when researchers are attempting to identify low-frequency variants or distinguish true biological differences from technical artifacts.

Short-read limitations

Short-read sequencing is highly effective for many applications, but short reads can make some genomic regions difficult to analyze.

Repetitive sequences and large structural rearrangements may be challenging to resolve when the available reads are too short.

Long-read sequencing can address some of these limitations, although it also has its own technical and analytical considerations.

Data interpretation

Generating a sequence is only one part of an NGS experiment.

The biological interpretation of the resulting data can be considerably more difficult.

For example, identifying a genetic variant does not automatically establish that the variant causes disease.

Researchers may need to consider population frequency, predicted molecular effects, experimental evidence, clinical information, and other factors before interpreting its significance.

This is particularly important in clinical genomics, where incorrect interpretation can have significant consequences.

Cost and infrastructure

Although the cost per sequenced base has decreased substantially compared with earlier sequencing technologies, NGS still requires specialized instruments, reagents, computational infrastructure, and trained personnel.

The total cost of an experiment therefore includes more than the sequencing run itself.

Sample preparation, library preparation, data storage, bioinformatics, quality control, and validation can all contribute to the overall cost.

Conclusion

Next generation sequencing has fundamentally changed modern genomics by allowing researchers to analyze enormous numbers of DNA or RNA molecules in parallel.

Unlike traditional Sanger sequencing, NGS can generate large quantities of sequence information from a single experiment. A typical workflow involves nucleic acid extraction, fragmentation or other library preparation steps, adapter addition, sequencing, signal detection, and computational analysis.

Several NGS technologies are available, including short-read platforms such as Illumina and long-read approaches such as Oxford Nanopore and PacBio. Each technology has distinct characteristics that make it suitable for particular applications.

NGS is now widely used in cancer research, whole-genome and whole-exome sequencing, RNA sequencing, genetic disease research, infectious disease research, biomarker discovery, pharmacogenomics, and precision medicine.

At the same time, NGS requires careful experimental design, quality control, computational analysis, and biological interpretation. The large volume of data generated by sequencing creates both opportunities and challenges for researchers.

- Advertisement -
Mohamed NAJID
Mohamed NAJID
Mohamed Najid is a PhD student in Cancer Cell Biology with a Master’s degree in Cancer Biology. His research focuses on circulating tumor cells (CTCs) in bladder cancer and their role as emerging diagnostic biomarkers.He creates clear, science-based content to help readers understand medical tests, cancer biology, and everyday health topics—without the confusion.ResearchGate: https://www.researchgate.net/profile/Mohamed-Najid-2 ORCID: https://orcid.org/0009-0002-7491-3366
RELATED ARTICLES
- Advertisment -

Related Articles