Recently, Dr Roy Tan, Director of the Americas Marketing Center at MGI, shared his insights on the development of sequencing technology and the future of large-scale population genomics.
With many years of experience in sequencing technology development and applications, Dr Tan discussed recent advances in sequencing methodologies, library preparation and data quality standards.
Advances in Sequencing Technology
Q: The development of sequencing technology is progressing rapidly. Can you tell us about the progress made in recent years?
A: Over the past few years, high-throughput parallel sequencing technologies have developed rapidly. MGI’s DNBSEQ™ sequencing technology is a strong example of this progress.
In recent years, our technological advances have mainly focused on improvements in detection methodologies, library preparation techniques and the development of large-scale parallel sequencing chips. At the same time, we have gained a deeper understanding of sequencing data quality.
Improving Sequencing Library Quality
Q: High-quality library preparation is essential for accurate sequencing. How can we obtain higher quality sequencing libraries?
A: There have been two important developments in DNA library preparation.
The first is preserving the original state of the DNA sample as much as possible throughout the sequencing process. By using high-fidelity processes and avoiding unnecessary amplification errors, we can retain the authentic genomic information of the sample.
The second development is the use of PCR-free sequencing methods. Our sequencing technology uses linear amplification and PCR-free library preparation, allowing us to obtain a more accurate representation of the genome.
PCR-free library preparation offers several advantages, including lower input requirements, longer read lengths and support for multiple samples. It also improves accuracy when sequencing regions with high GC content, producing longer and more reliable contigs that are less affected by GC bias.
stLFR Single-Tube Long Fragment Read Technology
Q: Can you introduce the stLFR technology developed by MGI?
A: Another rapidly developing technology is stLFR, which stands for single-tube long fragment read technology.
The main principle of stLFR is a dual-barcode labelling system. Humans have two sets of genomes, one inherited from the mother and one from the father. Ideally, we want to distinguish these genomes within the same sequencing experiment.
In stLFR, long DNA fragments are tagged with barcodes attached to microscopic beads. By identifying these barcodes, we can determine which short sequencing reads originate from the same long DNA fragment.
Using this approach, fragments of up to around 300 kilobases can be reconstructed, with most fragments averaging about 60 kilobases.
This technology is particularly useful for detecting structural variations and performing haplotype phasing, which allows researchers to determine whether fragments originate from the maternal or paternal genome.
We refer to genomes generated through PCR-free sequencing as the “real genome”, while genomes reconstructed using stLFR technology move closer to what we call the “perfect genome”. These technologies allow our sequencing platforms to detect long fragments while maintaining very high accuracy.
Increasing Sequencing Throughput
Q: As sequencing applications continue to grow, the need for higher throughput is becoming more urgent. What breakthroughs have been made and where is sequencing technology heading?
A: The development of sequencing applications is closely linked to advances in chip technology. To achieve large-scale sequencing, we need larger and more efficient sequencing chips.
One example is the chip used in the DNBSEQ-T7 sequencer, which can generate up to five billion reads in a single run.
Looking ahead, sequencing technology will likely develop in two main directions.
The first is optical improvements. By improving the performance of optical detection systems and reducing the spacing between sequencing sites on a chip, more reads can be generated from the same surface area.
The second direction is advances in sequencing chemistry. Improvements in chemistry could enable longer paired-end reads, such as PE300 or PE600, which would significantly increase total sequencing throughput.
The “676” Genome Quality Standard
Q: You mentioned earlier that there have been new developments in sequencing data quality. Can you explain the “676” standard proposed by MGI?
A: The “676” standard is a benchmark proposed by MGI for defining high-precision genome assemblies under modern sequencing conditions.
Under this standard, when sequencing depth reaches around 50× and duplication is minimal, the genome assembly should meet several criteria. These include a Contig N50 greater than 1 Mb, a Scaffold N50 greater than 10 Mb and a total assembled genome size greater than 6 Gb for the human genome.
It is also important to distinguish between resequencing and de novo genome assembly. Resequencing compares sequencing reads with an existing reference genome, while de novo assembly reconstructs the genome entirely from sequencing data without using a reference.
De novo assembly provides a more accurate representation of the true genome.
Using this approach, MGI researchers have assembled genomes for multiple plant and animal species, including more than 20 species of marine fish, demonstrating strong performance and reliability.
Sequencing Large Populations
Q: In recent years, several countries including the United Kingdom, China and the United States have launched large population sequencing programmes. Has population sequencing become a global trend?
A: Interpreting human genes and understanding complex diseases requires sequencing large populations. Ultimately, the largest population group is the entire human population.
There are currently around eight billion people worldwide. If we aim to sequence all human genomes within the next 50 years using the “676” high-precision genome standard, we would need to sequence approximately 240 million people each year.
Although this is an ambitious target, it is achievable with sufficient infrastructure. If there were 1,000 sequencing laboratories worldwide, and each laboratory sequenced 1,000 people per day, this goal could realistically be reached.
The Future of Personal Genomics
Q: What role will large-scale sequencing play in the future of healthcare?
A: Health is influenced by two main factors: environment and genetics. Environmental factors relate to the conditions we live in, while genetics represents the information inherited from our parents.
Throughout human history, we have learned how to better manage our environment. Now, advances in sequencing technology are helping us better understand genetics.
With continued technological progress, large-scale genomic data will help researchers identify disease risks earlier and develop more personalised healthcare strategies.
Ultimately, the goal is to make high-quality genomic information accessible to everyone, enabling better prevention, diagnosis and treatment of disease.




