We are sharing the results of a comprehensive benchmarking study evaluating MGI’s DNBSEQ platforms against three alternative sequencing platforms. Using the globally recognised commercial reference sample HG002/NA24385 under consistent 30× effective sequencing depth, the study demonstrates that MGI’s DNBSEQ platforms deliver superior, consistent performance across raw data quality, alignment metrics and variant calling.
The evaluation included three DNBSEQ platforms (T7+, T1+ and T7), compared against PCR-free whole-genome sequencing data from three comparable platforms (Element AVITI, Illumina NovaSeq 6000 and PacBio Onso). All datasets were analysed at identical effective sequencing depth to ensure a balanced comparison.
Key Findings
- Raw Data Quality
DNBSEQ platforms delivered exceptional and consistent Clean Q30 scores, exceeding 98%. Comparable platforms showed greater variability, ranging from 91% to 99%.
In contrast to competing platforms, DNBSEQ required only 93 Gbp to achieve 30× effective coverage, demonstrating more efficient data utilisation. Additionally, DNBSEQ platforms achieved Q40 scores above 94%, with internal testing showing Q50 scores exceeding 80%.
| Table 1. Comparison of Raw Data Quality Metrics Across Platforms | ||||||
|---|---|---|---|---|---|---|
| T7+ | T1+ | T7 | Element AVITI | NovaSeq 6000 | PacBio Onso | |
| Raw bases (Gbp) | 93 | 93 | 93 | 93 | 104 | 92 |
| Clean data rate (%) | 100 | 100 | 100 | 99.64 | 99.45 | 100 |
| Clean Q20 (%) | 99.09 | 99.48 | 99.47 | 97.31 | 96.77 | 99.7 |
| Clean Q30 (%) | 98.72 | 99.15 | 99.05 | 92.90 | 91.80 | 99.28 |
| GC content (%) | 40.65 | 40.68 | 39.98 | 40.82 | 41.34 | 40.13 |
Enhanced Alignment Performance
At 30× effective depth, DNBSEQ platforms matched or outperformed competitors across key alignment metrics:
- Duplicate rate: Maintained significantly lower duplication levels compared to alternatives, preserving effective data yield
- Mismatch rate: Achieved the lowest error rate at 0.34%, compared to over 0.5% for NovaSeq 6000 and PacBio Onso
- Coverage ≥20×: Exceeded 93% across all DNBSEQ platforms, indicating strong uniformity; PacBio Onso reached 88%
- Coverage ≥10× in homopolymer regions: Maintained >98% coverage, compared to 93% for PacBio Onso
| Table 2. Comparison of Alignment Metrics Across Platforms | ||||||
|---|---|---|---|---|---|---|
| T7+ | T1+ | T7 | Element AVITI |
NovaSeq 6000 |
PacBio Onso |
|
| Mapping rate (%) | 99.97 | 99.97 | 99.96 | 99.89 | 99.79 | 99.93 |
| Duplicate rate (%) | 0.62 | 2.97 | 1.09 | 0.52 | 9.17 | 1.36 |
| Mismatch rate (%) | 0.41 | 0.37 | 0.34 | 0.53 | 0.59 | 0.35 |
| Sequencing depth (x) | 31.41 | 30.78 | 31.58 | 31.39 | 30.13 | 31.55 |
| Coverage at least ≥20x (%) | 95.19 | 94.87 | 93.37 | 93.85 | 93.80 | 88.02 |
| Coverage at least ≥10x in homopolymers region (%) | 99.07 | 99.18 | 98.63 | 98.50 | 99.16 | 93.21 |
Variant Calling Accuracy
Variant calling performance was evaluated using NIST Reference Material v4.2.1 and the RTG Tools engine, following GA4GH-recommended methods. DNBSEQ platforms used the proprietary PanVariants pipeline, while comparator platforms used the open-source tool DeepVariant.
- SNP detection: Achieved precision of 99.80% and sensitivity of 99.78%, with an F-measure approaching 99.80%, outperforming Element AVITI (99.60%), NovaSeq 6000 (99.57%) and PacBio Onso (98.93%)
- Indel detection: Achieved an F-measure exceeding 99.55%, comparable to Element AVITI (99.52%) and higher than NovaSeq 6000 (99.42%) and PacBio Onso (98.38%)
| Table 3. Comparison of Variant Calling Metrics Across Platforms | ||||||
|---|---|---|---|---|---|---|
| Sample | T7+ | T1+ | T7 | Element AVITI |
NovaSeq 6000 |
PacBio Onso |
| SNP_True-pos | 3,357,811 | 3,357,798 | 3,357,744 | 3,343,564 | 3,341,522 | 3,302,760 |
| SNP_Precision (%) | 99.83 | 99.82 | 99.82 | 99.85 | 99.85 | 99.73 |
| SNP_Sensitivity (%) | 99.78 | 99.78 | 99.78 | 99.36 | 99.30 | 98.15 |
| SNP_F-measure (%) | 99.80 | 99.80 | 99.80 | 99.60 | 99.57 | 98.93 |
| Indel_True-pos | 522,701 | 522,972 | 522,773 | 521,786 | 521,240 | 512,596 |
| Indel_Precision (%) | 99.63 | 99.64 | 99.61 | 99.74 | 99.64 | 99.22 |
| Indel_Sensitivity (%) | 99.47 | 99.52 | 99.49 | 99.30 | 99.20 | 97.55 |
| Indel_F-measure (%) | 99.55 | 99.58 | 99.55 | 99.52 | 99.42 | 98.38 |
Summary
DNBSEQ platforms demonstrated strong and consistent performance across all evaluated metrics:
- 98% Clean Q30 scores
- 93% coverage uniformity at ≥20×
- 99.80% SNP detection F-measure
- 99.55% indel detection F-measure
Comparator platforms showed the following characteristics:
- Element AVITI: Greater variability in quality scores and higher mismatch rates
- NovaSeq 6000: Higher duplication and mismatch rates
- PacBio Onso: Lower coverage uniformity and reduced variant detection accuracy
Methodology Transparency
To ensure reproducibility, MGI provides detailed analytical workflows.
Data Processing
Due to differences in sequencing throughput across platforms, uniform subsampling was applied to FASTQ files to ensure comparability. Each dataset was subsampled to 30× coverage using seqtk (v1.2). Quality control was then performed using SOAPnuke (v2.1.9), including adapter trimming, low-quality filtering and removal of reads with high N content.
Alignment
Following subsampling and quality control, alignment was performed using the FPGA-accelerated MegaBOLT software, implementing the BWA-MEM2 algorithm against the GRCh38 reference genome. Genome stratification intervals published by NIST were used to evaluate performance in complex genomic regions.
Variant Calling
SNP and indel detection was conducted using the PanVariants pipeline for DNBSEQ data and DeepVariant (v1.9.0) for comparator platforms. Performance was evaluated using RTG Tools against the NIST v4.2.1 benchmark set.
- Precision = TP / (TP + FP)
- Sensitivity = TP / (TP + FN)
- F-measure = 2 × Precision × Sensitivity / (Precision + Sensitivity)
References:
- Benchmarking challenging small variants with linked and long reads. Cell Genom. 2022 May;2(5):100128. doi: 10.1016/j.xgen.2022.100128. PMID: 36452119; PMCID: PMC9706577.
- Comparing Variant Call Files for Performance Benchmarking of Next-Generation Sequencing Variant Calling Pipelines. bioRxiv 023754; doi: https://doi.org/10.1101/023754.
- Best practices for benchmarking germline small-variant calls in human genomes. Nat Biotechnol. 2019 May;37(5):555-560. doi: 10.1038/s41587-019-0054-x. Epub 2019 Mar 11. Erratum in: Nat Biotechnol. 2019 May;37(5):567. doi: 10.1038/s41587-019-0108-0. PMID: 30858580; PMCID: PMC6699627.
- A universal SNP and small-indel variant caller using deep neural networks. Nature Biotechnology 36, 983–987 (2018). . doi: https://doi.org/10.1038/nbt.4235. Shen W, Le S, Li Y, Hu F. SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation. PLoS One. 2016 Oct 5;11(10):e0163962. doi: 10.1371/journal.pone.0163962. PMID: 27706213; PMCID: PMC5051824.
- SOAPnuke: a MapReduce acceleration-supported software for integrated quality control and preprocessing of high-throughput sequencing data. Gigascience. 2018 Jan 1;7(1):1-6. doi: 10.1093/gigascience/gix120. PMID: 29220494; PMCID: PMC5788068.
- Efficient Architecture-Aware Acceleration of BWA-MEM for Multicore Systems. IEEE Parallel and Distributed Processing Symposium (IPDPS), 2019. 10.1109/IPDPS.2019.00041.
- The GIAB genomic stratifications resource for human reference genomes. Nat Commun 15, 9029 (2024). https://doi.org/10.1038/s41467-




