Google Study Validates DNBSEQ:
T7+ and T1+ Data Outperforms NovaSeq in AI-Driven Variant Calling
AI-Based Variant Calling Reveals Performance Advantages of DNBSEQ Platforms
Recently, the Google Research team hosted an online seminar titled “Scaling Genomics with Higher Throughput and AI-Driven Variant Calling,” showcasing the latest advances in a series of high-performance AI-based variant detection tools developed by Google, including DeepVariant, DeepConsensus and DeepSomatic. Notably, high-quality data from MGI/Complete Genomics’ DNBSEQ platform were featured prominently.
The team compared the mean identity of data from different sequencing platforms. When aligned using a pangenome method, data from MGI’s T7+ platform achieved a mean identity of 0.995999, outperforming the 0.993489 of Illumina NovaSeq.
The Google Research team then trained a DNBSEQ-specific model of DeepVariant using high-quality T7+ data, rather than a general model which was trained on other platforms. The training set for this model consisted of GIAB reference samples (HG001, HG002, HG004, HG005–HG007), with the HG003 sample and chromosome 20 data withheld to validate performance. The results were impressive: on the HG003 sample, the DNBSEQ-specific model produced a total of 14,183 false positive and false negative errors—significantly fewer than the 15,481 errors from the model trained on NovaSeq data.
For a more comprehensive evaluation, the team used the T2T (telomere-to-telomere) variant truth set for the HG002 sample. This truth set contains over 4.5 million variant sites—far more than older versions—enabling a more comprehensive assessment of performance. In this test, the T7+ DeepVariant model produced 64,116 total errors, significantly better than NovaSeq + DRAGEN v4.3 (71,854 errors) and NovaSeq + DeepVariant (73,213 errors).
When using the same cutting-edge AI tool, DeepVariant, the quality of the model differs markedly depending on the sequencing platform used to train it. The model trained on DNBSEQ platform data delivers higher performance, with fewer false positive and false negative variant calls. The seminar also shared data showing that the advantages of DNBSEQ are even more pronounced in certain genomic regions:
1 Homopolymer regions
In all homopolymer regions, DNBSEQ + DeepVariant improved indel detection accuracy by approximately 55% compared to NovaSeq + DRAGEN. This means that in challenging regions, DNBSEQ can more accurately determine whether insertions or deletions have occurred.
2 Complex structural variant regions
In segmental duplication and complex copy number variation (CNV) regions, DNBSEQ + DeepVariant reduced the number of error sites by about 30% compared to NovaSeq + DRAGEN.
The reason lies in the distinct sequencing chemistries of the two platforms (DNA nanoball and combinatorial probe anchor synthesis vs. reversible terminator sequencing), which give DNBSEQ naturally lower background error rates in these specific regions. This provides a cleaner “signal” for the AI model and leads to better variant detection performance. The seminar also evaluated another DNBSEQ sequencing platform T1+. Results showed that whether using data from the higher-throughput T7+ or the more flexible T1+, the variant detection performance of the resulting models remained consistently high and superior to the comparison approach (NovaSeq + DRAGEN).
This indicates that the DNBSEQ platform delivers stable, reliable, high-quality data across different models and throughput levels, meeting the needs of projects ranging from large-scale population genomics to small, rapid research studies.
Conclusion
This evaluation by Google Research demonstrates that the high accuracy and low error rates of the DNBSEQ sequencing platform can significantly enhance the performance of AI-based variant detection tools such as DeepVariant, especially in the most challenging regions of the genome. This provides a powerful technology combination for genomics researchers who demand the highest data quality and analytical precision.
Notably, the Google Research team, in collaboration with MGI and researchers from the University of Chinese Academy of Sciences, published a preprint on bioRxiv titled “PanVariants: Best Practice for Pangenome-based Variant Calling Pipeline and Framework.”
This study establishes a robust framework and best-practice pipeline for pangenome-based variant detection—named PanVariants—enabling sensitive discovery of novel variants and high-precision detection of single nucleotide variants (SNVs), insertions and deletions (indels), and structural variants (SVs), strongly supporting the future transition of genomics from linear to pangenome references.
Published 18 May 2026
Discover More
Sequencers
Explore our cutting-edge sequencers designed for high performance and accuracy.
CECs
Visit our Customer Experience Centres to see our workflows in action and discuss your projects.
Customer Stories
Discover how our customers are achieving remarkable results with our technologies.