Google Study Validates DNBSEQ:
T7+ and T1+ Data Outperforms NovaSeq in AI-Driven Variant Calling

AI-Based Variant Calling Reveals Performance Advantages of DNBSEQ Platforms

Recently, the Google Research team hosted an online seminar titled “Scaling Genomics with Higher Throughput and AI-Driven Variant Calling,” showcasing the latest advances in a series of high-performance AI-based variant detection tools developed by Google, including DeepVariant, DeepConsensus and DeepSomatic. Notably, high-quality data from MGI/Complete Genomics’ DNBSEQ platform were featured prominently.

Google Study Validates DNBSEQ: T7+ and T1+ Data Outperforms NovaSeq in AI-Driven Variant Calling

The team compared the mean identity of data from different sequencing platforms. When aligned using a pangenome method, data from MGI’s T7+ platform achieved a mean identity of 0.995999, outperforming the 0.993489 of Illumina NovaSeq.

T7+ mapping with VG-Giraffe

The Google Research team then trained a DNBSEQ-specific model of DeepVariant using high-quality T7+ data, rather than a general model which was trained on other platforms. The training set for this model consisted of GIAB reference samples (HG001, HG002, HG004, HG005–HG007), with the HG003 sample and chromosome 20 data withheld to validate performance. The results were impressive: on the HG003 sample, the DNBSEQ-specific model produced a total of 14,183 false positive and false negative errors—significantly fewer than the 15,481 errors from the model trained on NovaSeq data.

HG003 performance on GIAB v4.2.1

For a more comprehensive evaluation, the team used the T2T (telomere-to-telomere) variant truth set for the HG002 sample. This truth set contains over 4.5 million variant sites—far more than older versions—enabling a more comprehensive assessment of performance. In this test, the T7+ DeepVariant model produced 64,116 total errors, significantly better than NovaSeq + DRAGEN v4.3 (71,854 errors) and NovaSeq + DeepVariant (73,213 errors).

HG002 whole genome at 40x coverage

When using the same cutting-edge AI tool, DeepVariant, the quality of the model differs markedly depending on the sequencing platform used to train it. The model trained on DNBSEQ platform data delivers higher performance, with fewer false positive and false negative variant calls. The seminar also shared data showing that the advantages of DNBSEQ are even more pronounced in certain genomic regions:

1 Homopolymer regions
In all homopolymer regions, DNBSEQ + DeepVariant improved indel detection accuracy by approximately 55% compared to NovaSeq + DRAGEN. This means that in challenging regions, DNBSEQ can more accurately determine whether insertions or deletions have occurred.

Homopolymer INDEL accuracy with Complete T7+ is 50% more accurate

2 Complex structural variant regions
In segmental duplication and complex copy number variation (CNV) regions, DNBSEQ + DeepVariant reduced the number of error sites by about 30% compared to NovaSeq + DRAGEN.

Resolving variants in complex structural events with Complete T7+

The reason lies in the distinct sequencing chemistries of the two platforms (DNA nanoball and combinatorial probe anchor synthesis vs. reversible terminator sequencing), which give DNBSEQ naturally lower background error rates in these specific regions. This provides a cleaner “signal” for the AI model and leads to better variant detection performance. The seminar also evaluated another DNBSEQ sequencing platform T1+. Results showed that whether using data from the higher-throughput T7+ or the more flexible T1+, the variant detection performance of the resulting models remained consistently high and superior to the comparison approach (NovaSeq + DRAGEN).

This indicates that the DNBSEQ platform delivers stable, reliable, high-quality data across different models and throughput levels, meeting the needs of projects ranging from large-scale population genomics to small, rapid research studies.

Comparing errors at 30x on T2T-Q100 truth set

Conclusion

This evaluation by Google Research demonstrates that the high accuracy and low error rates of the DNBSEQ sequencing platform can significantly enhance the performance of AI-based variant detection tools such as DeepVariant, especially in the most challenging regions of the genome. This provides a powerful technology combination for genomics researchers who demand the highest data quality and analytical precision.

Notably, the Google Research team, in collaboration with MGI and researchers from the University of Chinese Academy of Sciences, published a preprint on bioRxiv titled “PanVariants: Best Practice for Pangenome-based Variant Calling Pipeline and Framework.”

PanVariants: Best Practice for Pangenome-based Variant Calling Pipeline and Framework.

This study establishes a robust framework and best-practice pipeline for pangenome-based variant detection—named PanVariants—enabling sensitive discovery of novel variants and high-precision detection of single nucleotide variants (SNVs), insertions and deletions (indels), and structural variants (SVs), strongly supporting the future transition of genomics from linear to pangenome references.

Published 18 May 2026

Discover More