AI Frontier Post
AI News

Illumina launches SpliceAI2: genomic AI that predicts RNA splicing from DNA alone — and beats AlphaGenome on the benchmarks

Illumina's BioInsight AI Lab has launched SpliceAI2, a genomic AI model that predicts how genetic variants alter RNA splicing from a DNA sequence alone — finding 17% more disease-relevant variants than rival models and outperforming Google DeepMind's AlphaGenome across three independent benchmarks.

Illumina's BioInsight AI Lab has launched SpliceAI2, a genomic AI model that predicts how genetic variants alter RNA splicing — from a DNA sequence alone. According to the company, the model identified 17% more disease-relevant variants than other splicing models when applied to a rare disease research dataset, and it outperformed Google DeepMind's AlphaGenome across three independent benchmarks, Unite.AI reported on the company's announcement.

What changed since the 2019 original

The first SpliceAI predicted whether a cell would splice at a given location. SpliceAI2 answers three harder questions: which positions in a gene are used as splice sites and how often, which splice sites connect through splice junctions, and which full-length RNA transcript isoforms get produced. The model takes only a DNA sequence as input — which Illumina says enables transcript-level analysis without RNA data from difficult-to-obtain tissues.

A DNA double helix illustration representing SpliceAI2, which takes only a DNA sequence as input.
DNA in, transcript-level splicing predictions out — no RNA sequencing required from hard-to-obtain tissues. Photo via Unsplash.

Under the hood: SpliceAI2 processes 196,608 base pairs of genomic context with roughly 13 million trainable parameters, integrating splice-site and junction predictions into a splice graph from which complete transcripts and their usage are inferred. It was trained end-to-end on 314,745 RNA sequencing samples spanning humans and nine other mammalian species — more than 46 million observed splice junctions after filtering — plus 330 long-read RNA sequencing samples from the public ENCODE project. The model exactly reconstructed the most abundant transcript for 82% of held-out genes, up from 78% without the long-read data. Illumina says the training set is more than 100 times larger than its predecessor's; the original has been cited in over 3,400 publications and sits in ClinGen's splice variant interpretation recommendations.

The benchmarks: AlphaGenome included

Across three independent benchmarks, the accompanying manuscript reports SpliceAI2 outperforming the original SpliceAI, Pangolin, and Google DeepMind's AlphaGenome — with the AlphaGenome comparisons run independently by collaborators at the University of Oxford. The reported numbers: an auPRC of 0.77 vs 0.66 for the next-best model on GTEx cryptic splice variant detection, a Spearman correlation of 0.63 vs 0.47 on splice site usage quantification, an auROC of 0.76 vs 0.73 on splicing quantitative trait locus classification, and a Spearman correlation of 0.59 vs 0.55 on the OpenSplice massively parallel reporter assay benchmark. Illumina's announcement puts the splice-site-usage improvement at 34% over the next-best model.

Diagram illustrating RNA splicing decisions: which splice sites are used, how junctions connect, and which transcript isoforms result.
The three questions SpliceAI2 answers: which splice sites are used and how often, which connect through junctions, and which full-length isoforms result. Diagram via Wikimedia Commons.

Rare disease is the real test

In an analysis of 7,504 probands from the Genomics England 100,000 Genomes Project, variants prioritized by SpliceAI2 were significantly enriched in phenotype-matched disease genes — identifying 17% more disease-associated variants than any other tested splicing model at matched confidence thresholds, and recovering 133 excess variants versus 114 for the next-best model at a fixed odds ratio of 2. Against the legacy SpliceAI, the gains were 33% more disease-relevant splice variants at a 2X confidence interval and 66% more at 4X. Roughly half of the cryptic splice variants identified sat deep within intronic regions; variants more than 50 base pairs into introns would typically be missed by exome sequencing. Predicted splice-altering variants accounted for 15% of the excess genetic burden in the cohort.

Population-scale validation drew on more than 627,000 genomes from gnomAD, TOPMed, and UK Biobank: the highest-scoring SpliceAI2 variants were strongly depleted, approaching the depletion seen for protein-truncating loss-of-function mutations. UK Biobank proteomic data from 36,764 participants showed carriers of higher-scoring variants had lower plasma protein levels (Pearson correlation −0.50, the strongest among models tested), and RNA sequencing from 5,435 Genomics England participants validated predictions in individual cases — including a PEX1 donor-loss variant associated with rod-cone dystrophy and a PKD1 cryptic donor variant associated with cystic kidney disease.

Tissue-specific splicing, learned rather than taught

The manuscript also describes a fine-tuning framework that conditions predictions on the expression levels of 147 RNA binding proteins, capturing tissue-specific splicing differences across 48 GTEx tissues in nearly 15 million differential splice site usage measurements. The authors report the model independently learned sequence motifs recognized by real splicing regulators without being explicitly taught those relationships, and that the framework was adapted to disease states including SF3B1-mutant tumors and myotonic dystrophy.

A three-model suite for the noncoding genome

SpliceAI2 joins PromoterAI and PrimateAI-3D as a suite covering splice, promoter, and missense variant effect types. Illumina says the three collectively let researchers identify up to twice as many variants with predicted biological impact. "Illumina is advancing AI to systematically shrink the portion of the genome that remains uninterpretable," said Kyle Farh, vice president of the BioInsight AI Lab. Rami Mehio, senior vice president and general manager of BioInsight, described variant effect prediction tools like SpliceAI2 as one of the lab's key focus areas — alongside generating genomic and multiomic data and training biological foundation models.

The model is accessible through Illumina's BioInsight Platform applications, including DRAGEN Annotation, Emedgene, and Illumina Connected Insights. Source code, trained weights, and precomputed predictions — 4 billion single-nucleotide variants within human gene bodies and 150 million indels observed in human populations — are available through the SpliceAI2 GitHub repository and Hugging Face for academic and non-commercial research use, with a separate contact for commercial licensing. The software installs via PyPI and needs a CUDA-capable GPU. Score thresholds: 0.1 for high recall, 0.25 for balanced precision and recall, 0.5 for high precision.

Sources: Unite.AI — Illumina Launches SpliceAI2 Model for Splice Variant Interpretation (Oct. 8, 2026), citing Illumina's company announcement, technical article, and the SpliceAI2 manuscript.