<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Promoterai on Genomics x AI</title><link>https://genomicsxai.github.io/tags/promoterai/</link><description>Recent content in Promoterai on Genomics x AI</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Tue, 29 Sep 2026 12:11:09 +0000</lastBuildDate><atom:link href="https://genomicsxai.github.io/tags/promoterai/index.xml" rel="self" type="application/rss+xml"/><item><title>promoterai-torch: a PyTorch port of Illumina's PromoterAI</title><link>https://genomicsxai.github.io/blogs/2026-014/</link><pubDate>Tue, 29 Sep 2026 00:00:00 +0000</pubDate><guid>https://genomicsxai.github.io/blogs/2026-014/</guid><description>&lt;img src="https://genomicsxai.github.io/" alt="Featured image of post promoterai-torch: a PyTorch port of Illumina's PromoterAI" /&gt;&lt;aside class="summary-box"&gt;&#10; &lt;h2 class="summary-box__title"&gt;Summary&lt;/h2&gt;&#10; &lt;div class="summary-box__body"&gt;&#10; PromoterAI [1] predicts how promoter variants alter gene expression, but the official release ships as a TensorFlow/Keras SavedModel [2]. &lt;a class="link" href="https://github.com/genomicsxai/promoterai-torch" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;promoterai-torch&lt;/code&gt;&lt;/a&gt; is an independent, numerically-equivalent PyTorch port that converts Illumina&amp;rsquo;s checkpoints and makes variant scoring, track prediction, embedding extraction, and DeepLIFT/SHAP attribution available through the PyTorch ecosystem, with training and fine-tuning scripts included for anyone who wants to reproduce or extend the model from scratch.&#10; &lt;/div&gt;&#10;&lt;/aside&gt;&#10;&#10;&lt;hr&gt;&#10;&lt;h2 id="overview"&gt;Overview&#10;&lt;/h2&gt;&lt;p&gt;Promoters are a critical class of non-coding regulatory DNA elements that set the baseline transcriptional output of a gene. A single nucleotide change just upstream of the transcription start site can create or destroy a transcription factor binding site without touching the coding sequence of a gene. The best-known example is &lt;em&gt;TERT&lt;/em&gt;: two recurrent, mutually exclusive promoter mutations, independently identified in melanoma [4, 5] and subsequently found in glioblastoma, bladder cancer, and dozens of other tumor types [6], each create a de novo ETS/GABP transcription-factor binding motif upstream of the transcription start site, driving aberrant telomerase re-expression [4–7]. Both are in the &lt;em&gt;TERT&lt;/em&gt; variant set checked below — C228T and C250T, chr5:1,295,113 G&amp;gt;A and chr5:1,295,135 G&amp;gt;A on the hg38 minus strand — and PromoterAI&amp;rsquo;s ensembled scores for them are 0.74 and 0.85 respectively: both comfortably past the paper&amp;rsquo;s ±0.5 &amp;ldquo;strong effect&amp;rdquo; threshold and positive, consistent with the gain-of-function reported in the literature.&lt;/p&gt;&#10;&#10;&#9;&#9;&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-014/TERT_promoter_track.png" class="image-link" data-pswp-width="2000" data-pswp-height="920"&gt;&#10;&#9;&#9;&lt;img src="https://genomicsxai.github.io/blogs/2026-014/TERT_promoter_track.png.optimized.webp" width="750px" height="345"loading="lazy"&#10;&#9;&#9;&#9;alt="Bar chart of PromoterAI’s saturation-mutagenesis scores across a 1 kb window of the TERT promoter (chr5, hg38), colored red for positive (over-expression) and blue for negative (under-expression) scores, with the transcription start site (arrow marking the minus-strand direction of transcription) and the C228T and C250T mutations annotated. Both mutations sit in a cluster of strongly positive-scoring positions and score well above the &amp;#43;0.5 strong-effect threshold."&#10;&#9;&#9;&#9;data-title-escaped="PromoterAI saturation-mutagenesis scores across the TERT promoter (chr5, hg38); bar height is the largest-magnitude signed score among the 3 possible alt alleles at each position. TERT is on the minus strand, so transcription proceeds toward decreasing coordinate (arrow). The recurrent C228T and C250T mutations (chr5:1,295,113 G&amp;amp;gt;A and chr5:1,295,135 G&amp;amp;gt;A) both score well past the paper&amp;amp;#39;s &amp;#43;0.5 strong-effect threshold."&gt;&#10;&#9;&#9;&lt;/a&gt;&lt;figcaption&gt;PromoterAI saturation-mutagenesis scores across the TERT promoter (chr5, hg38); bar height is the largest-magnitude signed score among the 3 possible alt alleles at each position. TERT is on the minus strand, so transcription proceeds toward decreasing coordinate (arrow). The recurrent C228T and C250T mutations (chr5:1,295,113 G&amp;gt;A and chr5:1,295,135 G&amp;gt;A) both score well past the paper&amp;rsquo;s +0.5 strong-effect threshold.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Prioritizing functional promoter variants more broadly is a hard, unsolved problem in variant interpretation. PromoterAI addressed this by training a sequence-to-function model on hundreds of regulatory tracks (histone marks, TF ChIP-seq, ATAC-seq, RNA-seq) across human and mouse promoters, then fine-tuning on expression outlier variants (with signed differences between reference and alternate predictions as a variant effect score).&lt;/p&gt;&#10;&lt;p&gt;However, the official release of PromoterAI is in TensorFlow/Keras, which has effectively walled PromoterAI off from the PyTorch-based S2F ecosystem since it shipped. Getting per-base attributions out of PromoterAI today means reimplementing DeepLIFT/SHAP&amp;rsquo;s gradient-correction rules against &lt;code&gt;tf.GradientTape&lt;/code&gt; — a substantial undertaking, and easy to get subtly wrong. Similarly, gradient-based sequence design with tools like &lt;a class="link" href="https://github.com/jmschrei/ledidi" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;Ledidi&lt;/code&gt;&lt;/a&gt; [8], which optimizes a sequence toward a target prediction by backpropagating through the model, needs the same kind of direct, differentiable access. Scoring the same variant set across PromoterAI and PyTorch-native models, to compare or ensemble their predictions, means round-tripping through disk rather than composing tensors directly. &lt;code&gt;promoterai-torch&lt;/code&gt; re-implements the architecture layer-for-layer in PyTorch and ships a converter that reads an existing Illumina SavedModel and produces a &lt;code&gt;.pt&lt;/code&gt; checkpoint, so PromoterAI becomes just another &lt;code&gt;nn.Module&lt;/code&gt; that plugs into existing PyTorch pipelines — attribution, design, and ensembling with the rest of the PyTorch S2F ecosystem all become drop-in rather than bespoke.&lt;/p&gt;&#10;&lt;p&gt;Porting sequence-to-function models to PyTorch this way is a familiar, worthwhile pattern for the community rather than a one-off: &lt;a class="link" href="https://github.com/lucidrains/enformer-pytorch" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;enformer-pytorch&lt;/code&gt;&lt;/a&gt; ports Enformer [9], Flashzoi provides an accelerated PyTorch reimplementation of Borzoi [10], and we previously covered &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-004/" target="_blank" rel="noopener"&#10; &gt;porting AlphaGenome to PyTorch&lt;/a&gt; on this blog [11]. &lt;code&gt;promoterai-torch&lt;/code&gt; follows the same playbook for PromoterAI.&lt;/p&gt;&#10;&lt;h3 id="architecture"&gt;Architecture&#10;&lt;/h3&gt;&lt;p&gt;PromoterAI&amp;rsquo;s backbone is a &amp;ldquo;MetaFormer&amp;rdquo;-style stack: a 1×1 convolution stem projects the one-hot DNA sequence into a &lt;code&gt;model_dim&lt;/code&gt;-wide channel space, followed by &lt;code&gt;num_blocks&lt;/code&gt; residual blocks. Each block alternates two mixing operations behind BatchNorm and a residual connection — a depthwise, dilated 1D convolution across positions (token mixing, with the dilation rate held at 1 for the first four blocks, then doubling every two blocks thereafter: 1, 1, 1, 1, 2, 2, 4, 4, …) and a two-layer feed-forward network across channels (channel mixing). This is the same token-mixing/channel-mixing split used by &amp;ldquo;MetaFormer&amp;rdquo;-family vision architectures, with a convolutional token mixer in place of pooling or attention.&lt;/p&gt;&#10;&lt;p&gt;Prediction happens through per-species output heads that read out from several depths of the backbone rather than just the final block: every &lt;code&gt;shortcut_layer_freq&lt;/code&gt;-th block&amp;rsquo;s hidden state is linearly projected to that species&amp;rsquo; track dimension, passed through a ReLU, and the resulting projections are averaged and center-cropped to the model&amp;rsquo;s output length. The released checkpoints carry two such heads sharing one backbone — a 498-track human head and a 472-track mouse head. Only the human head is unfrozen during PromoterAI&amp;rsquo;s own variant-effect fine-tuning; the rest of the network, including BatchNorm running statistics, stays in eval mode throughout.&lt;/p&gt;&#10;&#10; &lt;blockquote&gt;&#10; &lt;p&gt;This is &lt;strong&gt;not&lt;/strong&gt; an official Illumina product or publication. &lt;code&gt;promoterai-torch&lt;/code&gt;&amp;rsquo;s PyTorch code is an independent reimplementation, and its release should not be construed as endorsed by Illumina. The pretrained &lt;em&gt;weights&lt;/em&gt; remain gated under Illumina&amp;rsquo;s own license, independent of this port&amp;rsquo;s own code license — see &lt;a class="link" href="#license" &gt;License&lt;/a&gt; below for the full picture, and please don&amp;rsquo;t redistribute converted checkpoints.&lt;/p&gt;&#10;&#10; &lt;/blockquote&gt;&#10;&lt;h2 id="getting-started"&gt;Getting Started&#10;&lt;/h2&gt;&lt;p&gt;The core package installs without pulling in TensorFlow, HDF5/BigWig tooling, or attribution libraries:&lt;/p&gt;&#10;&lt;div class="code-block" data-lang="sh"&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;pip install promoterai-torch&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Converting a pretrained checkpoint requires the &lt;code&gt;[convert]&lt;/code&gt; extra and a copy of the official SavedModel from &lt;a class="link" href="https://github.com/Illumina/PromoterAI" target="_blank" rel="noopener"&#10; &gt;Illumina/PromoterAI&lt;/a&gt;. Note that Illumina gates the pretrained SavedModels and precomputed variant scores behind a signed academic-use license agreement (commercial licensing goes through &lt;code&gt;AI_licensing@illumina.com&lt;/code&gt;) — see their README for the request form first.&lt;/p&gt;&#10;&lt;div class="code-block" data-lang="sh"&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;pip install &lt;span class="s2"&gt;&amp;#34;promoterai-torch[convert]&amp;#34;&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;promoterai-torch convert &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --keras_model models/promoterAI_v1_hg38_mm10_finetune &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --output models/promoterAI_v1_hg38_mm10_finetune.pt &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --input_length &lt;span class="m"&gt;20480&lt;/span&gt; &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --output_length &lt;span class="m"&gt;4096&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Architecture hyperparameters (&lt;code&gt;num_blocks&lt;/code&gt;, &lt;code&gt;model_dim&lt;/code&gt;, &lt;code&gt;output_dims&lt;/code&gt;) are inferred automatically from the SavedModel, so this works for any PromoterAI-architecture checkpoint — not just Illumina&amp;rsquo;s four released checkpoints, but your own Keras fine-tunes produced by &lt;code&gt;promoterai.finetune&lt;/code&gt; as well.&lt;/p&gt;&#10;&lt;p&gt;From there, scoring a variant TSV (&lt;code&gt;chrom&lt;/code&gt;, &lt;code&gt;pos&lt;/code&gt;, &lt;code&gt;ref&lt;/code&gt;, &lt;code&gt;alt&lt;/code&gt;, &lt;code&gt;strand&lt;/code&gt;) is one command:&lt;/p&gt;&#10;&lt;div class="code-block" data-lang="sh"&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;promoterai-torch score &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --model_checkpoint models/promoterAI_v1_hg38_mm10_finetune.pt &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --var_file variants.tsv &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --fasta_file hg38.fa &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --input_length &lt;span class="m"&gt;20480&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Scores land in [−1, 1], with the same effect-size thresholds as the original paper (±0.1 weak, ±0.2 moderate, ±0.5 strong).&lt;/p&gt;&#10;&lt;h2 id="numerical-equivalence"&gt;Numerical Equivalence&#10;&lt;/h2&gt;&lt;p&gt;Porting a model is only useful if it actually reproduces the original, so most of the engineering effort went here.&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Track-level equivalence.&lt;/strong&gt; Running both the original TF/Keras SavedModel and the converted PyTorch checkpoint on the same sequences and comparing every output track gives errors of ~1e-7 at FP32 — within machine precision — across all four released checkpoints (&lt;code&gt;hg38&lt;/code&gt;, &lt;code&gt;hg38_mm10&lt;/code&gt;, &lt;code&gt;hg38_finetune&lt;/code&gt;, &lt;code&gt;hg38_mm10_finetune&lt;/code&gt;).&lt;/p&gt;&#10;&lt;p&gt;&lt;strong&gt;Variant-score equivalence.&lt;/strong&gt; On promoter variants at &lt;em&gt;TERT&lt;/em&gt; (&lt;em&gt;n&lt;/em&gt; = 6,006), &lt;em&gt;SFSWAP&lt;/em&gt; (&lt;em&gt;n&lt;/em&gt; = 3,003), and &lt;em&gt;DNAJC9&lt;/em&gt; (&lt;em&gt;n&lt;/em&gt; = 9,009), torch and TF/Keras variant scores are identical, including the ensembled score used in the paper and distributed in the official scores published by Illumina (Pearson &lt;em&gt;r&lt;/em&gt; = 1.0000, MAE = 0.0000). Note that the scoring script/CLI in both the official repo and this port round the score to 4 digits, which is why variant scores will generally be identical.&lt;/p&gt;&#10;&#10;&#9;&#9;&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-014/TERT_scatter.png" class="image-link" data-pswp-width="2390" data-pswp-height="499"&gt;&#10;&#9;&#9;&lt;img src="https://genomicsxai.github.io/blogs/2026-014/TERT_scatter.png.optimized.webp" width="700px" height="146"loading="lazy"&#10;&#9;&#9;&#9;alt="Five scatter plots comparing PromoterAI TERT variant scores across checkpoints and implementations — per-checkpoint TF versus torch scores, ensembled TF versus torch scores, and each against the officially published scores — all falling exactly on the identity line."&#10;&#9;&#9;&#9;data-title-escaped="TERT promoter variant scores (n = 6,006): per-checkpoint and ensembled TF/Keras versus PyTorch scores, and each versus the officially published PromoterAI scores. r = 1.000, MAE = 0.000 in every panel."&gt;&#10;&#9;&#9;&lt;/a&gt;&lt;figcaption&gt;TERT promoter variant scores (n = 6,006): per-checkpoint and ensembled TF/Keras versus PyTorch scores, and each versus the officially published PromoterAI scores. r = 1.000, MAE = 0.000 in every panel.&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;&lt;strong&gt;Benchmark equivalence.&lt;/strong&gt; Scoring the public benchmark variant sets released alongside the paper — &lt;code&gt;CAGI5_saturation&lt;/code&gt;, &lt;code&gt;GEL_RNA&lt;/code&gt;, &lt;code&gt;GTEx_eQTL&lt;/code&gt;, &lt;code&gt;GTEx_outlier&lt;/code&gt;, &lt;code&gt;MPRA_eQTL&lt;/code&gt;, &lt;code&gt;MPRA_saturation&lt;/code&gt;, and &lt;code&gt;UKBB_proteome&lt;/code&gt; (under/over/null variant categories per dataset) — with the torch checkpoints reproduces the TF/Keras ensemble&amp;rsquo;s under-vs-over, under-vs-null, and over-vs-null AUROCs to within ~1e-6:&lt;/p&gt;&#10;&lt;table&gt;&#10;&#9;&lt;thead&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;Dataset&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;&lt;em&gt;n&lt;/em&gt; (under/over/null)&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;under-vs-over&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;under-vs-null&lt;/th&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;th&gt;over-vs-null&lt;/th&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/thead&gt;&#10;&#9;&lt;tbody&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;CAGI5_saturation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;976 / 499 / 5,095&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8845&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7939&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7153&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;GEL_RNA&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;309 / 239 / 609&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.9002&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7757&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7802&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;GTEx_eQTL&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;191 / 218 / 393&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8697&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7876&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7503&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;GTEx_outlier&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;206 / 161 / 382&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8938&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7972&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7423&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;MPRA_eQTL&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;70 / 74 / 542&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.9004&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8069&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8278&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;MPRA_saturation&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;773 / 275 / 3,981&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8707&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.8675&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7010&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&#9;&#9;&lt;tr&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;UKBB_proteome&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;182 / 69 / 760&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.9116&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7718&lt;/td&gt;&#10;&#9;&#9;&#9;&#9;&#9;&lt;td&gt;0.7757&lt;/td&gt;&#10;&#9;&#9;&#9;&lt;/tr&gt;&#10;&#9;&lt;/tbody&gt;&#10;&lt;/table&gt;&#10;&lt;p&gt;(Torch AUROCs shown; the matching TF/Keras run agrees on every value to at least five decimal places.) The per-dataset and aggregate ensemble variant scores underlying these AUROCs also match nearly exactly between the two implementations (Pearson &lt;em&gt;r&lt;/em&gt; = 1.0000 for each of the seven datasets and for all 16,004 variants combined):&lt;/p&gt;&#10;&#10;&#9;&#9;&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-014/paper_benchmark_concordance.png" class="image-link" data-pswp-width="3000" data-pswp-height="3000"&gt;&#10;&#9;&#9;&lt;img src="https://genomicsxai.github.io/blogs/2026-014/paper_benchmark_concordance.png.optimized.webp" width="700px" height="700"loading="lazy"&#10;&#9;&#9;&#9;alt="Grid of eight scatter plots, one per benchmark dataset plus an aggregate panel, each showing PyTorch ensemble variant scores plotted against TF/Keras ensemble scores falling tightly on the identity line."&#10;&#9;&#9;&#9;data-title-escaped="PyTorch versus TF/Keras ensemble variant scores on each of the paper&amp;amp;#39;s released benchmark datasets (Pearson r = 1.0000 in every panel) and combined across all 16,004 variants (bottom centre)."&gt;&#10;&#9;&#9;&lt;/a&gt;&lt;figcaption&gt;PyTorch versus TF/Keras ensemble variant scores on each of the paper&amp;rsquo;s released benchmark datasets (Pearson r = 1.0000 in every panel) and combined across all 16,004 variants (bottom centre).&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;&lt;strong&gt;Training equivalence.&lt;/strong&gt; Inference equivalence doesn&amp;rsquo;t guarantee the training loop itself matches — a converter can produce an identical model while the from-scratch training and fine-tuning code silently diverges from Illumina&amp;rsquo;s Keras implementation. &lt;code&gt;train.py&lt;/code&gt; and &lt;code&gt;finetune.py&lt;/code&gt; clip gradients per parameter (matching Keras&amp;rsquo; &lt;code&gt;clipnorm&lt;/code&gt; semantics, as opposed to a single norm across all parameters jointly), set &lt;code&gt;BatchNorm&lt;/code&gt;&amp;rsquo;s momentum to the value equivalent to Keras&amp;rsquo; &lt;code&gt;momentum=0.99&lt;/code&gt;, and count &lt;code&gt;steps_per_epoch&lt;/code&gt; the same way Keras does. Multi-species training also matches Keras&amp;rsquo; handling of a batch&amp;rsquo;s inactive species: its loss term stays in the graph as a zero-weighted zero rather than being dropped, so weight decay still applies to every head on every step as it does in Keras, and each batch is drawn from a single species rather than mixed across species. A cross-framework test suite runs one training step through numerically identical converted weights in both frameworks — at toy scale, and, for all four released checkpoints, at the real published scale (&lt;code&gt;num_blocks=24&lt;/code&gt;, &lt;code&gt;model_dim=1024&lt;/code&gt;) on GPU against Illumina&amp;rsquo;s own SavedModels — and checks agreement on the loss, gradients, AdamW parameter deltas, and BatchNorm running-stat updates via a per-tensor cosine-similarity/relative-L2 pass rate rather than strict elementwise tolerances. All four real-checkpoint configurations pass; for the base (&lt;code&gt;hg38&lt;/code&gt;, &lt;code&gt;hg38_mm10&lt;/code&gt;) checkpoints, forward-pass prediction agreement is cosine = 1.0000 with relative L2 under 0.5%. The one remaining, characterized divergence is framework-inherent rather than a porting gap: Keras&amp;rsquo; &lt;code&gt;AdamW&lt;/code&gt; places its &lt;code&gt;epsilon&lt;/code&gt; term differently than PyTorch&amp;rsquo;s, transiently damping its first ~1,000 optimizer steps more strongly even with a matching &lt;code&gt;epsilon&lt;/code&gt;. See &lt;code&gt;notes/implementation.md&lt;/code&gt; and &lt;code&gt;docs/training.md&lt;/code&gt; in the repo for the full derivation.&lt;/p&gt;&#10;&lt;h2 id="what-can-you-do-with-this"&gt;What Can You Do With This?&#10;&lt;/h2&gt;&lt;p&gt;Beyond variant scoring, &lt;code&gt;load_pretrained()&lt;/code&gt; exposes the full model for anything you&amp;rsquo;d normally do with a PyTorch sequence model — track prediction, embeddings, and DeepLIFT/SHAP attribution [3]:&lt;/p&gt;&#10;&lt;div class="code-block" data-lang="python"&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;torch&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="nn"&gt;torch.nn&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nn"&gt;nn&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;tangermeme.deep_lift_shap&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;deep_lift_shap&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;promoterai_torch.dataset&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;onehot_encode&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="nn"&gt;promoterai_torch.utils&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_pretrained&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;load_pretrained&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;models/promoterAI_v1_hg38_mm10_finetune.pt&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;eval&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;seq&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;ACGT&amp;#34;&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;input_length&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# replace with your sequence&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;from_numpy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;onehot_encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;seq&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;unsqueeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# (1, L, 4)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;no_grad&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;predictions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# tuple of (1, output_length, n_tracks) per species head&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;embeddings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# (1, input_length, model_dim), final MetaFormer block output&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# DeepLIFT/SHAP via tangermeme: every non-linearity is a distinct, named nn.ReLU()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# instance, which is exactly what tangermeme&amp;#39;s deep_lift_shap requires.&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;PromoterAIWrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="fm"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;super&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="fm"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;forward&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="c1"&gt;# x: (B, 4, L) channels-first&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;self&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="c1"&gt;# PromoterAI expects (B, L, 4)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;dim&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;unsqueeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# (B, 1) — mean over positions and tracks&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;wrapper&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;PromoterAIWrapper&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;x_chfirst&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="c1"&gt;# (1, 4, input_length), channels-first for tangermeme&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;attributions&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;deep_lift_shap&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wrapper&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x_chfirst&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_shuffles&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;device&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;cuda&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;batch_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&#10;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# attributions: (1, 4, input_length) — per-position, per-base importance&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;Track prediction returns per-position predictions for all 498 human tracks the model was trained on (histone marks, TF ChIP-seq, ATAC-seq, RNA-seq), plus the 472-track mouse head; embeddings are the per-position hidden state after the final MetaFormer block, &lt;code&gt;(B, L, model_dim)&lt;/code&gt;. The &lt;code&gt;deep_lift_shap&lt;/code&gt; call above (transposing to channels-first, reducing the output heads to a scalar via the wrapper) is what produces per-base attribution maps like the one below:&lt;/p&gt;&#10;&#10;&#9;&#9;&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-014/deepliftshap.png" class="image-link" data-pswp-width="1324" data-pswp-height="446"&gt;&#10;&#9;&#9;&lt;img src="https://genomicsxai.github.io/blogs/2026-014/deepliftshap.png.optimized.webp" width="700px" height="235"loading="lazy"&#10;&#9;&#9;&#9;alt="DeepLIFT/SHAP contribution track across a 20 kb window around the SFSWAP promoter, with a zoomed-in per-base sequence logo over the 200 bp region of interest showing several high-contribution motif-like clusters."&#10;&#9;&#9;&#9;data-title-escaped="Per-base DeepLIFT/SHAP contribution scores at the SFSWAP promoter (chr12:131,700,849–131,721,329), zoomed into the 200 bp region of interest (chr12:131,710,989–131,711,189)."&gt;&#10;&#9;&#9;&lt;/a&gt;&lt;figcaption&gt;Per-base DeepLIFT/SHAP contribution scores at the SFSWAP promoter (chr12:131,700,849–131,721,329), zoomed into the 200 bp region of interest (chr12:131,710,989–131,711,189).&lt;/figcaption&gt;&lt;/figure&gt;&lt;p&gt;Fair warning on cost: DeepLIFT/SHAP on this model is not cheap — at TF32 with &lt;code&gt;n_shuffles=20&lt;/code&gt; and &lt;code&gt;batch_size=1&lt;/code&gt;, expect ~92s and ~71GB of VRAM per sequence on an A100 80GB. For an order-of-magnitude anchor (not a like-for-like benchmark — different attribution method, hardware, and input-window size), the &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-011/" target="_blank" rel="noopener"&#10; &gt;Cherimoya post&lt;/a&gt; [12] measured &lt;em&gt;in silico&lt;/em&gt; mutagenesis on a 1 kb locus at 0.08s (Cherimoya), 0.38s (ChromBPNet), 1.24s (AlphaGenome 2kb), 64.5s (Borzoi), and 324.3s (AlphaGenome) on an H200 GPU (bf16, &lt;code&gt;torch.compile&lt;/code&gt;); peak batch-1 VRAM for a full-sequence forward pass on the same hardware was 0.14GB, 0.19GB, 17.9GB, and 117GB for Cherimoya, ChromBPNet, Borzoi, and AlphaGenome (1Mb) respectively.&lt;/p&gt;&#10;&lt;h2 id="training-and-fine-tuning"&gt;Training and Fine-Tuning&#10;&lt;/h2&gt;&lt;p&gt;The repo also includes the full training pipeline, not just inference: HDF5 preprocessing of track and sequence data per chromosome, multi-GPU training via &lt;code&gt;torchrun&lt;/code&gt;, checkpoint/resume handling, and a fine-tuning script that trains only the first output head on a variant set (matching PromoterAI&amp;rsquo;s own fine-tuning protocol on GTEx outlier data) while keeping the rest of the backbone — including BatchNorm statistics — frozen in inference mode.&lt;/p&gt;&#10;&lt;div class="code-block" data-lang="sh"&gt;&#10;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-sh" data-lang="sh"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;promoterai-torch train &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --checkpoint_folder checkpoints/run1 &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --hdf5_human_folder data/hdf5/human &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --input_length &lt;span class="m"&gt;20480&lt;/span&gt; --output_length &lt;span class="m"&gt;4096&lt;/span&gt; &lt;span class="se"&gt;\&#10;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; --num_blocks &lt;span class="m"&gt;24&lt;/span&gt; --model_dim &lt;span class="m"&gt;1024&lt;/span&gt; --batch_size &lt;span class="m"&gt;32&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&#10;&lt;/div&gt;&#10;&lt;p&gt;This hasn&amp;rsquo;t been used to reproduce Illumina&amp;rsquo;s exact published model from scratch — that would require their full training corpus — but it has been verified to run end-to-end and to match the original&amp;rsquo;s documented training/fine-tuning behavior wherever that behavior is checkable.&lt;/p&gt;&#10;&lt;h2 id="code-and-tutorials"&gt;Code and Tutorials&#10;&lt;/h2&gt;&lt;ul&gt;&#10;&lt;li&gt;Repository: &lt;a class="link" href="https://github.com/genomicsxai/promoterai-torch" target="_blank" rel="noopener"&#10; &gt;github.com/genomicsxai/promoterai-torch&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;PyPI: &lt;a class="link" href="https://pypi.org/project/promoterai-torch/" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;promoterai-torch&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Worked examples (paper benchmark reproduction, track-level parity checks, TERT/SFSWAP/DNAJC9 notebooks): &lt;a class="link" href="https://github.com/genomicsxai/promoterai-torch/tree/main/examples" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;examples/&lt;/code&gt;&lt;/a&gt;&lt;/li&gt;&#10;&lt;/ul&gt;&#10;&lt;h2 id="license"&gt;License&#10;&lt;/h2&gt;&lt;p&gt;&lt;code&gt;promoterai-torch&lt;/code&gt; is an independent reimplementation, not an official Illumina product or publication. Its PyTorch code was written from PromoterAI&amp;rsquo;s published architecture description and released hyperparameters — the depth, width, dilation schedule, output-head structure, and naming scheme needed for the two implementations to be numerically compatible — rather than by transliterating Illumina&amp;rsquo;s TensorFlow/Keras source. &lt;code&gt;promoterai-torch&lt;/code&gt;&amp;rsquo;s own code is MIT-licensed.&lt;/p&gt;&#10;&lt;p&gt;The original PromoterAI codebase is released under the &lt;a class="link" href="https://polyformproject.org/licenses/strict/1.0.0/" target="_blank" rel="noopener"&#10; &gt;PolyForm Strict License 1.0.0&lt;/a&gt;, which permits noncommercial, research, and educational use but withholds the right to distribute the software or &amp;ldquo;changes or new works based on&amp;rdquo; it. Illumina&amp;rsquo;s pretrained weights and precomputed variant scores are gated separately, under Illumina&amp;rsquo;s own academic-only data license — see &lt;a class="link" href="https://github.com/Illumina/PromoterAI" target="_blank" rel="noopener"&#10; &gt;Illumina/PromoterAI&lt;/a&gt; for the license agreement and commercial-licensing contact (&lt;code&gt;AI_licensing@illumina.com&lt;/code&gt;). This repo contains no Illumina code, models, or scores; if you convert and use the original weights yourself, you are responsible for complying with Illumina&amp;rsquo;s terms. Converted checkpoints should not be redistributed.&lt;/p&gt;&#10;&lt;h2 id="acknowledgements"&gt;Acknowledgements&#10;&lt;/h2&gt;&lt;p&gt;This work builds directly on the architecture and training protocol described by Illumina&amp;rsquo;s PromoterAI team, and on &lt;a class="link" href="https://github.com/jmschrei/tangermeme" target="_blank" rel="noopener"&#10; &gt;&lt;code&gt;tangermeme&lt;/code&gt;&lt;/a&gt; for attribution tooling.&lt;/p&gt;&#10;&lt;h2 id="references"&gt;References&#10;&lt;/h2&gt;&lt;ol&gt;&#10;&lt;li&gt;Jaganathan, K., Ersaro, N., Novakovsky, G. et al. Predicting expression-altering promoter mutations with deep learning. &lt;em&gt;Science&lt;/em&gt; 389, eads7373 (2025). &lt;a class="link" href="https://doi.org/10.1126/science.ads7373" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1126/science.ads7373&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Illumina/PromoterAI (official TensorFlow implementation). &lt;a class="link" href="https://github.com/Illumina/PromoterAI" target="_blank" rel="noopener"&#10; &gt;https://github.com/Illumina/PromoterAI&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Schreiber, J. tangermeme: A toolkit for understanding cis-regulatory logic using deep learning models. &lt;em&gt;bioRxiv&lt;/em&gt; (2025). &lt;a class="link" href="https://www.biorxiv.org/content/10.1101/2025.08.08.669296v2" target="_blank" rel="noopener"&#10; &gt;https://www.biorxiv.org/content/10.1101/2025.08.08.669296v2&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Huang, F. W., Hodis, E., Xu, M. J., Kryukov, G. V., Chin, L. &amp;amp; Garraway, L. A. Highly recurrent TERT promoter mutations in human melanoma. &lt;em&gt;Science&lt;/em&gt; 339, 957–959 (2013). &lt;a class="link" href="https://doi.org/10.1126/science.1229259" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1126/science.1229259&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Horn, S., Figl, A., Rachakonda, P. S. et al. TERT promoter mutations in familial and sporadic melanoma. &lt;em&gt;Science&lt;/em&gt; 339, 959–961 (2013). &lt;a class="link" href="https://doi.org/10.1126/science.1230062" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1126/science.1230062&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Killela, P. J., Reitman, Z. J., Jiao, Y. et al. TERT promoter mutations occur frequently in gliomas and a subset of tumors derived from cells with low rates of self-renewal. &lt;em&gt;Proc Natl Acad Sci USA&lt;/em&gt; 110, 6021–6026 (2013). &lt;a class="link" href="https://doi.org/10.1073/pnas.1303607110" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1073/pnas.1303607110&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Bell, R. J. A., Rube, H. T., Xavier-Magalhães, A. et al. Understanding TERT Promoter Mutations: A Common Path to Immortality. &lt;em&gt;Mol Cancer Res&lt;/em&gt; 14, 315–323 (2016). &lt;a class="link" href="https://doi.org/10.1158/1541-7786.MCR-16-0003" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1158/1541-7786.MCR-16-0003&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Schreiber, J., Lorbeer, F. K., Heinzl, M., Reiter, F., Rafanel, B., Lu, Y., Stark, A. &amp;amp; Noble, W. S. Programmatic design and editing of cis-regulatory elements. &lt;em&gt;bioRxiv&lt;/em&gt; (2025). &lt;a class="link" href="https://doi.org/10.1101/2025.04.22.650035" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1101/2025.04.22.650035&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Avsec, Ž., Agarwal, V., Visentin, D. et al. Effective gene expression prediction from sequence by integrating long-range interactions. &lt;em&gt;Nat Methods&lt;/em&gt; 18, 1196–1203 (2021). &lt;a class="link" href="https://doi.org/10.1038/s41592-021-01252-x" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1038/s41592-021-01252-x&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Hingerl, J. C., Karollus, A. &amp;amp; Gagneur, J. Flashzoi: an enhanced Borzoi for accelerated genomic analysis. &lt;em&gt;Bioinformatics&lt;/em&gt; 41, btaf467 (2025). &lt;a class="link" href="https://doi.org/10.1093/bioinformatics/btaf467" target="_blank" rel="noopener"&#10; &gt;https://doi.org/10.1093/bioinformatics/btaf467&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Bredikhin, D., Buendia, A., Kjellberg, M., Zou, C., Tu, X. &amp;amp; Kundaje, A. Porting AlphaGenome to PyTorch. &lt;em&gt;Genomics x AI Blog&lt;/em&gt; (2026). &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-004/" target="_blank" rel="noopener"&#10; &gt;https://genomicsxai.github.io/blogs/2026-004/&lt;/a&gt;&lt;/li&gt;&#10;&lt;li&gt;Ramirez, C., Aruva, A. M., Weng, Z. &amp;amp; Schreiber, J. Cherimoya: a lightweight genomic sequence-to-function model built for large-scale design and understanding. &lt;em&gt;Genomics x AI Blog&lt;/em&gt; (2026). &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-011/" target="_blank" rel="noopener"&#10; &gt;https://genomicsxai.github.io/blogs/2026-011/&lt;/a&gt;&lt;/li&gt;&#10;&lt;/ol&gt;&#10;</description></item></channel></rss>