<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>DNase-Seq on Genomics x AI</title><link>https://genomicsxai.github.io/tags/dnase-seq/</link><description>Recent content in DNase-Seq on Genomics x AI</description><generator>Hugo -- gohugo.io</generator><language>en</language><lastBuildDate>Tue, 04 Aug 2026 17:38:26 +0000</lastBuildDate><atom:link href="https://genomicsxai.github.io/tags/dnase-seq/index.xml" rel="self" type="application/rss+xml"/><item><title>[ENCODE GRAMMAR] Quickstart: Accessing and using the ENCODE GRAMMAR collection</title><link>https://genomicsxai.github.io/blogs/2026-013/</link><pubDate>Tue, 04 Aug 2026 00:00:00 +0000</pubDate><guid>https://genomicsxai.github.io/blogs/2026-013/</guid><description>&lt;img src="https://genomicsxai.github.io/" alt="Featured image of post [ENCODE GRAMMAR] Quickstart: Accessing and using the ENCODE GRAMMAR collection" /&gt;&lt;aside class="summary-box"&gt;
 &lt;h2 class="summary-box__title"&gt;Summary&lt;/h2&gt;
 &lt;div class="summary-box__body"&gt;
 &lt;p&gt;&lt;strong&gt;ENCODE GRAMMAR&lt;/strong&gt; (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations and Rules) is a collection of 3,865 experiment-specific deep learning model sets and derived sequence annotations that connect the ENCODE project&amp;rsquo;s human gene regulation maps from extensive biochemical profiling experiments to the underlying DNA sequence features that drive regulatory activity.&lt;/p&gt;
&lt;p&gt;This five-minute &lt;strong&gt;quickstart guide&lt;/strong&gt; shows how to find an ENCODE GRAMMAR model-set and its associated annotation and load its most commonly used outputs into an interactive genome browser. By following the steps below, you will create a browser session displaying an experimentally observed regulatory profile, the corresponding model-predicted profile, and a base-resolution sequence-contribution map.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Contributions&lt;/strong&gt;: &lt;em&gt;(Author order does not represent relative contribution)&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Vivekanandan Ramalingam&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vir@stanford.edu" &gt;vir@stanford.edu&lt;/a&gt;): BPNet model optimization and training, data uploads, general analysis&lt;/li&gt;
&lt;li&gt;Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;): ChromBPNet model training, MotifCompendium analysis&lt;/li&gt;
&lt;li&gt;Vivian Hecht&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vhecht@stanford.edu" &gt;vhecht@stanford.edu&lt;/a&gt;): ChromBPNet model training, model resource uploads, project management&lt;/li&gt;
&lt;li&gt;Aman Patel&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:patelas@stanford.edu" &gt;patelas@stanford.edu&lt;/a&gt;): Model resource uploads&lt;/li&gt;
&lt;li&gt;Anusri Pampari&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:anusri@stanford.edu" &gt;anusri@stanford.edu&lt;/a&gt;): ChromBPNet model development, ChromBPNet model training, data uploads&lt;/li&gt;
&lt;li&gt;Ziwei Chen&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:ziwei75@stanford.edu" &gt;ziwei75@stanford.edu&lt;/a&gt;): ReporterNet model development, ReporterNet model training, ChromBPNet model training&lt;/li&gt;
&lt;li&gt;Kelly Cochran&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:kcochran@stanford.edu" &gt;kcochran@stanford.edu&lt;/a&gt;): ProCapNet model development and training&lt;/li&gt;
&lt;li&gt;Adam He&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:ayhe@stanford.edu" &gt;ayhe@stanford.edu&lt;/a&gt;): ProCapNet user guide&lt;/li&gt;
&lt;li&gt;Surag Nair&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:surag@stanford.edu" &gt;surag@stanford.edu&lt;/a&gt;), Zahoor Zafrulla&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:zahoor@stanford.edu" &gt;zahoor@stanford.edu&lt;/a&gt;), Alex Tseng&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:amtseng@stanford.edu" &gt;amtseng@stanford.edu&lt;/a&gt;): BPNet refactoring&lt;/li&gt;
&lt;li&gt;Avanti Shrikumar&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:avanti@stanford.edu" &gt;avanti@stanford.edu&lt;/a&gt;), Jacob Schreiber&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:jmschr@stanford.edu" &gt;jmschr@stanford.edu&lt;/a&gt;), Alex Tseng&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:amtseng@stanford.edu" &gt;amtseng@stanford.edu&lt;/a&gt;): TF-MoDISco methods development and optimization&lt;/li&gt;
&lt;li&gt;Austin Wang&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:atwang@stanford.edu" &gt;atwang@stanford.edu&lt;/a&gt;): FiNeMo methods development&lt;/li&gt;
&lt;li&gt;Salil Deshpande&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:salil512@stanford.edu" &gt;salil512@stanford.edu&lt;/a&gt;), Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;): MotifCompendium methods development&lt;/li&gt;
&lt;li&gt;Abhimanyu Banerjee&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:manyu@stanford.edu" &gt;manyu@stanford.edu&lt;/a&gt;), Georgi K. Marinov&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:marinovg@stanford.edu" &gt;marinovg@stanford.edu&lt;/a&gt;): Zinc finger transcription factor analysis&lt;/li&gt;
&lt;li&gt;Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;), Vivian Hecht&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vhecht@stanford.edu" &gt;vhecht@stanford.edu&lt;/a&gt;), Vivekanandan Ramalingam&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vir@stanford.edu" &gt;vir@stanford.edu&lt;/a&gt;): Blog posts&lt;/li&gt;
&lt;li&gt;Anshul Kundaje&lt;sup&gt;1&lt;/sup&gt;* (&lt;a class="link" href="mailto:akundaje@stanford.edu" &gt;akundaje@stanford.edu&lt;/a&gt;): Conceptualization, Project management, Mentoring, Funding, Blog post editing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;&lt;sup&gt;1&lt;/sup&gt;Stanford University, *Correspondence.&lt;/em&gt;&lt;/p&gt;

 &lt;/div&gt;
&lt;/aside&gt;

&lt;hr&gt;

 &lt;blockquote&gt;
 &lt;p&gt;This is the second post in a series on &lt;strong&gt;ENCODE GRAMMAR&lt;/strong&gt;. The series will cover:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-012/" target="_blank" rel="noopener"
 &gt;ENCODE GRAMMAR: The ENCODE deep learning model resource for decoding the DNA sequence logic of genomic regulatory elements&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Accessing and using the ENCODE GRAMMAR collection: A quickstart guide (this post)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Interpreting regulatory DNA with deep learning models&lt;/li&gt;
&lt;li&gt;The transcription factor binding GRAMMAR resource&lt;/li&gt;
&lt;li&gt;The chromatin accessibility GRAMMAR resource&lt;/li&gt;
&lt;li&gt;Predicting the effects of noncoding genetic variants&lt;/li&gt;
&lt;li&gt;MotifCompendium - a unified lexicon of regulatory sequence motifs&lt;/li&gt;
&lt;li&gt;Contrasting regulatory sequence codes across assays and cell types&lt;/li&gt;
&lt;li&gt;Building a production-scale model atlas in an academic setting&lt;/li&gt;
&lt;/ol&gt;

 &lt;/blockquote&gt;
&lt;h2 id="quick-start-guide-5-min"&gt;Quick-start guide (5 min)
&lt;/h2&gt;&lt;p&gt;ENCODE GRAMMAR transforms individual ENCODE experiments into experiment-specific BPNet-family model sets together with predicted regulatory profiles, sequence-contribution maps, predictive motifs, motif instances, and variant-effect predictions. Readers who are new to the resource may wish to begin with the &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-012/" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE GRAMMAR overview&lt;/strong&gt;&lt;/a&gt;, which explains the biological motivation, model families, interpretation workflow, and complete collection of released products.&lt;/p&gt;
&lt;p&gt;Below, we explain how to navigate an ENCODE GRAMMAR model-set annotation page and load several commonly used model outputs into the WashU Epigenome Browser. Visualizing these tracks is a useful first step before designing larger-scale quantitative analyses.&lt;/p&gt;
&lt;p&gt;We use the example of a &lt;strong&gt;ChromBPNet model set generated from an ATAC-seq experiment in K562 cells (ENCSR893SUD)&lt;/strong&gt;.&lt;/p&gt;
&lt;h3 id="step-1-find-the-model-set-annotation-page"&gt;Step 1: Find the model-set annotation page
&lt;/h3&gt;&lt;p&gt;Open the ChromBPNet model-set annotation for K562 ATAC-seq: &lt;a class="link" href="https://www.encodeproject.org/annotations/ENCSR893SUD/" target="_blank" rel="noopener"
 &gt;https://www.encodeproject.org/annotations/ENCSR893SUD/&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Model-set annotations associated with an experiment can also be found from the corresponding experiment summary page.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig1.png" class="image-link" data-pswp-width="1033" data-pswp-height="704"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig1.png" width="1033" height="704"loading="lazy"
			alt="Figure 1"
			title="ENCODE Portal annotation page for the K562 ATAC-seq ChromBPNet model set ENCSR893SUD." data-title-escaped="ENCODE Portal annotation page for the K562 ATAC-seq ChromBPNet model set ENCSR893SUD." class="gallery-image" data-flex-grow="146" data-flex-basis="352px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;A searchable list of all ENCODE annotations is available at &lt;a class="link" href="https://www.encodeproject.org/annotations/" target="_blank" rel="noopener"
 &gt;https://www.encodeproject.org/annotations/&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="step-2-find-the-files-of-interest"&gt;Step 2: Find the files of interest
&lt;/h3&gt;&lt;p&gt;Scroll to the middle of the annotation page and select the &lt;strong&gt;File details&lt;/strong&gt; tab to browse the available files.&lt;/p&gt;
&lt;p&gt;For an initial exploration of a ChromBPNet model set, we recommend viewing three complementary tracks:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;the &lt;strong&gt;normalized observed signal profile&lt;/strong&gt;, representing the experimentally measured chromatin-accessibility profile;&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;normalized predicted signal profile&lt;/strong&gt;, representing the regulatory profile predicted by ChromBPNet from DNA sequence; and&lt;/li&gt;
&lt;li&gt;the &lt;strong&gt;counts sequence-contribution scores&lt;/strong&gt;, estimating how much each DNA base contributes to the model&amp;rsquo;s prediction of total accessibility.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig2-3.png" class="image-link" data-pswp-width="3299" data-pswp-height="2250"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig2-3.png" width="3299" height="2250"loading="lazy"
			alt="Figure 2"
			title="File details tab for the K562 ATAC-seq ChromBPNet model set ENCSR893SUD." data-title-escaped="File details tab for the K562 ATAC-seq ChromBPNet model set ENCSR893SUD." class="gallery-image" data-flex-grow="146" data-flex-basis="351px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;The following steps require the URLs of these bigWig files. Right-click the &lt;strong&gt;download icon&lt;/strong&gt; beside a file and select &lt;strong&gt;Copy link address&lt;/strong&gt;. The links for this example are provided here:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Normalized observed signal profile&lt;/strong&gt;: &lt;a class="link" href="https://www.encodeproject.org/files/ENCFF880ZUI/@@download/ENCFF880ZUI.bigWig" target="_blank" rel="noopener"
 &gt;https://www.encodeproject.org/files/ENCFF880ZUI/@@download/ENCFF880ZUI.bigWig&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Normalized predicted signal profile&lt;/strong&gt;: &lt;a class="link" href="https://www.encodeproject.org/files/ENCFF296ICJ/@@download/ENCFF296ICJ.bigWig" target="_blank" rel="noopener"
 &gt;https://www.encodeproject.org/files/ENCFF296ICJ/@@download/ENCFF296ICJ.bigWig&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Counts sequence contribution scores&lt;/strong&gt;: &lt;a class="link" href="https://www.encodeproject.org/files/ENCFF407GCO/@@download/ENCFF407GCO.bigWig" target="_blank" rel="noopener"
 &gt;https://www.encodeproject.org/files/ENCFF407GCO/@@download/ENCFF407GCO.bigWig&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This quickstart focuses on these three tracks. Additional ENCODE GRAMMAR products, including predictive motif-instance annotations, can be accessed through the ENCODE Portal and the UCSC Track Hub linked below.&lt;/p&gt;
&lt;p&gt;For descriptions of the complete collection of files and model-derived products, see the &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-012/" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE GRAMMAR overview&lt;/strong&gt;&lt;/a&gt; and the &lt;a class="link" href="https://doi.org/10.64898/2026.07.06.731365" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE 4 preprint&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="step-3-load-the-bigwigs-into-the-washu-genome-browser"&gt;Step 3: Load the bigwigs into the WashU genome browser
&lt;/h3&gt;&lt;p&gt;Navigate to the &lt;a class="link" href="https://epigenomegateway.wustl.edu/browser2022/" target="_blank" rel="noopener"
 &gt;WashU Epigenome Browser&lt;/a&gt;. On the home page, find the &lt;strong&gt;Human&lt;/strong&gt; section and select &lt;strong&gt;hg38&lt;/strong&gt; to open a new browser session.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig3.png" class="image-link" data-pswp-width="1252" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig3.png" width="1252" height="550"loading="lazy"
			alt="Figure 3"
			title="WashU Epigenome Browser home page with the human hg38 genome assembly selected." data-title-escaped="WashU Epigenome Browser home page with the human hg38 genome assembly selected." class="gallery-image" data-flex-grow="227" data-flex-basis="546px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;In the new browser session, select the &lt;strong&gt;Tracks&lt;/strong&gt; icon at the top of the page and then choose &lt;strong&gt;Remote Tracks&lt;/strong&gt; from the dropdown menu.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig4.png" class="image-link" data-pswp-width="1224" data-pswp-height="270"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig4.png" width="1224" height="270"loading="lazy"
			alt="Figure 4"
			title="Opening the Remote Tracks menu in the WashU Epigenome Browser." data-title-escaped="Opening the Remote Tracks menu in the WashU Epigenome Browser." class="gallery-image" data-flex-grow="453" data-flex-basis="1088px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;A window for adding remote tracks will appear:&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig5.png" class="image-link" data-pswp-width="1214" data-pswp-height="259"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig5.png" width="1214" height="259"loading="lazy"
			alt="Figure 5"
			title="WashU Epigenome Browser window for adding remote tracks." data-title-escaped="WashU Epigenome Browser window for adding remote tracks." class="gallery-image" data-flex-grow="468" data-flex-basis="1124px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;First, add the &lt;strong&gt;normalized observed signal profile&lt;/strong&gt; and &lt;strong&gt;normalized predicted signal profile&lt;/strong&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Copy and paste the first URL from Step 2.&lt;/li&gt;
&lt;li&gt;Add an informative label, such as &lt;code&gt;Observed accessibility&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;Submit&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;Add another track&lt;/strong&gt; and repeat the process for the predicted profile, using a label such as &lt;code&gt;ChromBPNet predicted accessibility&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig6.png" class="image-link" data-pswp-width="98" data-pswp-height="60"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig6.png" width="98" height="60"loading="lazy"
			alt="Figure 6"
			title="Submitting a remote bigWig track to the WashU Epigenome Browser." data-title-escaped="Submitting a remote bigWig track to the WashU Epigenome Browser." class="gallery-image" data-flex-grow="163" data-flex-basis="392px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig7.png" class="image-link" data-pswp-width="261" data-pswp-height="59"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig7.png" width="261" height="59"loading="lazy"
			alt="Figure 7"
			title="Selecting Add another track to load an additional remote file." data-title-escaped="Selecting Add another track to load an additional remote file." class="gallery-image" data-flex-grow="442" data-flex-basis="1061px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;For the third bigWig—the counts sequence-contribution scores—change the track type to &lt;strong&gt;Dynseq&lt;/strong&gt;. Select the &lt;strong&gt;Track type&lt;/strong&gt; menu and choose &lt;strong&gt;Dynseq (dynamic sequence)&lt;/strong&gt; from the list.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig8.png" class="image-link" data-pswp-width="1211" data-pswp-height="318"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig8.png" width="1211" height="318"loading="lazy"
			alt="Figure 8"
			title="Selecting the Dynseq track type for a sequence-contribution bigWig." data-title-escaped="Selecting the Dynseq track type for a sequence-contribution bigWig." class="gallery-image" data-flex-grow="380" data-flex-basis="913px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;Dynseq displays each nucleotide as a letter whose height and direction reflect its contribution score. Positive scores indicate bases that increase the model prediction relative to the reference, whereas negative scores indicate bases that decrease it. The nucleotide letters appear only after zooming in sufficiently. Clusters of bases with large contribution scores often correspond to predictive regulatory sequence features, including transcription-factor motif instances.&lt;/p&gt;
&lt;p&gt;Add an informative label, such as &lt;code&gt;Counts sequence-contribution map&lt;/code&gt;, and submit the track.&lt;/p&gt;
&lt;p&gt;The completed browser session should resemble the example below, with the experimentally observed profile, model-predicted profile, and sequence-contribution map aligned at the same genomic locus:&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-013/Fig9.png" class="image-link" data-pswp-width="1275" data-pswp-height="502"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-013/Fig9.png" width="1275" height="502"loading="lazy"
			alt="Figure 9"
			title="WashU Epigenome Browser session displaying observed and ChromBPNet-predicted accessibility profiles together with a base-resolution sequence-contribution map." data-title-escaped="WashU Epigenome Browser session displaying observed and ChromBPNet-predicted accessibility profiles together with a base-resolution sequence-contribution map." class="gallery-image" data-flex-grow="253" data-flex-basis="609px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p&gt;For more information about configuring tracks, navigating loci, and sharing sessions, see the &lt;a class="link" href="https://epigenomegateway.readthedocs.io/en/latest/usage.html" target="_blank" rel="noopener"
 &gt;WashU Epigenome Browser documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="how-else-can-i-use-the-resources"&gt;How else can I use the resources?
&lt;/h2&gt;&lt;p&gt;All ENCODE data, model sets, and model-derived sequence annotations are openly available through the &lt;a class="link" href="https://www.encodeproject.org/search/?type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;annotation_type=ChromBPNet-model&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE Portal&lt;/strong&gt;&lt;/a&gt;. The complete ENCODE GRAMMAR resource contains 3,865 experiment-specific model sets spanning TF binding, chromatin accessibility, transcription initiation, and high-throughput reporter activity.&lt;/p&gt;
&lt;p&gt;Additional access points include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Model-sets and sequence annotations:&lt;/strong&gt; Browse released &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=BPNet&amp;amp;type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;status=released&amp;amp;assay_term_name=ChIP-seq" target="_blank" rel="noopener"
 &gt;BPNet model sets&lt;/a&gt;, &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ChromBPNet&amp;amp;type=Annotation&amp;amp;annotation_type=ChromBPNet-model&amp;amp;organism.scientific_name=Homo&amp;#43;sapiens&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;ChromBPNet model sets&lt;/a&gt;, &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ProCapNet&amp;amp;type=Annotation" target="_blank" rel="noopener"
 &gt;ProCapNet model sets&lt;/a&gt;, and &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ReporterNet&amp;amp;type=Annotation&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;ReporterNet model sets&lt;/a&gt; on the ENCODE Portal.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;HuggingFace model zoo:&lt;/strong&gt; Download ENCODE GRAMMAR models from &lt;a class="link" href="https://huggingface.co/collections/kundajelab/encode-bpnet-models" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;Hugging Face&lt;/strong&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Predictions and sequence annotations:&lt;/strong&gt; Explore model predictions, sequence-contribution maps, and predictive motif instances through the &lt;a class="link" href="https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38&amp;amp;hubUrl=https://kundajelab.github.io/ucsc-trackhub-encode.github.io/hub.txt" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;UCSC Track Hub&lt;/strong&gt;&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unified motif lexicon:&lt;/strong&gt; Access the &lt;a class="link" href="https://www.encodeproject.org/annotations/ENCSR091GRD/" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE Motif Compendium&lt;/strong&gt;&lt;/a&gt;, which organizes related predictive motifs discovered across all ENCODE GRAMMAR models.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Software:&lt;/strong&gt; Train and interpret new models using the open-source &lt;a class="link" href="https://github.com/kundajelab/bpnet/" target="_blank" rel="noopener"
 &gt;BPNet&lt;/a&gt;, &lt;a class="link" href="https://github.com/kundajelab/chrombpnet/" target="_blank" rel="noopener"
 &gt;ChromBPNet&lt;/a&gt;, and &lt;a class="link" href="https://github.com/kundajelab/ProCapNet/" target="_blank" rel="noopener"
 &gt;ProCapNet&lt;/a&gt; repositories.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;User guide:&lt;/strong&gt; We are developing an &lt;em&gt;interactive&lt;/em&gt; guide to help users navigate and interpret the resource (&lt;em&gt;work in progress&lt;/em&gt;).&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Overview and publications:&lt;/strong&gt; Read the &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-012/" target="_blank" rel="noopener"
 &gt;ENCODE GRAMMAR overview&lt;/a&gt;, &lt;a class="link" href="https://doi.org/10.64898/2026.07.06.731365" target="_blank" rel="noopener"
 &gt;ENCODE 4 preprint&lt;/a&gt;, the &lt;a class="link" href="https://doi.org/10.1038/s41588-021-00782-6" target="_blank" rel="noopener"
 &gt;BPNet paper&lt;/a&gt;, &lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;ChromBPNet preprint&lt;/a&gt; and &lt;a class="link" href="https://doi.org/10.1101/2024.05.28.596138" target="_blank" rel="noopener"
 &gt;ProCapNet preprint&lt;/a&gt; for additional detail.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This quickstart covers only one way to explore the resource. Future (weekly) posts in this series will describe how to interpret sequence-contribution maps, compare predictive motifs across cellular contexts, and use the models to estimate the molecular effects of noncoding genetic variants.&lt;/p&gt;
&lt;h2 id="references"&gt;References
&lt;/h2&gt;&lt;ol&gt;
&lt;li&gt;The ENCODE Project Consortium et al. The Encyclopedia of DNA Elements. &lt;em&gt;bioRxiv&lt;/em&gt; 2026.07.06.731365 (2026) (&lt;a class="link" href="https://doi.org/10.64898/2026.07.06.731365" target="_blank" rel="noopener"
 &gt;https://doi.org/10.64898/2026.07.06.731365&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Yun, C. M. et al. A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays. (2025) doi:10.5281/zenodo.17179111. (&lt;a class="link" href="https://doi.org/10.5281/zenodo.17179111" target="_blank" rel="noopener"
 &gt;https://doi.org/10.5281/zenodo.17179111&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Avsec, Ž. et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. &lt;em&gt;Nat Genet&lt;/em&gt; 53, 354—366 (2021). (&lt;a class="link" href="https://doi.org/10.1038/s41588-021-00782-6" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1038/s41588-021-00782-6&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Pampari, A. et al. ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants. &lt;em&gt;bioRxiv&lt;/em&gt; 2024.12.25.630221 (2024). (&lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1101/2024.12.25.630221&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Cochran, K. et al. Dissecting the cis-regulatory syntax of transcription initiation with deep learning. &lt;em&gt;bioRxiv&lt;/em&gt; 2024.05.28.596138 (2024). (&lt;a class="link" href="https://doi.org/10.1101/2024.05.28.596138" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1101/2024.05.28.596138&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Shrikumar, A., Greenside, P. &amp;amp; Kundaje, A. Learning Important Features Through Propagating Activation Differences. &lt;em&gt;arXIV&lt;/em&gt; (2019).(&lt;a class="link" href="https://doi.org/10.48550/arXiv.1704.02685" target="_blank" rel="noopener"
 &gt;https://doi.org/10.48550/arXiv.1704.02685&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Lundberg, S. M. &amp;amp; Lee, S.-I. A unified approach to interpreting model predictions. in &lt;em&gt;Proceedings of the 31st International Conference on Neural Information Processing Systems&lt;/em&gt; 4768–4777 (Curran Associates Inc., Red Hook, NY, USA, 2017). (&lt;a class="link" href="https://dl.acm.org/doi/10.5555/3295222.3295230" target="_blank" rel="noopener"
 &gt;https://dl.acm.org/doi/10.5555/3295222.3295230&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Shrikumar, A. et al. Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5. &lt;em&gt;arXiv&lt;/em&gt; (2020) (&lt;a class="link" href="https://doi.org/10.48550/arXiv.1811.00416" target="_blank" rel="noopener"
 &gt;https://doi.org/10.48550/arXiv.1811.00416&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;</description></item><item><title>ENCODE GRAMMAR: A deep learning model resource for decoding the DNA sequence logic of regulatory elements in the human genome</title><link>https://genomicsxai.github.io/blogs/2026-012/</link><pubDate>Tue, 04 Aug 2026 00:00:00 +0000</pubDate><guid>https://genomicsxai.github.io/blogs/2026-012/</guid><description>&lt;img src="https://genomicsxai.github.io/" alt="Featured image of post ENCODE GRAMMAR: A deep learning model resource for decoding the DNA sequence logic of regulatory elements in the human genome" /&gt;&lt;aside class="summary-box"&gt;
 &lt;h2 class="summary-box__title"&gt;Summary&lt;/h2&gt;
 &lt;div class="summary-box__body"&gt;
 &lt;p&gt;Over two decades, the Encyclopedia of DNA Elements (ENCODE) Consortium has used diverse functional genomics experiments to map millions of regions in the human and mouse genomes that help control when, where, and how strongly genes are turned on. These maps characterize the biochemical properties and activity of regulatory DNA elements across thousands of cell types and tissues. However, they do not fully explain how this activity is encoded in the DNA sequence itself: which individual DNA letters are important, how combinations and arrangements of letters cause a region to behave differently across cell types, or how a genetic variant might alter its activity.&lt;/p&gt;
&lt;p&gt;To help answer these questions, we developed a family of deep learning models that use DNA sequences as inputs to predict associated genome-wide biochemical activity measured by diverse experiments. We also developed a framework of model interpretation methods, allowing us to identify the sequence features and regulatory rules that drive model predictions.&lt;/p&gt;
&lt;p&gt;Today, we release &lt;strong&gt;GRAMMAR&lt;/strong&gt; (Genomic Regulatory Atlas of sequence Models, Motifs, Annotations and Rules), a collection of 3,865 experiment-specific model sets across the ENCODE data compendium spanning several layers of gene regulation, including the binding of regulatory proteins to DNA, chromatin accessibility, transcription initiation, and regulatory activity measured using high-throughput reporter assays. For each experiment, we also release model-predicted, de-noised biochemical activity profiles at single-base resolution; predicted contributions of individual DNA bases to the activity of each regulatory element in each cellular context; recurring predictive DNA patterns (motifs) learned by the models; the genomic locations of predictive motif instances; and genome-browser tracks that make these outputs easy to visualize and explore.&lt;/p&gt;
&lt;p&gt;Together, these resources transform thousands of ENCODE experiments into a reusable and interpretable atlas of the cell-context-specific DNA sequence rules that shape gene regulation in the human genome. In this first post of a broader series, we introduce the ENCODE GRAMMAR resource and show how predictions and sequence annotations derived from multiple models can be integrated to decode the sequence basis of regulatory element activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Contributions&lt;/strong&gt;: &lt;em&gt;(Author order does not represent relative contribution)&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Vivekanandan Ramalingam&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vir@stanford.edu" &gt;vir@stanford.edu&lt;/a&gt;): BPNet model optimization and training, data uploads, general analysis&lt;/li&gt;
&lt;li&gt;Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;): ChromBPNet model training, MotifCompendium analysis&lt;/li&gt;
&lt;li&gt;Vivian Hecht&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vhecht@stanford.edu" &gt;vhecht@stanford.edu&lt;/a&gt;): ChromBPNet model training, model resource uploads, project management&lt;/li&gt;
&lt;li&gt;Aman Patel&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:patelas@stanford.edu" &gt;patelas@stanford.edu&lt;/a&gt;): Model resource uploads&lt;/li&gt;
&lt;li&gt;Anusri Pampari&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:anusri@stanford.edu" &gt;anusri@stanford.edu&lt;/a&gt;): ChromBPNet model development, ChromBPNet model training, data uploads&lt;/li&gt;
&lt;li&gt;Ziwei Chen&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:ziwei75@stanford.edu" &gt;ziwei75@stanford.edu&lt;/a&gt;): ReporterNet model development, ReporterNet model training, ChromBPNet model training&lt;/li&gt;
&lt;li&gt;Kelly Cochran&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:kcochran@stanford.edu" &gt;kcochran@stanford.edu&lt;/a&gt;): ProCapNet model development and training&lt;/li&gt;
&lt;li&gt;Adam He&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:ayhe@stanford.edu" &gt;ayhe@stanford.edu&lt;/a&gt;): ProCapNet user guide&lt;/li&gt;
&lt;li&gt;Surag Nair&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:surag@stanford.edu" &gt;surag@stanford.edu&lt;/a&gt;), Zahoor Zafrulla&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:zahoor@stanford.edu" &gt;zahoor@stanford.edu&lt;/a&gt;), Alex Tseng&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:amtseng@stanford.edu" &gt;amtseng@stanford.edu&lt;/a&gt;): BPNet refactoring&lt;/li&gt;
&lt;li&gt;Avanti Shrikumar&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:avanti@stanford.edu" &gt;avanti@stanford.edu&lt;/a&gt;), Jacob Schreiber&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:jmschr@stanford.edu" &gt;jmschr@stanford.edu&lt;/a&gt;), Alex Tseng&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:amtseng@stanford.edu" &gt;amtseng@stanford.edu&lt;/a&gt;): TF-MoDISco methods development and optimization&lt;/li&gt;
&lt;li&gt;Austin Wang&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:atwang@stanford.edu" &gt;atwang@stanford.edu&lt;/a&gt;): FiNeMo methods development&lt;/li&gt;
&lt;li&gt;Salil Deshpande&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:salil512@stanford.edu" &gt;salil512@stanford.edu&lt;/a&gt;), Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;): MotifCompendium methods development&lt;/li&gt;
&lt;li&gt;Abhimanyu Banerjee&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:manyu@stanford.edu" &gt;manyu@stanford.edu&lt;/a&gt;), Georgi K. Marinov&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:marinovg@stanford.edu" &gt;marinovg@stanford.edu&lt;/a&gt;): Zinc finger transcription factor analysis&lt;/li&gt;
&lt;li&gt;Chang M. Yun&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:chang.m.yun@stanford.edu" &gt;chang.m.yun@stanford.edu&lt;/a&gt;), Vivian Hecht&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vhecht@stanford.edu" &gt;vhecht@stanford.edu&lt;/a&gt;), Vivekanandan Ramalingam&lt;sup&gt;1&lt;/sup&gt; (&lt;a class="link" href="mailto:vir@stanford.edu" &gt;vir@stanford.edu&lt;/a&gt;): Blog posts&lt;/li&gt;
&lt;li&gt;Anshul Kundaje&lt;sup&gt;1&lt;/sup&gt;* (&lt;a class="link" href="mailto:akundaje@stanford.edu" &gt;akundaje@stanford.edu&lt;/a&gt;): Conceptualization, Project management, Mentoring, Funding, Blog post editing&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;&lt;sup&gt;1&lt;/sup&gt;Stanford University, *Correspondence.&lt;/em&gt;&lt;/p&gt;

 &lt;/div&gt;
&lt;/aside&gt;


 &lt;blockquote&gt;
 &lt;p&gt;This is the first post in a series on &lt;strong&gt;ENCODE GRAMMAR&lt;/strong&gt;. The series will cover:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;ENCODE GRAMMAR: The ENCODE deep learning model resource for decoding the DNA sequence logic of genomic regulatory elements (this post)&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;Accessing and using the ENCODE GRAMMAR collection: A quickstart guide&lt;/li&gt;
&lt;li&gt;Interpreting regulatory DNA with deep learning models&lt;/li&gt;
&lt;li&gt;The transcription factor binding GRAMMAR resource&lt;/li&gt;
&lt;li&gt;The chromatin accessibility GRAMMAR resource&lt;/li&gt;
&lt;li&gt;Predicting the effects of noncoding genetic variants&lt;/li&gt;
&lt;li&gt;MotifCompendium - a unified lexicon of regulatory sequence motifs&lt;/li&gt;
&lt;li&gt;Contrasting regulatory sequence codes across assays and cell types&lt;/li&gt;
&lt;li&gt;Building a production-scale model atlas in an academic setting&lt;/li&gt;
&lt;/ol&gt;

 &lt;/blockquote&gt;
&lt;hr&gt;
&lt;h2 id="the-genome-encodes-a-regulatory-control-system"&gt;The genome encodes a regulatory control system
&lt;/h2&gt;&lt;p&gt;The &lt;strong&gt;human genome&lt;/strong&gt; is the complete set of instructions encoded in DNA and carried by nearly every cell in the body. It contains about 3.2 billion DNA base pairs, built from four nucleotide bases represented by the letters A, C, G, and T. Although human genomes are overwhelmingly similar, their DNA sequences differ at millions of positions across individuals. These differences, known as &lt;strong&gt;genetic variants&lt;/strong&gt;, contribute to human diversity and can influence traits and disease risk.&lt;/p&gt;
&lt;p&gt;The human body contains hundreds of distinct cell types—including neurons, liver cells, and immune cells—with very different morphology and functions. Yet nearly all cells within an individual contain essentially the same genome. How can the same DNA sequence produce such remarkable cellular diversity?&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Genes&lt;/strong&gt; are segments of DNA that contain instructions for producing &lt;strong&gt;RNA&lt;/strong&gt; molecules, some of which are translated into proteins. The transcription of genes into RNA is tightly regulated: different genes are expressed at different levels, in different cell types, and at different times. These distinct patterns of gene expression allow cells with the same genome to acquire different identities and perform specialized functions. Precise regulation of gene expression is therefore essential for development, normal cellular function, and responses to the environment.&lt;/p&gt;
&lt;p&gt;Much of this control is encoded in &lt;strong&gt;regulatory elements&lt;/strong&gt;—regions of DNA that help determine when, where, and how strongly genes are expressed. &lt;strong&gt;Promoters&lt;/strong&gt; are regulatory elements found at or near the sites where gene transcription begins and help recruit the molecular machinery that produces RNA. Other regulatory elements, such as distal &lt;strong&gt;enhancers&lt;/strong&gt;, can act over hundreds of kilobases, boosting transcription of target gene promoters they contact through three-dimensional folding of the genome.&lt;/p&gt;
&lt;p&gt;DNA is packaged with proteins into &lt;strong&gt;chromatin&lt;/strong&gt;, which influences how accessible different regions of the genome are to the cellular machinery. Regulatory proteins called &lt;strong&gt;transcription factors&lt;/strong&gt; (TFs), recognize and bind specific short DNA sequence patterns, or &lt;strong&gt;motifs&lt;/strong&gt;, within regulatory elements and help increase or decrease gene expression by recruiting or blocking the machinery that carries out transcription. Chromatin accessibility and TF binding influence one another: exposed DNA is easier for regulatory proteins to reach, while some proteins can also open or reorganize chromatin.&lt;/p&gt;
&lt;p&gt;Different cell types express different combinations of TFs and maintain different chromatin states across the genome. As a result, different cell types engage different repertoires of regulatory elements producing distinct patterns of gene expression, despite containing essentially the same genomic DNA sequence. Genetic variants within regulatory elements can alter transcription-factor binding or other regulatory activity, potentially changing gene expression in particular cellular contexts. Disruption of this regulatory system can interfere with development and cellular function and contribute to disease.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/intro.png" class="image-link" data-pswp-width="10166" data-pswp-height="5185"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/intro.png" width="600px" height="306"loading="lazy"
			alt="Figure: Introduction to genome regulation"
			title="How the same genome gives rise to diverse cell types and gene-expression programs: (1) Nearly all cells in an individual contain essentially the same genome, yet cell types such as neurons, hepatocytes, and immune cells have distinct structures and functions. (2) The human genome consists of approximately 3.2 billion DNA base pairs, encoded by the nucleotide bases A, C, G, and T and packaged with proteins into chromatin. Human genomes differ at millions of positions. The example illustrates a single-nucleotide genetic variant in which one DNA base is replaced by another. (3) Regulatory elements help determine when, where, and how strongly genes are expressed. Enhancers can increase transcription, promoters mark regions where transcription is initiated, and genes are transcribed into RNA. (4) Regulatory activity depends on both DNA sequence and chromatin state. In accessible chromatin, transcription factors (TFs) bind specific sequence motifs within regulatory elements and regulate RNA polymerase II which transcribes DNA into RNA. In less accessible chromatin, densely positioned nucleosomes can restrict TF and polymerase access, reducing or preventing transcription. The motif examples illustrate that different transcription factors recognize different DNA sequence patterns. (5) Because cell types express different TF combinations and maintain different chromatin states, the same regulatory elements and genes can have different activities across cellular contexts. The bar plots illustrate distinct expression levels of three genes in neurons, liver cells, and immune cells. Together, these mechanisms allow the same genome to generate diverse, cell-type-specific patterns of gene expression." data-title-escaped="How the same genome gives rise to diverse cell types and gene-expression programs: (1) Nearly all cells in an individual contain essentially the same genome, yet cell types such as neurons, hepatocytes, and immune cells have distinct structures and functions. (2) The human genome consists of approximately 3.2 billion DNA base pairs, encoded by the nucleotide bases A, C, G, and T and packaged with proteins into chromatin. Human genomes differ at millions of positions. The example illustrates a single-nucleotide genetic variant in which one DNA base is replaced by another. (3) Regulatory elements help determine when, where, and how strongly genes are expressed. Enhancers can increase transcription, promoters mark regions where transcription is initiated, and genes are transcribed into RNA. (4) Regulatory activity depends on both DNA sequence and chromatin state. In accessible chromatin, transcription factors (TFs) bind specific sequence motifs within regulatory elements and regulate RNA polymerase II which transcribes DNA into RNA. In less accessible chromatin, densely positioned nucleosomes can restrict TF and polymerase access, reducing or preventing transcription. The motif examples illustrate that different transcription factors recognize different DNA sequence patterns. (5) Because cell types express different TF combinations and maintain different chromatin states, the same regulatory elements and genes can have different activities across cellular contexts. The bar plots illustrate distinct expression levels of three genes in neurons, liver cells, and immune cells. Together, these mechanisms allow the same genome to generate diverse, cell-type-specific patterns of gene expression."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;How does the same genome give rise to diverse cell types and gene-expression programs?.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;To understand how the genome encodes gene regulation, we therefore need to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;map the genomic locations of candidate regulatory elements;&lt;/li&gt;
&lt;li&gt;measure their biochemical activity (e.g. transcription-factor binding and chromatin state), and associated gene expression across cell types and conditions;&lt;/li&gt;
&lt;li&gt;determine which DNA bases within regulatory elements are important and how their combinations and arrangements control different types of biochemical activity in different cell types; and&lt;/li&gt;
&lt;li&gt;determine how genetic variants alter biochemical activity in different cellular contexts.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="encode-an-encyclopedia-of-dna-elements-across-thousands-of-cell-types"&gt;ENCODE: An Encyclopedia of DNA Elements across thousands of cell types
&lt;/h2&gt;&lt;p&gt;Over the past two decades, the &lt;a class="link" href="https://www.encodeproject.org/" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;Encyclopedia of DNA Elements&lt;/strong&gt; (&lt;strong&gt;ENCODE&lt;/strong&gt;) Consortium&lt;/a&gt; has made major progress toward the first two goals. Using a broad range of genome-wide functional genomics experiments, ENCODE has mapped millions of candidate regulatory elements in the human and mouse genomes by characterizing their biochemical activity across diverse cell types, tissues, developmental stages, and conditions.&lt;/p&gt;
&lt;p&gt;These experiments measure complementary layers of gene regulation, including where TFs bind DNA, which regions of chromatin are accessible, where transcription begins, and which genes are expressed. Together, they provide detailed maps of where regulatory activity occurs and how it differs across cellular contexts. We briefly describe some of the key experimental assays below.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Transcription factor ChIP-seq&lt;/strong&gt; (transcription factor chromatin immunoprecipitation followed by sequencing) maps where a particular TF binds the genome in a specific cell type. Cells are treated so that proteins remain attached to the DNA they occupy, the DNA is fragmented, and an antibody is used to isolate fragments bound by the TF of interest. These fragments are sequenced on a high-throughput sequencer and mapped to the genome to identify their likely locations, which produces concentrations of reads, or &lt;strong&gt;peaks&lt;/strong&gt;, at genomic regions enriched for binding by that TF. TF ChIP-seq peak regions are often statistically enriched for recurring short DNA sequence patterns, called motifs, that typically mediate the binding of the TF to DNA. However, each TF ChIP–seq experiment profiles only one TF in one cellular context. Systematically measuring the binding of the roughly 2,000 human TFs across all cell types and conditions would therefore be prohibitively laborious and expensive.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/TFChIP.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/TFChIP.gif" width="600px" height="247"loading="lazy"
			alt="Figure: TF ChIP-seq"
			title="A transcription factor (TF) ChIP-seq experiment profiles genomic binding sites of a TF: (1) A TF binds specific regulatory DNA elements across the genome. (2) The DNA is randomly fragmented, (3) an antibody specific to the TF binds the resulting TF–DNA complexes. (3) These complexes are selectively isolated, (4) the associated DNA fragments are purified and sequenced, and (5) the sequencing reads are mapped back to the genome. (6) Regions where many reads accumulate appear as peaks centered around sites occupied by the transcription factor." data-title-escaped="A transcription factor (TF) ChIP-seq experiment profiles genomic binding sites of a TF: (1) A TF binds specific regulatory DNA elements across the genome. (2) The DNA is randomly fragmented, (3) an antibody specific to the TF binds the resulting TF–DNA complexes. (3) These complexes are selectively isolated, (4) the associated DNA fragments are purified and sequenced, and (5) the sequencing reads are mapped back to the genome. (6) Regions where many reads accumulate appear as peaks centered around sites occupied by the transcription factor."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Transcription factor ChIP-seq experiments.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DNase-seq&lt;/strong&gt; and &lt;strong&gt;ATAC-seq&lt;/strong&gt; experiments partly address this limitation by providing a genome-wide map of regions of accessible chromatin which are often regulatory elements occupied by combinations of TFs and other regulatory proteins. So a single accessibility experiment can highlight regulatory elements bound by many factors in a given cellular context, although it does not directly identify which proteins are bound. DNase–seq uses the enzyme DNase I to cut exposed DNA, whereas ATAC–seq uses the Tn5 transposase to insert sequencing adapters into accessible DNA. The resulting DNA fragments are sequenced and mapped to the genome. Regions containing many mapped fragments appear as peaks of chromatin accessibility. DNase I and Tn5 also have preferences for particular DNA sequences, so the observed signal profiles reflect both genuine chromatin accessibility and assay-specific sequence bias. Separating these components is especially important when interpreting the signal at single-base resolution.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/ChromatinAccessibility.gif" class="image-link" data-pswp-width="1920" data-pswp-height="1080"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/ChromatinAccessibility.gif" width="600px" height="337"loading="lazy"
			alt="Figure: DNase-seq, ATAC-seq"
			title="DNase-seq and ATAC-seq experiments profile regions of accessible chromatin: (1) DNA wrapped around histone proteins is relatively inaccessible, whereas (2) exposed DNA can be cut by DNase-I in DNase-seq or Tn5 transposase in ATAC-seq. These enzymes preferentially act on accessible DNA, and the resulting fragments are isolated, (3) sequenced, and mapped back to the genome. (4) Regions where many fragments accumulate appear as peaks of chromatin accessibility." data-title-escaped="DNase-seq and ATAC-seq experiments profile regions of accessible chromatin: (1) DNA wrapped around histone proteins is relatively inaccessible, whereas (2) exposed DNA can be cut by DNase-I in DNase-seq or Tn5 transposase in ATAC-seq. These enzymes preferentially act on accessible DNA, and the resulting fragments are isolated, (3) sequenced, and mapped back to the genome. (4) Regions where many fragments accumulate appear as peaks of chromatin accessibility."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;DNase-seq and ATAC-seq chromatin accessibility experiments.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;However, TF binding and chromatin accessibility do not necessarily lead to productive downstream regulatory effects such as transcription. Additional assays are therefore needed to measure where transcription initiates and whether candidate DNA sequences can directly drive regulatory activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;PRO-cap&lt;/strong&gt; experiments map the precise genomic positions at which transcription initiates across the genome. It enriches for and sequences the capped 5′ ends of newly synthesized RNAs, producing base-resolution, strand-specific maps of active transcription initiation. PRO-cap can identify initiation at both gene promoters and transcribed regulatory elements, while the number of reads beginning at a site provides a measure of its relative initiation activity.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Massively parallel reporter assays (MPRAs)&lt;/strong&gt; directly test the regulatory potential of thousands of DNA sequences in parallel. Each candidate sequence is placed alongside a reporter gene in a synthetic construct, typically together with a sequence barcode that identifies it. After the constructs are introduced into cells, regulatory activity is commonly measured by comparing the abundance of each barcode in reporter RNA with its abundance in the input DNA library. Sequences that produce more reporter RNA have greater regulatory activity in that assay and cellular context. Because the sequences are generally tested outside their native genomic locations, MPRAs measure regulatory potential in the reporter system rather than fully reproducing their endogenous functions.&lt;/p&gt;
&lt;p&gt;Together, these assays measure complementary layers of gene regulation. TF ChIP-seq identifies where individual TFs bind; DNase-seq and ATAC-seq reveal accessible regulatory DNA; PRO-cap pinpoints sites of active transcription initiation; and MPRAs directly test whether particular DNA sequences can drive regulatory activity.&lt;/p&gt;
&lt;p&gt;ENCODE has completed and released more than 16,000 genome-wide assays across thousands of biological samples, including cell lines, primary cells, tissues, differentiated cells, and experimentally perturbed samples from humans and mice. These data are processed using standardized pipelines and made publicly available through the &lt;a class="link" href="https://encodeproject.org" target="_blank" rel="noopener"
 &gt;ENCODE portal&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;By integrating evidence from these and many other assays, ENCODE has mapped more than 5 million accessible chromatin elements in the human genome. Approximately 2.4 million of these are classified as **&lt;a class="link" href="https://screen.wenglab.org/" target="_blank" rel="noopener"
 &gt;candidate cis-regulatory elements (cCREs)**&lt;/a&gt; because they are also supported by additional biochemical signatures of regulatory activity. The ENCODE consortium recently released a preprint describing the entire compendium developed over two decades including many new datasets and derived analysis products from the fourth and final phase of the project &lt;a class="link" href="https://www.biorxiv.org/content/10.64898/2026.07.06.731365v1" target="_blank" rel="noopener"
 &gt;ENCODE 4&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/ENCODE_cube.png" class="image-link" data-pswp-width="1440" data-pswp-height="810"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/ENCODE_cube.png" width="1440" height="810"loading="lazy"
			alt="Figure: ENCODE cube"
			title="The ENCODE data cube: ENCODE comprises 1000s of datasets spanning 3 dimensions. (1) Biochemical assays measure diverse layers of genome function, including TF binding, chromatin accessibility, histone modifications, transcription initiation and nascent transcription, RNA expression, DNA methylation, and 3D long-range chromatin interactions. (2) Genomic coordinates place each measurement across approximately 3 billion positions in the human genome, producing signal tracks that can be compared at the same loci. (3) Biological contexts include cell lines, primary cells, tissues, developmental stages, and experimentally perturbed samples. The grid is schematic and does not imply that every possible combination has been experimentally profiled." data-title-escaped="The ENCODE data cube: ENCODE comprises 1000s of datasets spanning 3 dimensions. (1) Biochemical assays measure diverse layers of genome function, including TF binding, chromatin accessibility, histone modifications, transcription initiation and nascent transcription, RNA expression, DNA methylation, and 3D long-range chromatin interactions. (2) Genomic coordinates place each measurement across approximately 3 billion positions in the human genome, producing signal tracks that can be compared at the same loci. (3) Biological contexts include cell lines, primary cells, tissues, developmental stages, and experimentally perturbed samples. The grid is schematic and does not imply that every possible combination has been experimentally profiled." class="gallery-image" data-flex-grow="177" data-flex-basis="426px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;The ENCODE data cube.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Together, these experiments address two foundational goals: mapping regulatory elements throughout the genome and characterizing their biochemical activity and properties across cellular contexts. Yet these maps do not explain how DNA sequence mediates the diverse types of biochemical activity across the genome and their cell-type specificity. Important questions remain:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Which individual DNA bases and sequence motifs within a regulatory element influence different types of biochemical activity (e.g. TF binding, accessibility) in a specific cell type?&lt;/li&gt;
&lt;li&gt;How do combinations and arrangements of these patterns determine biochemical activity?&lt;/li&gt;
&lt;li&gt;How does the same regulatory element behave differently across cell types?&lt;/li&gt;
&lt;li&gt;How might a genetic variant alter biochemical activity in a particular cellular context?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Answering these questions requires moving beyond mapping regulatory elements to decoding the sequence rules that govern their activity.&lt;/p&gt;
&lt;h2 id="the-bpnet-family-of-deep-learning-models-from-regulatory-maps-to-predictive-sequence-rules"&gt;The BPNet family of deep learning models: From regulatory maps to predictive sequence rules
&lt;/h2&gt;&lt;p&gt;We developed the &lt;strong&gt;BPNet family&lt;/strong&gt; of deep learning models to address these questions. These neural networks use stacks of dilated residual convolutional layers to learn sequence features including TF motifs and their higher-order combinations and arrangements (called regulatory syntax) that can predict the biochemical activity measured by each experiment at every base, using up to approximately 2 kilobases of local DNA sequence context. Rather than simply classifying a region as active or inactive, BPNet models predict both the total amount of activity and the shape of the experimental signal at &lt;strong&gt;base-pair resolution&lt;/strong&gt;. The deliberate choice of restricting the models to only use local-context makes them computationally efficient and amenable to robust sequence-level interpretation. Despite their compact architecture and restricted sequence context, these models are &lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;quite competitive&lt;/a&gt; with substantially larger models that use much longer genomic sequences.&lt;/p&gt;
&lt;p&gt;The ENCODE GRAMMAR model resource contains four related model families:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/bpnet" target="_blank" rel="noopener"
 &gt;BPNet&lt;/a&gt;:&lt;/strong&gt; &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=BPNet&amp;amp;type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;status=released&amp;amp;assay_term_name=ChIP-seq" target="_blank" rel="noopener"
 &gt;2,339 TF ChIP-seq model sets&lt;/a&gt; spanning 788 transcription-factor targets.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/chrombpnet" target="_blank" rel="noopener"
 &gt;ChromBPNet&lt;/a&gt;:&lt;/strong&gt; &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ChromBPNet&amp;amp;type=Annotation&amp;amp;annotation_type=ChromBPNet-model&amp;amp;organism.scientific_name=Homo&amp;#43;sapiens&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;1,512 DNase-seq and ATAC-seq model sets&lt;/a&gt; across 408 biosamples.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/procapnet" target="_blank" rel="noopener"
 &gt;ProCapNet&lt;/a&gt;:&lt;/strong&gt; &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ProCapNet&amp;amp;type=Annotation" target="_blank" rel="noopener"
 &gt;6 PRO-cap model sets&lt;/a&gt; that predict transcription-initiation profiles.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/reporternet" target="_blank" rel="noopener"
 &gt;ReporterNet&lt;/a&gt;:&lt;/strong&gt; &lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ReporterNet&amp;amp;type=Annotation&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;8 model sets&lt;/a&gt; trained on high-throughput reporter assays. Unlike the above three models, ReporterNet makes predictions at the resolution of the candidate sequences tested in the experiments.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Together, these comprise &lt;a class="link" href="https://www.encodeproject.org/search/?type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;annotation_type=ChromBPNet-model&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;3,865 experiment-specific model sets&lt;/strong&gt;&lt;/a&gt;, each trained and evaluated using five-fold cross-validation. They span diverse cell lines, primary cells, and tissues represented in ENCODE.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_architecture.png" class="image-link" data-pswp-width="4000" data-pswp-height="1650"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_architecture.png" width="600px" height="247"loading="lazy"
			alt="Figure: ChromBPNet model architecture"
			title="Neural network architecture schematic of the BPNet model family: BPNet uses ~2 kb of local DNA sequence as input and applies convolutional and dilated residual layers to learn predictive sequence features and their spatial organization. The model jointly predicts the base-resolution shape of the regulatory profile across a 1 kb region and the total experimental signal within that region." data-title-escaped="Neural network architecture schematic of the BPNet model family: BPNet uses ~2 kb of local DNA sequence as input and applies convolutional and dilated residual layers to learn predictive sequence features and their spatial organization. The model jointly predicts the base-resolution shape of the regulatory profile across a 1 kb region and the total experimental signal within that region."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Schematic of the neural network architecture of the BPNet model family.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The trained models are only one component of the resource. We also developed a suite of interpretation methods, described below, to interrogate each model and identify the cell-context-specific DNA sequence features that drive its predictions within biochemically active regulatory elements. For every experiment, ENCODE GRAMMAR provides several derived products:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Predicted regulatory profiles&lt;/strong&gt; at base-pair resolution which often reveal de-noised signal profiles compared to the sparse, noisy measured profiles especially from TF ChIP-seq experiments.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bias-corrected regulatory profiles&lt;/strong&gt; for chromatin-accessibility assays, eliminating distortions in the profiles due to sequence biases of the Tn5 and DNase I enzymes used in ATAC–seq and DNase–seq, respectively.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Sequence-contribution maps&lt;/strong&gt; estimating how much each DNA base contributes to a model’s prediction for individual regulatory sequences.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;De novo predictive sequence motifs&lt;/strong&gt; which are derived from recurrent patterns of high contribution scores with similar sequences across biochemically active regulatory sequences (e.g. peaks)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;ENCODE Motif Compendium&lt;/strong&gt; which is a unified lexicon of non-redundant sequence motifs derived from all ENCODE GRAMMAR models&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Predictive genomic motif instances&lt;/strong&gt; which map high-contribution sequence patterns in all biochemically active regulatory sequences to the unified motif lexicon.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Genome-browser tracks&lt;/strong&gt; that allow predictions, contribution scores, motifs, and motif instances to be explored together at any genomic locus.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Together, this collection of models and derived annotations transforms thousands of ENCODE experiments into a practical, sequence-resolved atlas of gene-regulatory activity and its underlying predictive DNA features.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/main.png" class="image-link" data-pswp-width="4000" data-pswp-height="2250"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/main.png" width="4000" height="2250"loading="lazy"
			alt="Figure: ENCODE GRAMMAR"
			title="ENCODE GRAMMAR resource that transforms the extensive ENCODE compendium of genome-wide biochemical profiling experiments into predictive models and interpretable regulatory sequence annotations: (1) ENCODE experiments measure complementary layers of gene regulation, including TF binding by TF ChIP–seq, chromatin accessibility by DNase-seq and ATAC-seq, transcription initiation by PRO-cap, and sequence-driven regulatory activity by MPRAs. (2) Deep learning models from the BPNet family (BPNet, ChromBPNet, ProCapNet, and ReporterNet) are trained separately for each experiment and cellular context to predict the corresponding biochemical signal directly from local DNA sequence. (3) Product resources released for each experiment include the trained models; predicted, base-resolution biochemical profiles; sequence-contribution maps identifying bases that drive model predictions; recurring predictive sequence motifs; genomic motif instances; and predicted effects of genetic variants obtained by comparing reference and alternate allele sequences." data-title-escaped="ENCODE GRAMMAR resource that transforms the extensive ENCODE compendium of genome-wide biochemical profiling experiments into predictive models and interpretable regulatory sequence annotations: (1) ENCODE experiments measure complementary layers of gene regulation, including TF binding by TF ChIP–seq, chromatin accessibility by DNase-seq and ATAC-seq, transcription initiation by PRO-cap, and sequence-driven regulatory activity by MPRAs. (2) Deep learning models from the BPNet family (BPNet, ChromBPNet, ProCapNet, and ReporterNet) are trained separately for each experiment and cellular context to predict the corresponding biochemical signal directly from local DNA sequence. (3) Product resources released for each experiment include the trained models; predicted, base-resolution biochemical profiles; sequence-contribution maps identifying bases that drive model predictions; recurring predictive sequence motifs; genomic motif instances; and predicted effects of genetic variants obtained by comparing reference and alternate allele sequences." class="gallery-image" data-flex-grow="177" data-flex-basis="426px"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;The ENCODE GRAMMAR resource transforms the extensive ENCODE compendium of genome-wide biochemical profiling experiments into predictive models and interpretable regulatory sequence annotations.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="overview-of-the-workflow-for-generating-the-encode-grammar-resource"&gt;Overview of the workflow for generating the ENCODE GRAMMAR resource
&lt;/h2&gt;&lt;p&gt;We describe the main steps of our workflow below. More detailed descriptions of the model architecture, evaluations and applications are available in the &lt;a class="link" href="https://doi.org/10.1038/s41588-021-00782-6" target="_blank" rel="noopener"
 &gt;BPNet&lt;/a&gt;, &lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;ChromBPNet&lt;/a&gt;, and &lt;a class="link" href="https://doi.org/10.1101/2024.05.28.596138" target="_blank" rel="noopener"
 &gt;ProCapNet&lt;/a&gt; manuscripts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(1) Train a set of sequence-to-profile models for each experiment:&lt;/strong&gt; For each experiment, we train a deep learning model to predict the measured signal profile from the local DNA sequence surrounding biochemically active regions, together with background regions matched for overall sequence composition. The model is trained on a subset of chromosomes and evaluated on held-out chromosomes containing sequences never seen during training. Strong performance indicates that it has learned generalizable sequence features, including TF motifs and their combinations, spacing, and arrangement. We typically train at least five models per experiment using different chromosome splits, allowing us to estimate performance variability and obtain more stable predictions and interpretations by averaging across models. The resulting ensemble constitutes the experiment’s model set.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig1.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig1.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Train a model"
			title="A deep learning model is trained to predict experimentally measured biochemical profiles from local DNA sequence context." data-title-escaped="A deep learning model is trained to predict experimentally measured biochemical profiles from local DNA sequence context."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Train a sequence-to-profile deep learning model.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(2) Predict sequence and variant effects:&lt;/strong&gt; Once trained, a model can predict biochemical activity profiles for previously unseen DNA sequences. Because each model is trained on a single assay in a specific cellular context and receives only DNA sequence as input, it learns the sequence-to-activity relationship particular to that experiment. Its predictions should therefore not be extrapolated directly to other assays or cell types. Generalization is strongest for sequences whose regulatory syntax resembles that represented in the training data, although the models can often tolerate modest departures from this distribution. Genetic variants provide an important example. Although the models are trained only on reference-genome sequences and receive no explicit variant-effect labels, they can often predict the molecular effects of genetic variants quite effectively (&lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;See Fig. 6 in the ChromBPNet paper&lt;/a&gt;). We estimate a variant’s effect by comparing predictions for the reference and alternate sequences. Predictions are typically averaged across all models in the corresponding model set to improve robustness and stability, while variation among models provides an empirical estimate of uncertainty. The resulting change in signal quantifies the variant’s predicted effect on the specific biochemical activity measured by that experiment, not on downstream phenotypes such as disease risk.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig3.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig3.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Predict mutations"
			title="Predict the biochemical effect of unseen sequences and genetic variants: A trained model can predict the biochemical activity profile of an unseen DNA sequence. To estimate the effect of a genetic variant, predictions for the reference and alternate sequences are compared. The resulting change in signal quantifies the variant’s predicted effect on the molecular activity measured by the experiment." data-title-escaped="Predict the biochemical effect of unseen sequences and genetic variants: A trained model can predict the biochemical activity profile of an unseen DNA sequence. To estimate the effect of a genetic variant, predictions for the reference and alternate sequences are compared. The resulting change in signal quantifies the variant’s predicted effect on the molecular activity measured by the experiment."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Predict the biochemical effect of unseen sequences and genetic variants.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(3) Separate biological signal from assay-specific sequence bias:&lt;/strong&gt; ATAC-seq and DNase-seq profiles reflect both genuine chromatin accessibility distorted by the sequence preferences of the DNase I and Tn5 enzymes. ChromBPNet explicitly models these components. A sequence bias model learns the enzyme-driven signal, while the main model learns the remaining sequence-dependent accessibility signal. This &lt;strong&gt;bias-factorized prediction&lt;/strong&gt; provides a cleaner estimate of the underlying regulatory profile, particularly at base-pair resolution.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig2.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig2.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Remove bias"
			title="Separate biological signal from assay-specific sequence bias: For DNase-seq and ATAC-seq, ChromBPNet first learns sequence-dependent enzyme bias from the data and subtracts it out of the measured profiles to make sequence based bias-corrected predictions that more closely represents the underlying biological accessibility profile." data-title-escaped="Separate biological signal from assay-specific sequence bias: For DNase-seq and ATAC-seq, ChromBPNet first learns sequence-dependent enzyme bias from the data and subtracts it out of the measured profiles to make sequence based bias-corrected predictions that more closely represents the underlying biological accessibility profile."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Separate biological signal from assay-specific sequence bias.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(4) Estimate quantitative base-resolution sequence contributions that drive predictions:&lt;/strong&gt; How is the model making a particular prediction? One useful way to answer this is to ask which bases in the input sequence are responsible for the predicted activity and by how much. We use the &lt;strong&gt;&lt;a class="link" href="https://proceedings.mlr.press/v70/shrikumar17a.html" target="_blank" rel="noopener"
 &gt;DeepLIFT&lt;/a&gt;/&lt;a class="link" href="https://github.com/kundajelab/shap" target="_blank" rel="noopener"
 &gt;DeepSHAP&lt;/a&gt;&lt;/strong&gt; feature attribution method which efficiently approximates the model’s predictions for any input sequence as an additive sum of its base-resolution contributions. For a selected input sequence and its associated model prediction, DeepLIFT assigns a contribution score to each base such that the scores sum to the difference between the model’s prediction for the input sequence and its prediction for dinucleotide-shuffled reference sequences. Positive scores identify bases that increase the predicted activity relative to this reference, whereas negative scores identify bases that decrease it. For each sequence, we average contribution scores across all models in the corresponding model set to obtain more stable and robust estimates. The resulting &lt;strong&gt;sequence-contribution map&lt;/strong&gt; highlights the individual bases and short sequence patterns that drive the model’s prediction. Many high-contribution patterns correspond to known TF binding sites, while others may reveal previously unrecognized predictive sequence features.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig4.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig4.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Understand the importance sequences for the model"
			title="Estimate the contribution of each base in a query DNA sequences to a model’s prediction of biochemical activity: DeepLIFT/DeepSHAP assigns a quantitative contribution score to each base in the input sequence, indicating how strongly it increases or decreases the predicted activity relative to shuffled reference sequences. The resulting sequence-contribution map highlights the bases and short sequence patterns used by the model to make the specific prediction." data-title-escaped="Estimate the contribution of each base in a query DNA sequences to a model’s prediction of biochemical activity: DeepLIFT/DeepSHAP assigns a quantitative contribution score to each base in the input sequence, indicating how strongly it increases or decreases the predicted activity relative to shuffled reference sequences. The resulting sequence-contribution map highlights the bases and short sequence patterns used by the model to make the specific prediction."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Estimate the contribution of each base in a query DNA sequences to a model’s prediction of biochemical activity.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(5) Discover recurring predictive sequence motifs:&lt;/strong&gt;. A single model can identify thousands of high-contribution sequence instances across biochemically active peaks across the genome. To summarize these recurring patterns, we developed the &lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/tfmodisco" target="_blank" rel="noopener"
 &gt;TF-MoDISco&lt;/a&gt;&lt;/strong&gt; algorithm which samples highly contributing subsequences across peak regions, aligns and clusters similar subsequences and averages clusters into position specific contribution-weighted summary motifs. These motifs provide a compact representation of the sequence features learned by the model. Many can be matched to the known binding preferences of TF, helping identify regulators that may influence the experimental signal. Others often reveal previously unrecognized motifs, context-specific sequence preferences of TFs, or novel composite patterns recognized by complexes of multiple TFs.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig5.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig5.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Aggregate elements into motifs"
			title="Discover recurring predictive sequence motifs: TF-MoDISco samples highly contributing subsequences from many biochemically active genomic regions, aligns and clusters similar examples, and summarizes each cluster as a contribution-weighted motif." data-title-escaped="Discover recurring predictive sequence motifs: TF-MoDISco samples highly contributing subsequences from many biochemically active genomic regions, aligns and clusters similar examples, and summarizes each cluster as a contribution-weighted motif."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Discover recurring predictive sequence motifs.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(6) Map predictive motif instances across the genome:&lt;/strong&gt; Each TF-MoDISco motif summarizes a cluster of similar high-contribution subsequences, but does not by itself identify all predictive instances of that motif in genomic or user-designed sequences with high sensitivity and specificity. This is challenging because motifs can overlap, share similar subpatterns, and compete to explain the same contribution signal. We developed &lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/Fi-NeMo" target="_blank" rel="noopener"
 &gt;FiNeMo&lt;/a&gt;&lt;/strong&gt; to efficiently scan genome-wide sequence-contribution maps with a collection of TF-MoDISco motifs and jointly assign predictive sequence instances to the motifs that best explain them. The result is a quantitative genome-wide map of predictive motif occurrences from each model within biochemically active regulatory elements identified by the corresponding experiment.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig6.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig6.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Identify all genomics motif instances"
			title="Map predictive motif instances: FiNeMo scans sequence-contribution maps with the motifs discovered by TF-MoDISco and jointly identifies the motif instances that best explain the contribution scores. This produces quantitative maps of predictive motif occurrences across genomic or other query sequences" data-title-escaped="Map predictive motif instances: FiNeMo scans sequence-contribution maps with the motifs discovered by TF-MoDISco and jointly identifies the motif instances that best explain the contribution scores. This produces quantitative maps of predictive motif occurrences across genomic or other query sequences"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Map predictive motif instances .&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;(7) Unify motifs across all models into a non-redundant motif lexicon:&lt;/strong&gt; Each model produces its own TF-MoDISco motifs, making it difficult to distinguish patterns that are shared across assays and cell types from those that are context specific. Comparing hundreds of thousands of contribution-weighted motifs is challenging because similarity must account for shifts, reverse complements, partial overlaps, and differences in contribution patterns. We developed &lt;strong&gt;&lt;a class="link" href="https://github.com/kundajelab/motifcompendium" target="_blank" rel="noopener"
 &gt;MotifCompendium&lt;/a&gt;&lt;/strong&gt; to efficiently cluster motifs discovered from all our trained models into a scalable, non-redundant lexicon, link them to known TF motifs, classify distinct motif types, and retain the models and contexts in which each was discovered.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig7.gif" class="image-link" data-pswp-width="1333" data-pswp-height="550"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/BPNet_Fig7.gif" width="600px" height="247"loading="lazy"
			alt="Figure: Combine motifs across models"
			title="Unify motifs across models: MotifCompendium compares and clusters related motifs discovered by models trained across different assays and cell contexts into a unified motif lexicon" data-title-escaped="Unify motifs across models: MotifCompendium compares and clusters related motifs discovered by models trained across different assays and cell contexts into a unified motif lexicon"&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Unify motifs across models.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;h2 id="case-study-decoding-a-myc-enhancer-with-the-encode-grammar-resource"&gt;Case study: Decoding a &lt;em&gt;MYC&lt;/em&gt; enhancer with the ENCODE GRAMMAR resource
&lt;/h2&gt;&lt;p&gt;To illustrate how the ENCODE GRAMMAR resource can be used to uncover the sequence basis of gene regulation, we examine a distal enhancer of the &lt;em&gt;MYC&lt;/em&gt; gene in K562, a leukemia cell line. &lt;em&gt;MYC&lt;/em&gt; encodes a transcription factor that promotes cell growth and proliferation, and dysregulation of &lt;em&gt;MYC&lt;/em&gt; is a common feature of many cancers. We focus on a CRISPRi-validated enhancer of &lt;em&gt;MYC&lt;/em&gt; at [chr8:127,898,412—127,899,647] and analyze its sequence using 15 independently trained models spanning chromatin accessibility and TF binding assays in K562.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/MYC_fig0.png" class="image-link" data-pswp-width="3748" data-pswp-height="588"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/MYC_fig0.png" width="600px" height="94"loading="lazy"
			alt="Figure: MYC enhancer"
			title="A distal MYC enhancer in K562 cells: Overview of the *MYC* locus showing a CRISPRi-validated enhancer located approximately 162 kb from the gene (green marker). The green wedge indicates the region shown at higher resolution below, where observed DNase-seq and ATAC-seq profiles reveal strong chromatin accessibility. Track labels report the displayed signal ranges." data-title-escaped="A distal MYC enhancer in K562 cells: Overview of the *MYC* locus showing a CRISPRi-validated enhancer located approximately 162 kb from the gene (green marker). The green wedge indicates the region shown at higher resolution below, where observed DNase-seq and ATAC-seq profiles reveal strong chromatin accessibility. Track labels report the displayed signal ranges."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;A distal enhancer of the MYC gene in K562 leukemia cells.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We first examine chromatin accessibility measured by DNase-seq and ATAC-seq. A separate ChromBPNet model was trained for each experiment and then used to predict its corresponding experimental profile. Both models closely recapitulate the broad shape and fine-scale structure of the observed signal at the enhancer, demonstrating that local DNA sequence contains substantial information about its accessibility in K562 cells.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/MYC_fig1.png" class="image-link" data-pswp-width="2053" data-pswp-height="217"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/MYC_fig1.png" width="600px" height="63"loading="lazy"
			alt="Figure: MYC - Observed and predicted profile"
			title="Observed and ChromBPNet-predicted chromatin-accessibility profiles at the MYC enhancer: Experimentally observed and model-predicted DNase-seq and ATAC-seq profiles across the enhancer in K562 cells. The independently trained ChromBPNet models recapitulate the broad structure and many fine-scale features of their corresponding experimental signals. Despite measuring the same underlying property, the DNase-seq and ATAC-seq profiles differ substantially in shape, reflecting assay-specific effects such as the distinct sequence preferences of DNase I and Tn5. Track labels indicate the displayed signal ranges." data-title-escaped="Observed and ChromBPNet-predicted chromatin-accessibility profiles at the MYC enhancer: Experimentally observed and model-predicted DNase-seq and ATAC-seq profiles across the enhancer in K562 cells. The independently trained ChromBPNet models recapitulate the broad structure and many fine-scale features of their corresponding experimental signals. Despite measuring the same underlying property, the DNase-seq and ATAC-seq profiles differ substantially in shape, reflecting assay-specific effects such as the distinct sequence preferences of DNase I and Tn5. Track labels indicate the displayed signal ranges."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Observed and ChromBPNet-predicted chromatin-accessibility profiles at the *MYC* enhancer.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Closer inspection reveals that the raw observed and predicted DNase-seq and ATAC-seq profiles differ substantially, even though both assays measure chromatin accessibility. Much of this discrepancy arises because DNase I and Tn5 have distinct sequence preferences. ChromBPNet models the regulatory signal and assay-specific enzyme bias separately, producing &lt;strong&gt;bias-corrected accessibility profiles&lt;/strong&gt; that more closely approximates the underlying biological signal. Although the DNase-seq and ATAC-seq models were trained independently, their bias-corrected predictions converge on a much more similar accessibility profile at the enhancer, helping reconcile the two assays.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/MYC_fig2.png" class="image-link" data-pswp-width="2053" data-pswp-height="107"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/MYC_fig2.png" width="600px" height="31"loading="lazy"
			alt="Figure: MYC - Bias-corrected profile"
			title="Bias-corrected ChromBPNet accessibility profiles at the *MYC* enhancer: ChromBPNet separates assay-specific enzyme bias from the predicted regulatory signal in DNase-seq and ATAC-seq. Although the raw profiles differ substantially, the independently derived bias-corrected predictions converge on a similar accessibility profile across the enhancer. Track labels indicate the displayed signal ranges." data-title-escaped="Bias-corrected ChromBPNet accessibility profiles at the *MYC* enhancer: ChromBPNet separates assay-specific enzyme bias from the predicted regulatory signal in DNase-seq and ATAC-seq. Although the raw profiles differ substantially, the independently derived bias-corrected predictions converge on a similar accessibility profile across the enhancer. Track labels indicate the displayed signal ranges."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Bias-corrected ChromBPNet accessibility profiles at the *MYC* enhancer.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We next interrogate the ChromBPNet models using DeepLIFT to identify the DNA bases that influence their bias-corrected accessibility predictions. The sequence-contribution maps highlight multiple predictive motif instances associated with TFs active in K562, including GATA, AP-1, SP, ETV/ETS, and CEBP family proteins. The contribution maps are also highly reproducible across the DNase-seq and ATAC-seq models. These annotations move beyond identifying the enhancer as accessible and help nominate the specific sequence features that may help establish and maintain its accessibility.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/MYC_fig3.png" class="image-link" data-pswp-width="4000" data-pswp-height="766"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/MYC_fig3.png" width="600px" height="114"loading="lazy"
			alt="Figure: MYC - Sequence contribution scores"
			title="Highly concordant sequence-contribution maps reveal shared regulatory features at the *MYC* enhancer: ChromBPNet contribution scores from independently trained DNase-seq and ATAC-seq models show strong concordance and highlight many of the same predictive sequence features. Zoomed views identify shared motif instances associated with GATA, SP, AP-1, ETV/ETS, and CEBP family TFs. Shaded regions indicate selected high-contribution sites, and red bars mark the corresponding annotated motif instances." data-title-escaped="Highly concordant sequence-contribution maps reveal shared regulatory features at the *MYC* enhancer: ChromBPNet contribution scores from independently trained DNase-seq and ATAC-seq models show strong concordance and highlight many of the same predictive sequence features. Zoomed views identify shared motif instances associated with GATA, SP, AP-1, ETV/ETS, and CEBP family TFs. Shaded regions indicate selected high-contribution sites, and red bars mark the corresponding annotated motif instances."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;Highly concordant sequence-contribution maps reveal shared regulatory features at the *MYC* enhancer.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Finally, we compare the accessibility-derived annotations from ChromBPNet with sequence-contribution maps from BPNet models trained on TF ChIP–seq experiments. For each motif class identified by ChromBPNet, the corresponding TF-specific BPNet model assigns high contribution to the same genomic instances. For example, GATA sites are highlighted by the GATA2 model, AP-1 sites by the JUND (an AP-1 TF family member) model, and ETV/ETS sites by the GABPB1 (an ETS TF family member) model. Collectively, the TF-binding models account for most of the predictive motif instances identified by the accessibility models. This cross-assay agreement links individual sequence features to both chromatin accessibility and binding by specific TFs, providing a more detailed view of the enhancer’s regulatory sequence logic.&lt;/p&gt;
&lt;figure&gt;&lt;a href="https://genomicsxai.github.io/blogs/2026-012/MYC_fig4.png" class="image-link" data-pswp-width="4000" data-pswp-height="903"&gt;
		&lt;img src="https://genomicsxai.github.io/blogs/2026-012/MYC_fig4.png" width="600px" height="135"loading="lazy"
			alt="Figure: MYC - Chromatin accessibility vs. transcription factor binding"
			title="TF ChIP–seq BPNet models link ChromBPNet motif instances to specific TFs at the *MYC* enhancer: Sequence-contribution maps from independently trained GATA2, SP1, CEBPB, JUND, and GABPB1 models highlight the corresponding GATA, SP, CEBP, AP-1, and ETV/ETS motif instances identified by the DNase-seq and ATAC-seq ChromBPNet models. Collectively, these TF-specific models account for most of the predictive motif instances highlighted by ChromBPNet, providing cross-assay support for the inferred regulatory sequence architecture. Red bars mark the annotated motif instances." data-title-escaped="TF ChIP–seq BPNet models link ChromBPNet motif instances to specific TFs at the *MYC* enhancer: Sequence-contribution maps from independently trained GATA2, SP1, CEBPB, JUND, and GABPB1 models highlight the corresponding GATA, SP, CEBP, AP-1, and ETV/ETS motif instances identified by the DNase-seq and ATAC-seq ChromBPNet models. Collectively, these TF-specific models account for most of the predictive motif instances highlighted by ChromBPNet, providing cross-assay support for the inferred regulatory sequence architecture. Red bars mark the annotated motif instances."&gt;
		&lt;/a&gt;&lt;/figure&gt;&lt;p align="center"&gt;&lt;em&gt;TF ChIP–seq BPNet models link ChromBPNet motif instances to specific TFs at the *MYC* enhancer.&lt;br&gt;(Click the figure to see detailed legend)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The example illustrates a central strength of ENCODE GRAMMAR: models trained on complementary ENCODE assays can be integrated at the same locus to connect experimental profiles with the individual DNA bases, motifs, and TFs that may drive different types of biochemical activity.
The interactive browser below allows the experimental data, model predictions, sequence-contribution maps, and motif annotations to be explored together at the MYC enhancer:
&lt;div class="gxai-igv-panel" data-igv-panel="myc" data-igv-src="https://genomicsxai.github.io/blogs/2026-012/myc-igv-panel.json"&gt;&lt;/div&gt;
&lt;link rel="stylesheet" href="https://genomicsxai.github.io/css/igv-panel.css"&gt;
&lt;script src="https://genomicsxai.github.io/js/vendor/igv-3.8.0.min.js" defer&gt;&lt;/script&gt;
&lt;script src="https://genomicsxai.github.io/js/igv-panel.js" defer&gt;&lt;/script&gt;
&lt;/p&gt;
&lt;h2 id="what-can-researchers-do-with-the-encode-grammar-resource"&gt;What can researchers do with the ENCODE GRAMMAR resource?
&lt;/h2&gt;&lt;p&gt;GRAMMAR is designed to support analyses that would otherwise require training and interpreting thousands of models from scratch. Researchers can use it to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;explore base-resolution regulatory predictions and sequence-contribution maps in a genome browser;&lt;/li&gt;
&lt;li&gt;identify candidate motifs, motif combinations and other sequence features that influence different types of biochemical activity of a candidate regulatory element in many cell types;&lt;/li&gt;
&lt;li&gt;compare local and globally predictive sequence features across assays, cell types, and tissues mapped to a common unified lexicon;&lt;/li&gt;
&lt;li&gt;predict the effects of non-coding genetic variants in diverse assay and cell contexts;&lt;/li&gt;
&lt;li&gt;interpret what sequence features noncoding genetic variants may be disrupting to alter context-specific regulatory activity;&lt;/li&gt;
&lt;li&gt;reuse trained models and derived annotations in new computational methods.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These models are most informative when interpreted together with experimental data and appropriate biological context. Their outputs provide testable hypotheses about the sequence determinants of regulatory activity, not substitutes for perturbation experiments or evidence of causal effects on downstream phenotypes.&lt;/p&gt;
&lt;h2 id="accessing-the-encode-grammar-resource"&gt;Accessing the ENCODE GRAMMAR resource
&lt;/h2&gt;&lt;p&gt;All ENCODE data, models, and model-derived sequence annotations are openly available through the &lt;a class="link" href="https://www.encodeproject.org/search/?type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;annotation_type=ChromBPNet-model&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;ENCODE portal&lt;/strong&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=BPNet&amp;amp;type=Annotation&amp;amp;annotation_type=BPNet-model&amp;amp;status=released&amp;amp;assay_term_name=ChIP-seq" target="_blank" rel="noopener"
 &gt;2,339 TF ChIP-seq BPNet model sets&lt;/a&gt; spanning 788 transcription-factor targets.&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ChromBPNet&amp;amp;type=Annotation&amp;amp;annotation_type=ChromBPNet-model&amp;amp;organism.scientific_name=Homo&amp;#43;sapiens&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;1,512 DNase-seq and ATAC-seq ChromBPNet model sets&lt;/a&gt; across 408 biosamples.&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ProCapNet&amp;amp;type=Annotation" target="_blank" rel="noopener"
 &gt;6 ProCapNet model sets&lt;/a&gt; that predict transcription-initiation profiles.&lt;/li&gt;
&lt;li&gt;&lt;a class="link" href="https://www.encodeproject.org/search/?searchTerm=ReporterNet&amp;amp;type=Annotation&amp;amp;status=released" target="_blank" rel="noopener"
 &gt;8 ReporterNet model sets&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Additional access points include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Trained models:&lt;/strong&gt; &lt;a class="link" href="https://huggingface.co/collections/kundajelab/encode-bpnet-models" target="_blank" rel="noopener"
 &gt;ENCODE GRAMMAR models on Hugging Face&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Predictions and sequence annotations:&lt;/strong&gt; &lt;a class="link" href="https://genome.ucsc.edu/cgi-bin/hgTracks?db=hg38&amp;amp;hubUrl=https://kundajelab.github.io/ucsc-trackhub-encode.github.io/hub.txt" target="_blank" rel="noopener"
 &gt;UCSC Track Hub&lt;/a&gt;, including model predictions, sequence-contribution maps, and motif instances.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Unified motif lexicon:&lt;/strong&gt; &lt;a class="link" href="https://www.encodeproject.org/annotations/ENCSR091GRD/" target="_blank" rel="noopener"
 &gt;ENCODE annotation ENCSR091GRD&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Software:&lt;/strong&gt; &lt;a class="link" href="https://github.com/kundajelab/bpnet/" target="_blank" rel="noopener"
 &gt;BPNet code repo&lt;/a&gt;, &lt;a class="link" href="https://github.com/kundajelab/chrombpnet/" target="_blank" rel="noopener"
 &gt;ChromBPNet code repo&lt;/a&gt;, and &lt;a class="link" href="https://github.com/kundajelab/ProCapNet/" target="_blank" rel="noopener"
 &gt;ProCapNet code repo&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://doi.org/10.64898/2026.07.06.731365" target="_blank" rel="noopener"
 &gt;ENCODE 4 preprint&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://doi.org/10.1038/s41588-021-00782-6" target="_blank" rel="noopener"
 &gt;BPNet paper&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;ChromBPNet preprint&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a class="link" href="https://doi.org/10.1101/2024.05.28.596138" target="_blank" rel="noopener"
 &gt;ProCapNet preprint&lt;/a&gt;&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Please check out a detailed quick-start guide in our &lt;a class="link" href="https://genomicsxai.github.io/blogs/2026-013/" target="_blank" rel="noopener"
 &gt;&lt;strong&gt;next blog post&lt;/strong&gt; &lt;em&gt;(out now)&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="from-encode-maps-to-regulatory-sequence-rules"&gt;From ENCODE maps to regulatory sequence rules
&lt;/h2&gt;&lt;p&gt;Over two decades, ENCODE has created an unprecedented map of regulatory elements across the human and mouse genomes. GRAMMAR adds a complementary layer of predictive models and sequence annotations that connect these experimental measurements back to the underlying DNA sequence.&lt;/p&gt;
&lt;p&gt;By releasing nearly 4,000 experiment-specific model sets together with predictions, contribution maps, motifs, and motif instances, we hope to make regulatory sequence analysis more accessible, reproducible, and scalable. The goal is not to only predict where regulatory activity occurs, but to help researchers ask more mechanistic questions about &lt;strong&gt;which DNA bases, motifs, and their syntactic arrangements influence different types of biochemical activity of every regulatory element in every cellular context?&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="whats-next"&gt;What’s next?
&lt;/h2&gt;&lt;p&gt;We still have so much to share about the resource! We are planning to regularly share the many different ways you can use the resource (~every week) for the foreseeable future, so give us a follow and be on the lookout for more.&lt;/p&gt;
&lt;h2 id="references"&gt;References
&lt;/h2&gt;&lt;ol&gt;
&lt;li&gt;The ENCODE Project Consortium et al. The Encyclopedia of DNA Elements. &lt;em&gt;bioRxiv&lt;/em&gt; 2026.07.06.731365 (2026) (&lt;a class="link" href="https://doi.org/10.64898/2026.07.06.731365" target="_blank" rel="noopener"
 &gt;https://doi.org/10.64898/2026.07.06.731365&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Yun, C. M. et al. A unified lexicon of predictive DNA sequence motifs from ENCODE transcription factor binding and chromatin accessibility assays. (2025) doi:10.5281/zenodo.17179111. (&lt;a class="link" href="https://doi.org/10.5281/zenodo.17179111" target="_blank" rel="noopener"
 &gt;https://doi.org/10.5281/zenodo.17179111&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Avsec, Ž. et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. &lt;em&gt;Nat Genet&lt;/em&gt; 53, 354—366 (2021). (&lt;a class="link" href="https://doi.org/10.1038/s41588-021-00782-6" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1038/s41588-021-00782-6&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Pampari, A. et al. ChromBPNet: bias factorized, base-resolution deep learning models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants. &lt;em&gt;bioRxiv&lt;/em&gt; 2024.12.25.630221 (2024). (&lt;a class="link" href="https://doi.org/10.1101/2024.12.25.630221" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1101/2024.12.25.630221&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Cochran, K. et al. Dissecting the cis-regulatory syntax of transcription initiation with deep learning. &lt;em&gt;bioRxiv&lt;/em&gt; 2024.05.28.596138 (2024). (&lt;a class="link" href="https://doi.org/10.1101/2024.05.28.596138" target="_blank" rel="noopener"
 &gt;https://doi.org/10.1101/2024.05.28.596138&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Shrikumar, A., Greenside, P. &amp;amp; Kundaje, A. Learning Important Features Through Propagating Activation Differences. &lt;em&gt;arXIV&lt;/em&gt; (2019).(&lt;a class="link" href="https://doi.org/10.48550/arXiv.1704.02685" target="_blank" rel="noopener"
 &gt;https://doi.org/10.48550/arXiv.1704.02685&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Lundberg, S. M. &amp;amp; Lee, S.-I. A unified approach to interpreting model predictions. in &lt;em&gt;Proceedings of the 31st International Conference on Neural Information Processing Systems&lt;/em&gt; 4768–4777 (Curran Associates Inc., Red Hook, NY, USA, 2017). (&lt;a class="link" href="https://dl.acm.org/doi/10.5555/3295222.3295230" target="_blank" rel="noopener"
 &gt;https://dl.acm.org/doi/10.5555/3295222.3295230&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Shrikumar, A. et al. Technical Note on Transcription Factor Motif Discovery from Importance Scores (TF-MoDISco) version 0.5.6.5. &lt;em&gt;arXiv&lt;/em&gt; (2020) (&lt;a class="link" href="https://doi.org/10.48550/arXiv.1811.00416" target="_blank" rel="noopener"
 &gt;https://doi.org/10.48550/arXiv.1811.00416&lt;/a&gt;).&lt;/li&gt;
&lt;/ol&gt;</description></item></channel></rss>