Cheat sheet

gLM cheat sheet: Nucleotide Transformer v3

One reference card for Nucleotide Transformer v3: architecture, tokenization, size, context, training data, what it was shown to do and what it does not do. One card in a set that covers one genomic language model at a time.

GenomicsLLMMachine Learning

Genomic language model · Cheat sheet

Nucleotide Transformer v3

One backbone that reads DNA, predicts function and designs regulatory sequence

First public
bioRxiv, December 2025
Published
Preprint. Not yet peer reviewed.
Lab
InstaDeep, IMP Vienna, Cornell Tech, Cold Spring Harbor
Architecture
U-Net. A convolutional tower compresses the sequence, a Transformer bottleneck adds global context, and a matching tower expands it back.
Tokenization
Single nucleotide, single-base tokenization
Parameters
650M
Context
1,000,000 bases, with predictions restored at 1 base resolution.
Training data
9 trillion base pairs from OpenGenome2, over 128,000 species genomes, plus about 16,000 functional tracks and annotations across 24 species.
Used for
Reading DNA, predicting function and annotation, guided generation

Headline result

State of the art on the paper's own NTv3 Benchmark, 106 long-range tasks built to rule out overlap with training data, scored by correlation on functional tracks and MCC on annotation. Retrained to generate, the same backbone designed enhancers to a set activity level, and 1,000 were built and measured in cells.

What it does not do

The variant work shows its internals separate harmful from harmless expression variants, not that it outscores other variant tools. Still a preprint.

Commonly confused

The "60x" runs the other way from how it is repeated. The largest NTv3 model runs faster than competing models with up to 60x fewer parameters. Faster despite being bigger, not 60x smaller.

Download this card PNG, 1080 x 1350, sized for a post or a slide