Cheat sheet

gLM cheat sheet: Evo

One reference card for Evo: architecture, tokenization, size, context, training data, what it was shown to do and what it does not do. One card in a set that covers one genomic language model at a time.

GenomicsLLMMachine Learning

Genomic language model · Cheat sheet

Evo

The first genomic language model whose designed sequences were built and tested in a lab

First public
bioRxiv, February 2024
Published
Science, November 2024
Lab
Arc Institute, Stanford, TogetherAI
Architecture
StripedHyena. Long convolutions with a few attention layers. First genomic model on it; the architecture came from text modelling in 2023.
Tokenization
Single nucleotide, byte-level tokenizer
Parameters
7B
Context
131,072 tokens. Pretrained at 8,192, then extended in a second stage.
Training data
OpenGenome. 300 billion nucleotides from bacteria, archaea, phage and plasmid.
Used for
Prediction and generation, from molecular to genome scale

Headline result

Function prediction with no task-specific training, competitive with models built for each task. Designed gene editing systems (CRISPR-Cas) and mobile genetic elements, and the designs were built and tested: one Evo-designed guide RNA cut DNA in a test tube about as well as the standard SpCas9 enzyme.

What it does not do

Bacteria and archaea only. No eukaryotic genomes in training, which is exactly what Evo 2 changed.

Commonly confused

Often mixed up with Evo 2. The 1 million token context belongs to Evo 2, not to Evo. Evo's megabase generations were checked by computer, not built.

Download this card PNG, 1080 x 1350, sized for a post or a slide