structural-biology/ai produced the result/Advanced Science 2026 · v2
Attention patterns inside a protein language model used to cut sequences into reusable units
Researchers split protein sequences into short segments they call protein words, using the internal attention patterns of the ESM2 language model, then trained a second model to link those words to molecular functions.
spectrum · one line per step, placed by what the step does · bright lines used AI
Automatically Defining Protein Words for Diverse Functional Predictions Based on Attention Analysis of a Protein Language Model
Advanced Science, 2026
doi:10.1002/advs.202521970 · record aix-00072 v2 · checked 2026-10-08
- AI was for
- Segmentation, Classification
- Model family
- Protein language model, Transformer
- Checked by
- Benchmark182 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are long chains of amino acids, and a chain's order determines what the protein does. Only some stretches matter for a given job: a few residues may grip a metal ion, bind DNA or form the pocket where a chemical reaction happens. Finding those stretches in a new sequence is hard. The classic approach is a dictionary of motifs, short patterns written down by curators from proteins already studied in the laboratory. That works when a new protein resembles an old one, but much of the protein world does not. Whole families are labelled domains of unknown function, meaning nobody has yet pinned a job to them.
The researchers set out to build such units automatically, without human curation. The idea was to divide a sequence into short segments, keep the ones that recur across many proteins, collect them into dictionaries, and then test whether those segments land on residues already known to matter and whether they can be mapped to functions.
Where AI came in
The segmentation came from a protein language model, ESM2-650M, which had been trained on large numbers of sequences to predict residues from their surroundings. Such a model builds attention matrices, internal tables recording which positions in a sequence it treats as related. Each sequence was passed through the model as supplied, without further tuning, yielding 660 such matrices and a 1280-number description of each residue. The matrices were then thresholded and cut into residue groups by a graph algorithm, Louvain, which finds clusters of densely connected points. Groups of 5 to 20 residues were kept as raw words. The model stood in for the human judgement behind curated motif dictionaries.
A second model, Word2Function, was trained from scratch on descriptions of these words, drawn from 56 299 sequences carrying 65 labels for molecular function, to predict those labels for a protein. An attribution method, Integrated Gradients, then scored how much each word contributed to a prediction, and the top-scoring words were written into a table keyed to the 65 functions. Performance was compared with the curated motif tools PROSITE and InterProScan and with other trained predictors.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study parses protein sequences into units it calls 'protein words' by binarizing the 660 attention matrices that the ESM2-650M protein language model produces for a sequence and clustering residues with the Louvain community detection algorithm, then compiles dictionaries of high-occurrence words from 1 million UniRef50 sequences and from individual Pfam families (20 762 family dictionaries). A supervised transformer model, Word2Function, was trained on word embeddings from 56 299 sequences labelled with 65 GO molecular-function terms, and Integrated Gradients attributions were used to assign individual words to GO terms in a table called WordTableGO65. On 182 proteins from the ProteinGym deep mutational scanning set, median functional residue coverage was 0.900 with the combined Pfam and UniRef50 dictionaries, compared with 0.045 for the motif-based tool PROSITE and 0.017 for InterProScan; for whole-protein GO-term prediction on the PWNet dataset, mean functional MCC was 0.284 for Word2Function and 0.150 for PROSITE. The word table was also matched against 1 494 146 InterPro sequences annotated as domains of unknown function and against MHC peptides from the MHC Motif Atlas, where the UniRef50 dictionary matched 2832 MHC-I peptides versus 373 for a random k-mer dictionary.
How AI was used
A pre-trained protein language model, ESM2-650M, was run without fine-tuning on each input sequence (length up to 1024, canonical residues only) to obtain 660 attention matrices and 1280-dimensional per-residue embeddings. Each attention matrix was binarized with a dual cutoff (a fixed threshold of 0.1, plus a proportional threshold discarding the lowest 20% of scores in heads with diffuse attention), converted to a directed residue graph, and partitioned with the Louvain algorithm (resolution 1.0, modularity gain threshold 1e-7); communities of 5–20 residues, contiguous or with up to two gaps, were kept as raw words. Running this in multi-sequence mode over 1 million UniRef50 sequences (50 per each of 20 000 Pfam families) and over individual families produced a common dictionary and family dictionaries, with the 20 amino acids collapsed into 12 degenerate types and length-stratified occurrence thresholds; in single-sequence mode, raw words for an analyte sequence were degenerated and matched exactly against these dictionaries to prioritize words. For function mapping, word embeddings were formed by averaging the ESM2 residue embeddings inside each word and fed to transformer layers plus a fully connected layer (embedding and hidden dimension 1280, 20 attention heads) trained with binary cross-entropy, Adam, batch size 256, weight decay 1e-5, initial learning rate 1e-3, cosine decay and early stopping, to predict 65 GO terms over the ExpGO65 sequences. Integrated Gradients was then applied to the trained model to score each word's contribution, and words whose cumulative positive attribution reached a cutoff of half the total were written to the WordTableGO65 table, which was subsequently used for whole-protein annotation by exact word matching. Baselines were obtained by submitting sequences to ScanProsite and InterProScan, retraining ProtNote on the same ExpGO65 splits, and querying the GPSFun web server.
The shape of the work
Structural · the record, drawn
no AI
Assemble sequence and annotation datasets
Obtaining raw data, whether by measurement, download or retrieval.
Protein sequences were downloaded from UniProt (2024/01)where the paper describes this · verbatim
AI
Extract ESM2 attention matrices and residue embeddings
Encoding data into features, descriptors, embeddings or graphs.
By inputting an analyte protein sequence into ESM2, we obtain 660 attention matrices from the corresponding attention heads.where the paper describes this · verbatim
no AI
Binarize attention and cluster residues into raw words
Extracting understanding from model behaviour.
We then use the Louvain algorithm, a graph‐based community detection algorithm, to segment the binary matrix into communities.where the paper describes this · verbatim
no AI
Compile degenerate UniRef50 and Pfam dictionaries
Cleaning, filtering, normalising or labelling data already obtained.
We stratified the dictionary by length, retaining high occurrence raw words for each length.where the paper describes this · verbatim
no AI
Prioritize words for an analyte sequence by dictionary lookup
Reducing a candidate set by filtering or ranking, in a single pass.
we convert raw words into their degenerate forms and then search for exact matcheswhere the paper describes this · verbatim
AI
Train Word2Function on word embeddings to predict GO terms
Fitting model parameters, including fine-tuning an existing model. The AI stood in for manual curation.
During training, we employed binary cross‐entropy loss to optimize the model, using a batch size of 256where the paper describes this · verbatim
AI
Attribute words to GO terms and populate WordTableGO65
Extracting understanding from model behaviour. The AI stood in for manual curation.
we use the Integrated Gradients algorithm to quantify the contribution of each word to the predictionwhere the paper describes this · verbatim
AI
Evaluate coverage and GO prediction against baselines
Testing outputs against ground truth.
The median and mean values for functional residue coverage using PROSITE were 0.045 and 0.393where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The object of the paper is the definition of 'protein words' obtained from ESM2 attention matrices and the GO-term mapping learned on those words; every reported result derives from the model's attention or from the trained Word2Function model.
The original ESM2‐650 M variant was run for all input sequences (without specific finetuning) to generate attention matriceswhere the paper describes this · verbatim
When applied to 182 proteins in the DMS dataset with sequence lengths between 50 and 1024 amino acidswhere the paper describes this · verbatim
The data that support the findings of this study are openly available in ProteinWordwise atwhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Word2FunctionWhich version of the model was used is not stated.
- Version of ProtNoteWhich version of the model was used is not stated.
- Version of GPSFunWhich version of the model was used is not stated.
- What step 2 replacedThe paper gives no basis for what the AI stood in for.
- What step 8 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00072, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error