structural-biology/ai produced the result/arXiv 2024 · v2
A shared test set for predicting which enzyme carries out a reaction
Researchers assembled CARE, a set of curated enzyme datasets and held-out test splits, then ran machine-learning models on them, including CREEP, a model they built by fine-tuning existing language models of proteins and reactions.
spectrum · one line per step, placed by what the step does · bright lines used AI
CARE: a Benchmark Suite for the Classification and Retrieval of Enzymes
arXiv, 2024
doi:10.48550/arxiv.2406.15669 · record aix-00115 v2 · checked 2026-10-08
- AI was for
- Classification, Structure determination
- Model family
- Transformer, Protein language model, Large language model, Graph neural network
- Checked by
- Held-out
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Enzymes are proteins that speed up chemical reactions in living things. Biologists label what an enzyme does with an enzyme commission number, or EC number, a four-part code that narrows from a broad class of chemistry down to a specific reaction. The trouble is that most proteins found by sequencing have never been tested in a laboratory, so their EC number is a guess. The usual way to guess is to look for a known protein with a similar sequence of amino acids. That works when a close relative is already labelled, and gets shaky when the nearest known protein is only distantly similar, or when one enzyme handles several different reactions.
The authors built CARE, a collection of curated data and test splits for two questions. The first is: given a protein sequence, what is its EC number? The second runs the other way: given a chemical reaction, which EC number of enzyme performs it? The protein set holds 185,995 sequence-EC pairs and the reaction set 61,766 reaction-EC pairs, covering 4,960 EC numbers. The test splits are designed to be awkward on purpose, holding back sequences that share under 30 per cent and 30 to 50 per cent of their sequence with the training data, plus sequences that earlier work had labelled wrongly and enzymes that do more than one job.
Where AI came in
The reported results are the scores of machine-learning models on these splits, so they exist only by running those models. The authors retrained three existing models, CLEAN, Pika and CLIPZyme, on their training data and asked them for the EC number of each held-out sequence or reaction. They also prompted two off-the-shelf systems, ChatGPT and ChemCrow, directly through their programming interfaces. And they introduced CREEP, which was made by fine-tuning two pretrained models, rxnfp for reactions and ProtT5 for protein sequences, so that a reaction and the proteins that carry it out land near each other in a shared numerical space. A version of CREEP adds written descriptions of EC numbers, encoded by a third model, SciBERT.
AI also filled a gap in the comparison baselines. One standard non-learned method searches for proteins of similar three-dimensional shape rather than similar sequence, which needs a structure for every protein. Where no structure existed in the AlphaFold database, ESMFold predicted one from the sequence, standing in for experimental structure determination. Those structures fed the Foldseek shape-search baseline, which was scored alongside random ordering, a sequence search with BLASTp and a chemical-similarity measure of reactions. On classification, the paper reports CLEAN doing best for distantly related sequences, with the sequence and shape searches close behind and shape search best on the previously misclassified split. On the harder retrieval splits, CREEP with text beat the chemical-similarity baseline, though accuracies there stayed low for every method.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
CARE is a benchmark and dataset suite for enzyme function prediction, built from a curated protein2EC set of 185,995 sequence-EC pairs and a reaction2EC set of 61,766 reaction-EC pairs covering 4,960 unique EC numbers. It defines two tasks: classifying a protein sequence by its enzyme commission number, with test splits at <30% and 30-50% sequence identity plus previously misclassified (Price et al.) and promiscuous-enzyme splits, and retrieving an EC number from a query reaction, with easy, medium and hard reaction splits. The authors retrained CLEAN, Pika and CLIPZyme on these splits, queried ChatGPT (gpt-4o-mini) and ChemCrow, and introduced CREEP, a contrastive model aligning reaction, protein sequence and EC text representations. On the classification task the paper reports that CLEAN performs best for sequences with low identity to known sequences while the BLASTp and Foldseek baselines perform similarly, with Foldseek best on the Price split; on the harder retrieval splits CREEP with text exceeds the DRFP similarity baseline, while accuracies on those splits remain low for all methods.
How AI was used
Protein sequences with EC annotations were drawn from Swiss-Prot and reaction-EC pairs from EnzymeMap supplemented by ECReact, then filtered and clustered with MMseqs2 to construct train-test splits that hold out sequences by sequence-identity band, by prior misclassification and by promiscuity, and hold out reactions at EC level 4 or EC level 3. For the classification task, CLEAN (a supervised contrastive model over ESM embeddings) and Pika (an LLM finetuned with a protein encoder) were retrained on the training split and then queried for the EC number of each test sequence, while ChatGPT and ChemCrow were prompted directly through their APIs without output cleaning. For the retrieval task, CREEP was built by finetuning the pretrained rxnfp reaction language model and the ProtT5 protein language model, with an optional third modality of gene-ontology EC descriptions encoded by SciBERT, using an EBM-NCE contrastive objective that projects all modalities into a 256-dimensional shared space and batches triplets from distinct EC numbers over 40 epochs; CLIPZyme was retrained on the same splits with its EGNN structure encoder. At inference, a query reaction (or text) representation was compared to per-EC centroids of protein representations to produce an EC ranking. ESMFold supplied structures absent from the AlphaFold database so that a Foldseek structural-search baseline could be run, and non-learned baselines (random ordering, Diamond BLASTp, Foldseek, DRFP reaction similarity) were scored alongside the models using k=1 accuracy at each EC level.
The shape of the work
Structural · the record, drawn
no AI
Curate protein-EC and reaction-EC datasets
Cleaning, filtering, normalising or labelling data already obtained.
Protein sequence data, paired to EC number(s), were downloaded from UniProt, selecting only reviewed sequences in Swiss-Protwhere the paper describes this · verbatim
no AI
Build train-test splits for both tasks
Cleaning, filtering, normalising or labelling data already obtained.
Clustering using MMseqs2 was performed at 30%, 50%, 70%, and 90% sequence identitywhere the paper describes this · verbatim
AI
Predict structures for sequences lacking database structures
Running a trained model over new data to predict, classify or score.
Missing structures were folded with ESMFoldwhere the paper describes this · verbatim
AI
Retrain sequence classification models on Task 1 train split
Fitting model parameters, including fine-tuning an existing model.
we retrained CLEAN using our training set with only one example from each cluster at 50% identitywhere the paper describes this · verbatim
AI
Train contrastive reaction-enzyme retrieval models
Fitting model parameters, including fine-tuning an existing model. The AI stood in for new capability.
We train for 40 epochs, and in each epoch, we loop over each EC numberwhere the paper describes this · verbatim
AI
Classify held-out sequences by EC number
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
a query protein is passed through a trained model to predict an EC numberwhere the paper describes this · verbatim
AI
Retrieve EC numbers for held-out query reactions
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
a query reaction is passed through a trained model to perform retrieval to an EC numberwhere the paper describes this · verbatim
no AI
Score predictions against test labels and non-learned baselines
Testing outputs against ground truth.
Accuracy is calculated as the number of true answers over the total number of exampleswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's reported findings are the performance numbers of machine-learned models on the curated splits, and the CREEP model is itself a contribution, so the results exist only through running models
CREEP leverages finetuning of pretrained language models, rxnfp and ProtT5 to learn aligned representations of reactions and proteinswhere the paper describes this · verbatim
Model training can use any of the data in the train split, and each model is evaluated on the associated test splitwhere the paper describes this · verbatim
The curated datasets and splits used in CARE can be accessed at https://github.com/jsunn-y/CARE/.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of CREEPWhich version of the model was used is not stated.
- Version of rxnfpWhich version of the model was used is not stated.
- Version of ProtT5Which version of the model was used is not stated.
- Version of SciBERTWhich version of the model was used is not stated.
- Version of CLIPZymeWhich version of the model was used is not stated.
- Version of CLEANWhich version of the model was used is not stated.
- Version of PikaWhich version of the model was used is not stated.
- Version of ESMFoldWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00115, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error