~/aixsci
200 records · all checked

structural-biology/ai produced the result/arXiv 2024 · v2

A shared test set for predicting which enzyme carries out a reaction

Researchers assembled CARE, a set of curated enzyme datasets and held-out test splits, then ran machine-learning models on them, including CREEP, a model they built by fine-tuning existing language models of proteins and reactions.

1. Curate protein-EC and reaction-EC datasets2. Build train-test splits for both tasks3. Predict structures for sequences lacking database structures4. Retrain sequence classification models on Task 1 train split5. Train contrastive reaction-enzyme retrieval models6. Classify held-out sequences by EC number7. Retrieve EC numbers for held-out query reactions8. Score predictions against test labels and non-learned baselines

spectrum · one line per step, placed by what the step does · bright lines used AI

CARE: a Benchmark Suite for the Classification and Retrieval of Enzymes
arXiv, 2024

doi:10.48550/arxiv.2406.15669 · record aix-00115 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Classification, Structure determination
Model family
Transformer, Protein language model, Large language model, Graph neural network
Checked by
Held-out
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Enzymes are proteins that speed up chemical reactions in living things. Biologists label what an enzyme does with an enzyme commission number, or EC number, a four-part code that narrows from a broad class of chemistry down to a specific reaction. The trouble is that most proteins found by sequencing have never been tested in a laboratory, so their EC number is a guess. The usual way to guess is to look for a known protein with a similar sequence of amino acids. That works when a close relative is already labelled, and gets shaky when the nearest known protein is only distantly similar, or when one enzyme handles several different reactions.

The authors built CARE, a collection of curated data and test splits for two questions. The first is: given a protein sequence, what is its EC number? The second runs the other way: given a chemical reaction, which EC number of enzyme performs it? The protein set holds 185,995 sequence-EC pairs and the reaction set 61,766 reaction-EC pairs, covering 4,960 EC numbers. The test splits are designed to be awkward on purpose, holding back sequences that share under 30 per cent and 30 to 50 per cent of their sequence with the training data, plus sequences that earlier work had labelled wrongly and enzymes that do more than one job.

Where AI came in

The reported results are the scores of machine-learning models on these splits, so they exist only by running those models. The authors retrained three existing models, CLEAN, Pika and CLIPZyme, on their training data and asked them for the EC number of each held-out sequence or reaction. They also prompted two off-the-shelf systems, ChatGPT and ChemCrow, directly through their programming interfaces. And they introduced CREEP, which was made by fine-tuning two pretrained models, rxnfp for reactions and ProtT5 for protein sequences, so that a reaction and the proteins that carry it out land near each other in a shared numerical space. A version of CREEP adds written descriptions of EC numbers, encoded by a third model, SciBERT.

AI also filled a gap in the comparison baselines. One standard non-learned method searches for proteins of similar three-dimensional shape rather than similar sequence, which needs a structure for every protein. Where no structure existed in the AlphaFold database, ESMFold predicted one from the sequence, standing in for experimental structure determination. Those structures fed the Foldseek shape-search baseline, which was scored alongside random ordering, a sequence search with BLASTp and a chemical-similarity measure of reactions. On classification, the paper reports CLEAN doing best for distantly related sequences, with the sequence and shape searches close behind and shape search best on the previously misclassified split. On the harder retrieval splits, CREEP with text beat the chemical-similarity baseline, though accuracies there stayed low for every method.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

CARE is a benchmark and dataset suite for enzyme function prediction, built from a curated protein2EC set of 185,995 sequence-EC pairs and a reaction2EC set of 61,766 reaction-EC pairs covering 4,960 unique EC numbers. It defines two tasks: classifying a protein sequence by its enzyme commission number, with test splits at <30% and 30-50% sequence identity plus previously misclassified (Price et al.) and promiscuous-enzyme splits, and retrieving an EC number from a query reaction, with easy, medium and hard reaction splits. The authors retrained CLEAN, Pika and CLIPZyme on these splits, queried ChatGPT (gpt-4o-mini) and ChemCrow, and introduced CREEP, a contrastive model aligning reaction, protein sequence and EC text representations. On the classification task the paper reports that CLEAN performs best for sequences with low identity to known sequences while the BLASTp and Foldseek baselines perform similarly, with Foldseek best on the Price split; on the harder retrieval splits CREEP with text exceeds the DRFP similarity baseline, while accuracies on those splits remain low for all methods.

How AI was used

Protein sequences with EC annotations were drawn from Swiss-Prot and reaction-EC pairs from EnzymeMap supplemented by ECReact, then filtered and clustered with MMseqs2 to construct train-test splits that hold out sequences by sequence-identity band, by prior misclassification and by promiscuity, and hold out reactions at EC level 4 or EC level 3. For the classification task, CLEAN (a supervised contrastive model over ESM embeddings) and Pika (an LLM finetuned with a protein encoder) were retrained on the training split and then queried for the EC number of each test sequence, while ChatGPT and ChemCrow were prompted directly through their APIs without output cleaning. For the retrieval task, CREEP was built by finetuning the pretrained rxnfp reaction language model and the ProtT5 protein language model, with an optional third modality of gene-ontology EC descriptions encoded by SciBERT, using an EBM-NCE contrastive objective that projects all modalities into a 256-dimensional shared space and batches triplets from distinct EC numbers over 40 epochs; CLIPZyme was retrained on the same splits with its EGNN structure encoder. At inference, a query reaction (or text) representation was compared to per-EC centroids of protein representations to produce an EC ranking. ESMFold supplied structures absent from the AlphaFold database so that a Foldseek structural-search baseline could be run, and non-learned baselines (random ordering, Diamond BLASTp, Foldseek, DRFP reaction similarity) were scored alongside the models using k=1 accuracy at each EC level.

The shape of the work

Structural · the record, drawn

PREPARATIONPREPARATIONINFERENCETRAININGTRAININGINFERENCEINFERENCEVALIDATION12345678AIAIAIAIAICurate protein-ECand reaction-ECdatasetsBuild train-testsplits for bothtasksPredictstructures forsequences lackin…Retrain sequenceclassificationmodels on Task 1…Train contrastivereaction-enzymeretrieval modelsClassify held-outsequences by ECnumberRetrieve ECnumbers forheld-out query r…Score predictionsagainst testlabels and non-l…↤ new capability↤ conventional algorithm↤ conventional algorithm
AI stepNo AI↤ what the AI stood in for
1Preparation
no AI

Curate protein-EC and reaction-EC datasets

Cleaning, filtering, normalising or labelling data already obtained.

Protein sequence data, paired to EC number(s), were downloaded from UniProt, selecting only reviewed sequences in Swiss-Protwhere the paper describes this · verbatim
in the paper
2Preparation
no AI

Build train-test splits for both tasks

Cleaning, filtering, normalising or labelling data already obtained.

Clustering using MMseqs2 was performed at 30%, 50%, 70%, and 90% sequence identitywhere the paper describes this · verbatim
in the paper
3Inference
AI

Predict structures for sequences lacking database structures

Running a trained model over new data to predict, classify or score.

Missing structures were folded with ESMFoldwhere the paper describes this · verbatim
in the paper
4Training
AI

Retrain sequence classification models on Task 1 train split

Fitting model parameters, including fine-tuning an existing model.

we retrained CLEAN using our training set with only one example from each cluster at 50% identitywhere the paper describes this · verbatim
in the paper
5Training
AI

Train contrastive reaction-enzyme retrieval models

Fitting model parameters, including fine-tuning an existing model. The AI stood in for new capability.

We train for 40 epochs, and in each epoch, we loop over each EC numberwhere the paper describes this · verbatim
in the paper
6Inference
AI

Classify held-out sequences by EC number

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

a query protein is passed through a trained model to predict an EC numberwhere the paper describes this · verbatim
in the paper
7Inference
AI

Retrieve EC numbers for held-out query reactions

Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.

a query reaction is passed through a trained model to perform retrieval to an EC numberwhere the paper describes this · verbatim
in the paper
8Validation
no AI

Score predictions against test labels and non-learned baselines

Testing outputs against ground truth.

Accuracy is calculated as the number of true answers over the total number of exampleswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's reported findings are the performance numbers of machine-learned models on the curated splits, and the CREEP model is itself a contribution, so the results exist only through running models

+What the AI was for
CREEP leverages finetuning of pretrained language models, rxnfp and ProtT5 to learn aligned representations of reactions and proteinswhere the paper describes this · verbatim
+How it was taught
Self-supervisedSupervisedTransfer / fine-tuningZero-shotin the paper
+Models named
CREEP · Fine-tunedrxnfp · Fine-tunedProtT5 · Fine-tunedSciBERT · Fine-tunedCLIPZyme · Trained from scratchCLEAN · Trained from scratchPika · Trained from scratchChatGPT gpt-4o-mini · Off the shelfChemCrow public version · Off the shelfESMFold · Off the shelfin the paper
+How results were checked
Held-outin the paper
Model training can use any of the data in the train split, and each model is evaluated on the associated test splitwhere the paper describes this · verbatim
+Code · weights · data
code availableweights not reporteddata availablein the paper
The curated datasets and splits used in CARE can be accessed at https://github.com/jsunn-y/CARE/.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 13 items
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • How many were testedThe paper gives no count of what was tested.
  • Version of CREEPWhich version of the model was used is not stated.
  • Version of rxnfpWhich version of the model was used is not stated.
  • Version of ProtT5Which version of the model was used is not stated.
  • Version of SciBERTWhich version of the model was used is not stated.
  • Version of CLIPZymeWhich version of the model was used is not stated.
  • Version of CLEANWhich version of the model was used is not stated.
  • Version of PikaWhich version of the model was used is not stated.
  • Version of ESMFoldWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00115, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error