~/aixsci
200 records · all checked

structural-biology/ai produced the result/Briefings in Bioinformatics 2025 · v2

Model predicts how fast an enzyme works on a substrate, with uncertainty attached

Researchers built IECata, a neural network that reads an enzyme's amino acid sequence and a chemical's structure and predicts catalytic efficiency. The model also reports how confident it is, and highlights which residues and atoms it attended to.

1. Collect kinetic entries from databases and literature2. Clean, filter, transform and split the dataset3. Encode enzymes and substrates4. Train the bilinear attention and evidential model5. Predict kcat/Km with uncertainty6. Evaluate against held-out data and a baseline7. Interpret attention over residues and substrate atoms8. Rank mutants by prediction plus uncertainty

spectrum · one line per step, placed by what the step does · bright lines used AI

IECata: interpretable bilinear attention network and evidential deep learning improve the catalytic efficiency prediction of enzymes
Briefings in Bioinformatics, 2025

doi:10.1093/bib/bbaf283 · record aix-00065 v2 · checked 2026-10-08

ai-resultrole of AI
AI was for
Property prediction, Experimental design
Model family
Protein language model, Transformer, Graph neural network, Multilayer perceptron, Convolutional neural network
Checked by
Held-out806 tested
Code
available

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Enzymes are proteins that speed up chemical reactions in living things. How well a particular enzyme handles a particular molecule, its substrate, is usually summarised by two numbers: kcat, how fast it turns substrate into product, and Km, roughly how much substrate it needs to get going. The ratio kcat/Km is a standard measure of catalytic efficiency. Getting that ratio normally means purifying the enzyme and measuring reaction rates in the laboratory, one enzyme-substrate pair at a time. There are vast numbers of possible pairs, so most have never been measured, and anyone designing an enzyme is working largely in the dark.

The researchers set out to predict kcat/Km directly from an enzyme's amino acid sequence and a text description of the substrate's chemical structure, known as a SMILES string. They wanted predictions that came with an estimate of their own reliability, and that could point to which parts of the enzyme and the substrate the prediction rested on. They gathered 11,815 measured entries from the BRENDA and SABIO-RK databases for training, and curated a separate set of 806 entries from published papers to test on data unlike the training material.

Where AI came in

The model is the result here. Every predicted efficiency value, every uncertainty figure and every residue-level interpretation comes out of the learned system. An existing protein language model, ProtT5, read each amino acid sequence and turned it into numbers; it had already learned patterns from protein sequences in general and its settings were left untouched. A graph network encoded each substrate as a network of atoms and bonds. A bilinear attention layer then compared every residue against every substrate atom, and an evidential output layer produced both the predicted value and a measure of doubt, split into uncertainty from gaps in the training data and uncertainty inherent to the measurements.

The prediction stands in for a bench experiment. The attention weights were read out as a guide to which residues mattered, and compared against binding sites known from crystal structures of 45 sesquiterpene synthases; among the top 25 attention residues the mean overlap was 1.911, against 1.454 for randomly chosen pockets. Predictions plus their uncertainty were also used to rank mutant enzymes, putting the ones worth testing first. Results were measured against UniKP, an earlier prediction model, both as released and retrained on the same data.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

IECata predicts enzyme catalytic efficiency (kcat/Km) directly from an enzyme amino acid sequence and a substrate SMILES string, and returns an uncertainty estimate alongside each prediction. The authors assembled 11 815 kcat/Km entries from BRENDA and SABIO-RK for training and an out-of-domain set of 806 entries curated from the literature. Reported Pearson correlation was 0.778 on the in-domain test split and 0.663 on the out-of-domain set, against 0.652 for UniKP retrained on the same training data and 0.463 for the released UniKP model; under five-fold cross-validation on the whole dataset IECata reported higher R and PCC and lower MAE and RMSE than UniKP. Bilinear attention weights were compared with crystallographic binding-site residues across 45 sesquiterpene synthase structures, giving a mean overlap of 1.911 among the top 25 attention residues against 1.454 for randomly sampled pockets.

How AI was used

Entries pairing an enzyme sequence, a substrate SMILES string and a measured kcat/Km value were pulled from BRENDA and SABIO-RK, deduplicated against PubChem standard SMILES, filtered by unit, sequence length and SMILES length, log10 transformed and split 8:1:1. Each enzyme sequence was embedded per residue by the frozen ProtT5 protein language model, whose fixed parameters were used without further training, and the embedding was passed through a light attention module and a multilayer perceptron; each substrate was built as a molecular graph with atom-type, degree, hydrogen-count and chirality features and encoded by a graph convolutional network. A bilinear attention network formed a pairwise enzyme-substrate interaction map and a bilinear pooling layer produced a joint representation, which fed an evidential output layer emitting the four parameters of a normal-inverse-gamma distribution, from which the predicted value and its epistemic, aleatoric and total uncertainty were derived. The model was trained with a negative log-likelihood term plus an evidential regulariser whose coefficient was chosen by calibration, using both the 8:1:1 split and five-fold cross-validation. Ablations swapped the enzyme encoder for integer-plus-CNN and ProtT5-plus-CNN and the joint representation for one-side attention or linear concatenation, and UniKP was retrained on the same training data as a baseline. Attention weights were then read out over residues and substrate atoms for cocrystal structures and for a curated set of sesquiterpene synthase structures, and predictions were ranked by predicted value plus uncertainty over a mutant dataset to compute hit ratios.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONTRAININGINFERENCEVALIDATIONINTERPRETATIONSCREENING12345678AIAIAIAIAICollect kineticentries fromdatabases and li…Clean, filter,transform andsplit the datasetEncode enzymesand substratesTrain thebilinearattention and ev…Predict kcat/Kmwith uncertaintyEvaluate againstheld-out data anda baselineInterpretattention overresidues and sub…Rank mutants byprediction plusuncertainty↤ physical experiment↤ expert judgement
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Collect kinetic entries from databases and literature

Obtaining raw data, whether by measurement, download or retrieval.

The whole dataset was retrieved from the BRENDA and SABIO-RK enzymatic databases using a custom script on 1 March 2023.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Clean, filter, transform and split the dataset

Cleaning, filtering, normalising or labelling data already obtained.

All kcat/Km values were log10-transformed. The final dataset for constructing the IECata model comprised 11 815 high-quality entrieswhere the paper describes this · verbatim
in the paper
3Representation
AI

Encode enzymes and substrates

Encoding data into features, descriptors, embeddings or graphs.

enzyme sequence information was extracted by the widely used protein language model ProtT5, a transformer-based self-supervised autocoder of protein sequenceswhere the paper describes this · verbatim
in the paper
4Training
AI

Train the bilinear attention and evidential model

Fitting model parameters, including fine-tuning an existing model.

Evidential models were trained using a dual-objective loss functionwhere the paper describes this · verbatim
in the paper
5Inference
AI

Predict kcat/Km with uncertainty

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

the joint enzyme–substrate representation was fed into the EDL layer, outputting the predicted kcat/Km values as well as the uncertainty of the predictionwhere the paper describes this · verbatim
in the paper
6Validation
AI

Evaluate against held-out data and a baseline

Testing outputs against ground truth.

we compared IECata with the current SOTA kcat/Km prediction model, UniKPwhere the paper describes this · verbatim
in the paper
7Interpretation
AI

Interpret attention over residues and substrate atoms

Extracting understanding from model behaviour. The AI stood in for expert judgement.

the SMILES of the substrates were inputted to the IECata model to generate the corresponding attention weight matriceswhere the paper describes this · verbatim
in the paper
8Screening
no AI

Rank mutants by prediction plus uncertainty

Reducing a candidate set by filtering or ranking, in a single pass.

Three sorting strategies were employed to evaluate the hit ratio (HR) of the predicted labelswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the predictive model itself: all reported kcat/Km values, uncertainties and residue-level interpretations are produced by the learned model

+What the AI was for
IECata used the ProtTrans (hereafter ProtT5) pretrained language model and light attention (LA) to extract enzyme featureswhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedin the paper
+Models named
IECata · Trained from scratchProtT5 (ProtTrans) · Off the shelfUniKP, retrained on the IECata training dataset · Trained from scratchUniKP, released trained model · Off the shelfin the paper
+How results were checked
Held-out806 testedin the paper
the final out-of-domain test dataset included 806 kcat/Km entrieswhere the paper describes this · verbatim
+Code · weights · data
code availableweights availabledata availablein the paper
The model saved after training can be found at https://github.com/zhaoyanpeng208/IECata.where the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 8 items
  • ComputeThe hardware or time used is not stated.
  • Version of IECataWhich version of the model was used is not stated.
  • Version of ProtT5 (ProtTrans)Which version of the model was used is not stated.
  • Version of UniKP, retrained on the IECata training datasetWhich version of the model was used is not stated.
  • Version of UniKP, released trained modelWhich version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 6 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00065, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error