structural-biology/ai produced the result/Briefings in Bioinformatics 2025 · v2
Model predicts how fast an enzyme works on a substrate, with uncertainty attached
Researchers built IECata, a neural network that reads an enzyme's amino acid sequence and a chemical's structure and predicts catalytic efficiency. The model also reports how confident it is, and highlights which residues and atoms it attended to.
spectrum · one line per step, placed by what the step does · bright lines used AI
IECata: interpretable bilinear attention network and evidential deep learning improve the catalytic efficiency prediction of enzymes
Briefings in Bioinformatics, 2025
doi:10.1093/bib/bbaf283 · record aix-00065 v2 · checked 2026-10-08
- AI was for
- Property prediction, Experimental design
- Model family
- Protein language model, Transformer, Graph neural network, Multilayer perceptron, Convolutional neural network
- Checked by
- Held-out806 tested
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Enzymes are proteins that speed up chemical reactions in living things. How well a particular enzyme handles a particular molecule, its substrate, is usually summarised by two numbers: kcat, how fast it turns substrate into product, and Km, roughly how much substrate it needs to get going. The ratio kcat/Km is a standard measure of catalytic efficiency. Getting that ratio normally means purifying the enzyme and measuring reaction rates in the laboratory, one enzyme-substrate pair at a time. There are vast numbers of possible pairs, so most have never been measured, and anyone designing an enzyme is working largely in the dark.
The researchers set out to predict kcat/Km directly from an enzyme's amino acid sequence and a text description of the substrate's chemical structure, known as a SMILES string. They wanted predictions that came with an estimate of their own reliability, and that could point to which parts of the enzyme and the substrate the prediction rested on. They gathered 11,815 measured entries from the BRENDA and SABIO-RK databases for training, and curated a separate set of 806 entries from published papers to test on data unlike the training material.
Where AI came in
The model is the result here. Every predicted efficiency value, every uncertainty figure and every residue-level interpretation comes out of the learned system. An existing protein language model, ProtT5, read each amino acid sequence and turned it into numbers; it had already learned patterns from protein sequences in general and its settings were left untouched. A graph network encoded each substrate as a network of atoms and bonds. A bilinear attention layer then compared every residue against every substrate atom, and an evidential output layer produced both the predicted value and a measure of doubt, split into uncertainty from gaps in the training data and uncertainty inherent to the measurements.
The prediction stands in for a bench experiment. The attention weights were read out as a guide to which residues mattered, and compared against binding sites known from crystal structures of 45 sesquiterpene synthases; among the top 25 attention residues the mean overlap was 1.911, against 1.454 for randomly chosen pockets. Predictions plus their uncertainty were also used to rank mutant enzymes, putting the ones worth testing first. Results were measured against UniKP, an earlier prediction model, both as released and retrained on the same data.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
IECata predicts enzyme catalytic efficiency (kcat/Km) directly from an enzyme amino acid sequence and a substrate SMILES string, and returns an uncertainty estimate alongside each prediction. The authors assembled 11 815 kcat/Km entries from BRENDA and SABIO-RK for training and an out-of-domain set of 806 entries curated from the literature. Reported Pearson correlation was 0.778 on the in-domain test split and 0.663 on the out-of-domain set, against 0.652 for UniKP retrained on the same training data and 0.463 for the released UniKP model; under five-fold cross-validation on the whole dataset IECata reported higher R and PCC and lower MAE and RMSE than UniKP. Bilinear attention weights were compared with crystallographic binding-site residues across 45 sesquiterpene synthase structures, giving a mean overlap of 1.911 among the top 25 attention residues against 1.454 for randomly sampled pockets.
How AI was used
Entries pairing an enzyme sequence, a substrate SMILES string and a measured kcat/Km value were pulled from BRENDA and SABIO-RK, deduplicated against PubChem standard SMILES, filtered by unit, sequence length and SMILES length, log10 transformed and split 8:1:1. Each enzyme sequence was embedded per residue by the frozen ProtT5 protein language model, whose fixed parameters were used without further training, and the embedding was passed through a light attention module and a multilayer perceptron; each substrate was built as a molecular graph with atom-type, degree, hydrogen-count and chirality features and encoded by a graph convolutional network. A bilinear attention network formed a pairwise enzyme-substrate interaction map and a bilinear pooling layer produced a joint representation, which fed an evidential output layer emitting the four parameters of a normal-inverse-gamma distribution, from which the predicted value and its epistemic, aleatoric and total uncertainty were derived. The model was trained with a negative log-likelihood term plus an evidential regulariser whose coefficient was chosen by calibration, using both the 8:1:1 split and five-fold cross-validation. Ablations swapped the enzyme encoder for integer-plus-CNN and ProtT5-plus-CNN and the joint representation for one-side attention or linear concatenation, and UniKP was retrained on the same training data as a baseline. Attention weights were then read out over residues and substrate atoms for cocrystal structures and for a curated set of sesquiterpene synthase structures, and predictions were ranked by predicted value plus uncertainty over a mutant dataset to compute hit ratios.
The shape of the work
Structural · the record, drawn
no AI
Collect kinetic entries from databases and literature
Obtaining raw data, whether by measurement, download or retrieval.
The whole dataset was retrieved from the BRENDA and SABIO-RK enzymatic databases using a custom script on 1 March 2023.where the paper describes this · verbatim
no AI
Clean, filter, transform and split the dataset
Cleaning, filtering, normalising or labelling data already obtained.
All kcat/Km values were log10-transformed. The final dataset for constructing the IECata model comprised 11 815 high-quality entrieswhere the paper describes this · verbatim
AI
Encode enzymes and substrates
Encoding data into features, descriptors, embeddings or graphs.
enzyme sequence information was extracted by the widely used protein language model ProtT5, a transformer-based self-supervised autocoder of protein sequenceswhere the paper describes this · verbatim
AI
Train the bilinear attention and evidential model
Fitting model parameters, including fine-tuning an existing model.
Evidential models were trained using a dual-objective loss functionwhere the paper describes this · verbatim
AI
Predict kcat/Km with uncertainty
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
the joint enzyme–substrate representation was fed into the EDL layer, outputting the predicted kcat/Km values as well as the uncertainty of the predictionwhere the paper describes this · verbatim
AI
Evaluate against held-out data and a baseline
Testing outputs against ground truth.
we compared IECata with the current SOTA kcat/Km prediction model, UniKPwhere the paper describes this · verbatim
AI
Interpret attention over residues and substrate atoms
Extracting understanding from model behaviour. The AI stood in for expert judgement.
the SMILES of the substrates were inputted to the IECata model to generate the corresponding attention weight matriceswhere the paper describes this · verbatim
no AI
Rank mutants by prediction plus uncertainty
Reducing a candidate set by filtering or ranking, in a single pass.
Three sorting strategies were employed to evaluate the hit ratio (HR) of the predicted labelswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's result is the predictive model itself: all reported kcat/Km values, uncertainties and residue-level interpretations are produced by the learned model
IECata used the ProtTrans (hereafter ProtT5) pretrained language model and light attention (LA) to extract enzyme featureswhere the paper describes this · verbatim
the final out-of-domain test dataset included 806 kcat/Km entrieswhere the paper describes this · verbatim
The model saved after training can be found at https://github.com/zhaoyanpeng208/IECata.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- ComputeThe hardware or time used is not stated.
- Version of IECataWhich version of the model was used is not stated.
- Version of ProtT5 (ProtTrans)Which version of the model was used is not stated.
- Version of UniKP, retrained on the IECata training datasetWhich version of the model was used is not stated.
- Version of UniKP, released trained modelWhich version of the model was used is not stated.
- What step 3 replacedThe paper gives no basis for what the AI stood in for.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 6 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00065, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error