structural-biology/ai produced the result/Nature Communications 2021 · v2
Machine learning picked which engineered enzymes to build for fatty alcohol production
Researchers stitched three bacterial enzymes into a library of 4,374 hybrids. Gaussian process models chose which hybrids to build and test across ten rounds, ending with a variant that made 54 ± 11 mg/L of fatty alcohols in E. coli.
spectrum · one line per step, placed by what the step does · bright lines used AI
Machine learning-guided acyl-ACP reductase engineering for improved in vivo fatty alcohol production
Nature Communications, 2021
doi:10.1038/s41467-021-25831-w · record aix-00024 v2 · checked 2026-10-07
- AI was for
- Property prediction, Classification, Experimental design
- Model family
- Gaussian process, Probabilistic graphical model
- Checked by
- Experimental
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Some bacteria carry enzymes that turn fatty acid building blocks into fatty alcohols, the kind of molecule used in detergents and cosmetics. Engineers would like E. coli to make them instead. One way to improve an enzyme is recombination: take several related natural versions, chop each into blocks, and swap the blocks between them to make hybrids, or chimeras. The trouble is arithmetic. Eight blocks drawn from three parents give thousands of possible combinations, and here the library held 4,374 of them. Measuring what each one produces means growing cells and running their contents through a gas chromatograph, a slow instrument that separates a mixture into its components one sample at a time.
So the full library could not be tested. The researchers set out to find high-producing chimeras by building and assaying only a small fraction of them, letting a statistical model decide which fraction that should be.
Where AI came in
The learning sat in the choosing. Twenty starting sequences were picked to be as informative as possible about the whole library, then built and measured. Their titres trained two models: a Gaussian Naïve Bayes classifier, which sorts sequences into active and inactive, and Gaussian process regression, which predicts a number and also how uncertain that prediction is. Both were then applied to every untested chimera. Predicted duds were dropped, and a rule favouring sequences that were either predicted high or still poorly understood selected ten to twelve to build next. The new measurements retrained the models, and the cycle ran ten times.
In place of testing everything, the models stood in for the untested majority of the library. A Gaussian process model was later fitted to all the collected data to estimate how much each sequence block contributed to output. The enzyme purification, the kinetic measurements and the docking calculations that followed did not involve learning.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
Three natural alcohol-forming fatty acyl reductases were recombined by SCHEMA into a library of 4374 chimeric acyl-thioester reductase domains, and a Gaussian process model with an upper-confidence-bound criterion chose which chimeras to build and assay over ten design-test-learn rounds, beginning from 20 seed sequences. The top chimera, ATR-83, produced a total titer of 54 ± 11 mg/L fatty alcohols in E. coli, 4.9-fold the titer of MA-ACR and about twofold that of the best natural sequence, MB-ACR. In vitro kinetics on palmitoyl-ACP showed a larger turnover number for ATR-83 than for parent B, with similar KM and no significant differences in enzyme expression. A Gaussian process model fitted to the sequence-function data was used to estimate each block's contribution, and net charge near the docked ACP interface correlated with titer.
How AI was used
Machine learning was used to search a combinatorial chimera space that could not be assayed exhaustively with the low-throughput gas chromatography readout. A greedy Gaussian-entropy criterion first picked 20 seed sequences that maximised mutual information with the full 4374-sequence library; these were assembled by Golden Gate, expressed in E. coli and measured. A Gaussian Naïve Bayes classifier on one-hot encoded sequences was then trained to separate active from inactive variants, and Gaussian process regression with a linear kernel over Hamming or contact-pair encodings was trained on the active sequences' titers, with the variance hyperparameter chosen by leave-one-out cross-validation. In each of ten rounds, both models were applied to all untested chimeras, inactive-predicted sequences were excluded, and a batch-mode upper-confidence-bound rule (mean plus one standard deviation, with the chosen sequence's predicted titer fed back as pseudo-data before reselecting) picked 10–12 sequences to construct and assay; the new titers retrained the models for the following round. After optimisation, a Gaussian process regression model fitted to the collected sequence-function data was used to estimate the contribution of each sequence block, alongside non-learned RosettaDock docking and interface charge calculation.
The shape of the work
Structural · the record, drawn
no AI
Design chimeric ATR library by SCHEMA recombination
Producing candidate objects that did not previously exist.
we used SCHEMA-RASPP to determine 7 additional crossover locations within the ATR domainwhere the paper describes this · verbatim
no AI
Select informative seed set by greedy entropy maximisation
Reducing a candidate set by filtering or ranking, in a single pass.
We sought to identify the set of 20 chimera sequences that is maximally informative of the full chimera landscape.where the paper describes this · verbatim
no AI
Assemble genes and measure in vivo fatty alcohol titers
Physical execution, by hand or by robot.
We then constructed these sequences and experimentally measured their fatty alcohol titers in three E. coli strainswhere the paper describes this · verbatim
AI
Train activity classifier and titer regression model
Fitting model parameters, including fine-tuning an existing model.
The fatty alcohol titer data from these 20 initial sequences was used to train Gaussian process (GP) sequence-function modelswhere the paper describes this · verbatim
AI
Predict untested chimeras and pick next batch by UCB
Iterative search over a space. The AI stood in for exhaustive search. Its result feeds back into an earlier step.
We then applied the GNB and GP models to make functional predictions over all untested chimeras.where the paper describes this · verbatim
no AI
Characterise expression and in vitro kinetics of top enzymes
Testing outputs against ground truth.
Next, we purified the enzymes and measured their kinetic properties on palmitoyl-ACPwhere the paper describes this · verbatim
AI
Model block contributions to activity
Extracting understanding from model behaviour.
We trained a GP regression model to predict fatty alcohol titers from sequence.where the paper describes this · verbatim
no AI
Dock ACP and compute interface net charge
Numerical or physics simulation, including where a learned surrogate replaces it.
We ran 1000 docking simulations and selected a model based on minimizing the total energy and the interface score.where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The improved enzymes the paper reports are the sequences chosen by the GP/UCB search; every tested variant after the seed round was selected by the model.
a Gaussian Naïve Bayes (GNB) classifier to distinguish inactive versus active sequences and Gaussian process (GP) regression to model a sequence’s fatty alcohol titerwhere the paper describes this · verbatim
measured each strain’s fatty alcohol titer using gas chromatographywhere the paper describes this · verbatim
Source data underlying Fig. 2c are provided as a Source Data file.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Gaussian process regression (linear/Hamming and structure kernels)Which version of the model was used is not stated.
- Version of Gaussian Naïve Bayes active/inactive classifier (scikit-learn)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
- What step 7 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00024, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error