~/aixsci
200 records · all checked

structural-biology/ai produced the result/PLoS Computational Biology 2017 · v2

Gaussian process models pick light-sensitive ion channels that reach mammalian cell membranes

Researchers measured how 218 stitched-together channelrhodopsin proteins behaved in human cells, then trained Gaussian process models on those measurements to predict, out of 118,098 possible variants, which untested ones would reach the cell membrane and which to build next.

1. Design SCHEMA recombination library2. Synthesize and assay training-set chimeras in HEK cells3. Encode sequence and contact-map features4. Train GP classification and regression models5. Predict across the library and select exploration and verification sets6. Select optimal chimeras and CsCbChR1 block swaps by confidence bounds7. Synthesize and measure model-selected ChR variants8. Weight sequence and contact features for interpretation

spectrum · one line per step, placed by what the step does · bright lines used AI

Machine learning to design integral membrane channelrhodopsins for efficient eukaryotic expression and plasma membrane localization
PLoS Computational Biology, 2017

doi:10.1371/journal.pcbi.1005786 · record aix-00009 v2 · checked 2026-10-07

ai-resultrole of AI
AI was for
Property prediction, Classification, Experimental design
Model family
Gaussian process, Linear model
Checked by
Experimental11 tested, 11 worked
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Crystal structure of a channelrhodopsin with colored spheres and sticks marking residues and contacts linked to membrane localization.
Structural features from three parent channelrhodopsins predicting membrane localization, shown on the C1C2 crystal structure.Fig 5 from Bedbrook et al., PLoS Computational Biology 2017 · source · CC BY · resized

Channelrhodopsins are proteins from algae that sit in a cell's outer membrane and open to let ions through when light hits them. Biologists put them into other cells so that light can switch those cells on or off. For that to work, the protein must first be made by the host cell and then travel to the outer membrane, called the plasma membrane. Many natural channelrhodopsins do neither well in mammalian cells. One way to find better ones is to shuffle pieces of several parent proteins into hybrids, known as chimeras. The trouble is arithmetic: the shuffling scheme here allows 118,098 combinations, far more than anyone can build and test.

The researchers took three parent channelrhodopsins, used a crystal structure of a related protein to decide where the sequences could be cut and rejoined, and built a library of possible hybrids on paper. They then made 218 of them, expressed them in human embryonic kidney cells, and measured two things with fluorescent tags: how much protein was made, and how much of it reached the plasma membrane.

Where AI came in

Those 218 measurements became training data. The team fitted Gaussian process models, a statistical method that predicts a value for an untried case and also reports how uncertain that prediction is. Proteins were described to the models as lists of which amino acids were present and which pairs of residues touch each other in the known structure. Some models sorted variants into 'high' or 'low' for expression and membrane localisation; another predicted the localisation level as a number. Simpler linear regressions picked out which sequence and contact features mattered, and later assigned weights that were drawn onto the structure.

The models then stood in for the screening the researchers could not do by hand. Applied across the whole library, they chose a set of varied hybrids and a separate set to check their own accuracy; the model sorted all eleven of the latter correctly. Using its uncertainty estimates, the regression model nominated four hybrids in the top 0.1% of its predictions, and nominated single-block swaps into a natural channelrhodopsin. New measurements were fed back to retrain the models. The variants tested were the ones the models selected.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

Crystal structure of a channelrhodopsin with colored spheres and sticks marking residues and contacts linked to membrane localization.
Structural features from three parent channelrhodopsins predicting membrane localization, shown on the C1C2 crystal structure.Fig 5 from Bedbrook et al., PLoS Computational Biology 2017 · source · CC BY · resized

The authors measured expression and plasma-membrane localization in HEK cells for 218 channelrhodopsin chimeras drawn from a 118,098-variant library designed by SCHEMA recombination of three parent channelrhodopsins. Gaussian process classification and regression models were trained on those measurements using kernels built from sequence and residue-residue contact features, then used to predict which untested chimeras express and localize and to pick further variants for synthesis. Four chimeras whose lower-confidence-bound predictions ranked in the top 0.1% of the library were built and all localized as well as or better than the best-localizing parent, CsChrimR, and model-selected single-block swaps raised the localization of a natural channelrhodopsin, CbChR1, that does not localize in mammalian cells. L1- and L2-regularized linear regression on the same features was used to weight the residues and contacts associated with localization.

How AI was used

Expression and localization data from 218 SCHEMA-recombination chimeras were used to train Gaussian process binary classification models (Laplace approximation for the intractable posterior) for 'high' versus 'low' expression, localization and localization efficiency, and a Gaussian process regression model for localization level. Sequences were compared through kernels over binary sequence feature vectors and residue-residue contact-map feature vectors derived from the C1C2 crystal structure; linear, squared exponential and Matern kernels were compared, and kernel form and hyperparameters were set by maximizing the marginal likelihood. For regression, L1-regularized linear regression first selected a subset of sequence and contact features, and the GP was trained on the non-zero-weight features; regularization strength and model performance were assessed by leave-one-out cross-validation. The classification model was applied across the full library to choose a diverse exploration set and a verification set for synthesis, and those measurements were added to retrain the models. The regression model was then applied with a lower confidence bound acquisition to pick library chimeras for synthesis and with an upper confidence bound acquisition to pick single-block swaps into the natural variant CsCbChR1. Separately, Bayesian ridge regression on the L1-selected features produced feature weights that were mapped onto the C1C2 structure. Modelling used open-source SciPy-ecosystem packages and scikit-learn.

The shape of the work

Structural · the record, drawn

GENERATIONEXPERIMENTREPRESENTATIONTRAININGINFERENCESCREENINGVALIDATIONINTERPRETATION12345678AIAIAIAIDesign SCHEMArecombinationlibrarySynthesize andassaytraining-set chi…Encode sequenceand contact-mapfeaturesTrain GPclassificationand regression m…Predict acrossthe library andselect explorati…Select optimalchimeras andCsCbChR1 block s…Synthesize andmeasuremodel-selected C…Weight sequenceand contactfeatures for int…↤ physical experiment↤ physical experiment↤ expert judgementloops back
AI stepNo AI↤ what the AI stood in for
1Generation
no AI

Design SCHEMA recombination library

Producing candidate objects that did not previously exist.

These designs generate 118,098 possible chimeraswhere the paper describes this · verbatim
in the paper
2Experiment
no AI

Synthesize and assay training-set chimeras in HEK cells

Physical execution, by hand or by robot.

Genes for these sequences were synthesized and expressed in human embryonic kidney (HEK) cells, and their expression and membrane localization properties were measuredwhere the paper describes this · verbatim
in the paper
3Representation
no AI

Encode sequence and contact-map features

Encoding data into features, descriptors, embeddings or graphs.

The contact-map can be encoded as a binary feature vector xst that indicates the presence or absence of each possible contacting pair.where the paper describes this · verbatim
in the paper
4Training
AI

Train GP classification and regression models

Fitting model parameters, including fine-tuning an existing model.

The training set data (S1 Fig) were used to build a GP classification modelwhere the paper describes this · verbatim
in the paper
5Inference
AI

Predict across the library and select exploration and verification sets

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

We therefore used the localization classification model to identify multi-block-swap chimeras from the librarywhere the paper describes this · verbatim
in the paper
6Screening
AI

Select optimal chimeras and CsCbChR1 block swaps by confidence bounds

Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.

We used the localization regression model to predict ChR chimeras with optimal localization using the Lower Confidence Bound (LCB) algorithmwhere the paper describes this · verbatim
in the paper
7Validation
no AI

Synthesize and measure model-selected ChR variants

Testing outputs against ground truth. Its result feeds back into an earlier step.

These were constructed and testedwhere the paper describes this · verbatim
in the paper
8Interpretation
AI

Weight sequence and contact features for interpretation

Extracting understanding from model behaviour. The AI stood in for expert judgement.

L2-regularized linear regression was used to calculate the positive and negative feature weightswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's claims are the predictions and designs produced by the Gaussian process models; the selected exploration, verification, optimal and CsCbChR1 variants were all chosen by the models.

+What the AI was for
We applied Gaussian process (GP) classification and regressionwhere the paper describes this · verbatim
+Model families
+How it was taught
SupervisedActive learningin the paper
+Models named
Gaussian process binary classification model (expression, localization, localization efficiency) · Trained from scratchGaussian process regression model (localization) · Trained from scratchL1-regularized linear regression (feature selection) · Trained from scratchBayesian ridge linear regression (scikit-learn) · Trained from scratchin the paper
+How results were checked
Experimental11 tested, 11 workedin the paper
The model perfectly classifies the eleven chimeras as either 'high' or 'low' for each propertywhere the paper describes this · verbatim
−Code · weights · data
code not reportedweights not reporteddata not reportednot reported
−Compute
not reportednot reported

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 9 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of Gaussian process binary classification model (expression, localization, localization efficiency)Which version of the model was used is not stated.
  • Version of Gaussian process regression model (localization)Which version of the model was used is not stated.
  • Version of L1-regularized linear regression (feature selection)Which version of the model was used is not stated.
  • Version of Bayesian ridge linear regression (scikit-learn)Which version of the model was used is not stated.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00009, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error