structural-biology/ai produced the result/PLoS Computational Biology 2017 · v2
Gaussian process models pick light-sensitive ion channels that reach mammalian cell membranes
Researchers measured how 218 stitched-together channelrhodopsin proteins behaved in human cells, then trained Gaussian process models on those measurements to predict, out of 118,098 possible variants, which untested ones would reach the cell membrane and which to build next.
spectrum · one line per step, placed by what the step does · bright lines used AI
Machine learning to design integral membrane channelrhodopsins for efficient eukaryotic expression and plasma membrane localization
PLoS Computational Biology, 2017
doi:10.1371/journal.pcbi.1005786 · record aix-00009 v2 · checked 2026-10-07
- AI was for
- Property prediction, Classification, Experimental design
- Model family
- Gaussian process, Linear model
- Checked by
- Experimental11 tested, 11 worked
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about

Channelrhodopsins are proteins from algae that sit in a cell's outer membrane and open to let ions through when light hits them. Biologists put them into other cells so that light can switch those cells on or off. For that to work, the protein must first be made by the host cell and then travel to the outer membrane, called the plasma membrane. Many natural channelrhodopsins do neither well in mammalian cells. One way to find better ones is to shuffle pieces of several parent proteins into hybrids, known as chimeras. The trouble is arithmetic: the shuffling scheme here allows 118,098 combinations, far more than anyone can build and test.
The researchers took three parent channelrhodopsins, used a crystal structure of a related protein to decide where the sequences could be cut and rejoined, and built a library of possible hybrids on paper. They then made 218 of them, expressed them in human embryonic kidney cells, and measured two things with fluorescent tags: how much protein was made, and how much of it reached the plasma membrane.
Where AI came in
Those 218 measurements became training data. The team fitted Gaussian process models, a statistical method that predicts a value for an untried case and also reports how uncertain that prediction is. Proteins were described to the models as lists of which amino acids were present and which pairs of residues touch each other in the known structure. Some models sorted variants into 'high' or 'low' for expression and membrane localisation; another predicted the localisation level as a number. Simpler linear regressions picked out which sequence and contact features mattered, and later assigned weights that were drawn onto the structure.
The models then stood in for the screening the researchers could not do by hand. Applied across the whole library, they chose a set of varied hybrids and a separate set to check their own accuracy; the model sorted all eleven of the latter correctly. Using its uncertainty estimates, the regression model nominated four hybrids in the top 0.1% of its predictions, and nominated single-block swaps into a natural channelrhodopsin. New measurements were fed back to retrain the models. The variants tested were the ones the models selected.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record

The authors measured expression and plasma-membrane localization in HEK cells for 218 channelrhodopsin chimeras drawn from a 118,098-variant library designed by SCHEMA recombination of three parent channelrhodopsins. Gaussian process classification and regression models were trained on those measurements using kernels built from sequence and residue-residue contact features, then used to predict which untested chimeras express and localize and to pick further variants for synthesis. Four chimeras whose lower-confidence-bound predictions ranked in the top 0.1% of the library were built and all localized as well as or better than the best-localizing parent, CsChrimR, and model-selected single-block swaps raised the localization of a natural channelrhodopsin, CbChR1, that does not localize in mammalian cells. L1- and L2-regularized linear regression on the same features was used to weight the residues and contacts associated with localization.
How AI was used
Expression and localization data from 218 SCHEMA-recombination chimeras were used to train Gaussian process binary classification models (Laplace approximation for the intractable posterior) for 'high' versus 'low' expression, localization and localization efficiency, and a Gaussian process regression model for localization level. Sequences were compared through kernels over binary sequence feature vectors and residue-residue contact-map feature vectors derived from the C1C2 crystal structure; linear, squared exponential and Matern kernels were compared, and kernel form and hyperparameters were set by maximizing the marginal likelihood. For regression, L1-regularized linear regression first selected a subset of sequence and contact features, and the GP was trained on the non-zero-weight features; regularization strength and model performance were assessed by leave-one-out cross-validation. The classification model was applied across the full library to choose a diverse exploration set and a verification set for synthesis, and those measurements were added to retrain the models. The regression model was then applied with a lower confidence bound acquisition to pick library chimeras for synthesis and with an upper confidence bound acquisition to pick single-block swaps into the natural variant CsCbChR1. Separately, Bayesian ridge regression on the L1-selected features produced feature weights that were mapped onto the C1C2 structure. Modelling used open-source SciPy-ecosystem packages and scikit-learn.
The shape of the work
Structural · the record, drawn
no AI
Design SCHEMA recombination library
Producing candidate objects that did not previously exist.
These designs generate 118,098 possible chimeraswhere the paper describes this · verbatim
no AI
Synthesize and assay training-set chimeras in HEK cells
Physical execution, by hand or by robot.
Genes for these sequences were synthesized and expressed in human embryonic kidney (HEK) cells, and their expression and membrane localization properties were measuredwhere the paper describes this · verbatim
no AI
Encode sequence and contact-map features
Encoding data into features, descriptors, embeddings or graphs.
The contact-map can be encoded as a binary feature vector xst that indicates the presence or absence of each possible contacting pair.where the paper describes this · verbatim
AI
Train GP classification and regression models
Fitting model parameters, including fine-tuning an existing model.
The training set data (S1 Fig) were used to build a GP classification modelwhere the paper describes this · verbatim
AI
Predict across the library and select exploration and verification sets
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
We therefore used the localization classification model to identify multi-block-swap chimeras from the librarywhere the paper describes this · verbatim
AI
Select optimal chimeras and CsCbChR1 block swaps by confidence bounds
Reducing a candidate set by filtering or ranking, in a single pass. The AI stood in for physical experiment.
We used the localization regression model to predict ChR chimeras with optimal localization using the Lower Confidence Bound (LCB) algorithmwhere the paper describes this · verbatim
no AI
Synthesize and measure model-selected ChR variants
Testing outputs against ground truth. Its result feeds back into an earlier step.
These were constructed and testedwhere the paper describes this · verbatim
AI
Weight sequence and contact features for interpretation
Extracting understanding from model behaviour. The AI stood in for expert judgement.
L2-regularized linear regression was used to calculate the positive and negative feature weightswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's claims are the predictions and designs produced by the Gaussian process models; the selected exploration, verification, optimal and CsCbChR1 variants were all chosen by the models.
We applied Gaussian process (GP) classification and regressionwhere the paper describes this · verbatim
The model perfectly classifies the eleven chimeras as either 'high' or 'low' for each propertywhere the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of Gaussian process binary classification model (expression, localization, localization efficiency)Which version of the model was used is not stated.
- Version of Gaussian process regression model (localization)Which version of the model was used is not stated.
- Version of L1-regularized linear regression (feature selection)Which version of the model was used is not stated.
- Version of Bayesian ridge linear regression (scikit-learn)Which version of the model was used is not stated.
- What step 4 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00009, version 2, checked by a person on 2026-10-07. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY-4.0; quotations are at most 25 words. How we work · Report an error