structural-biology/ai produced the result/Journal of Environmental and Public Health 2022 · v2
Machine learning sorts mitochondrial proteins into three compartments from sequence alone
Researchers built a classifier that guesses which part of the mitochondrion a protein sits in, using only its amino acid sequence. A convolutional neural network did the sorting, with gradient boosting trimming the input features.
spectrum · one line per step, placed by what the step does · bright lines used AI
[Retracted] Construction of Mitochondrial Protection and Monitoring Model of Lon Protease Based on Machine Learning under Myocardial Ischemia Environment
Journal of Environmental and Public Health, 2022
doi:10.1155/2022/4805009 · record aix-00121 v2 · checked 2026-10-08
- AI was for
- Classification
- Model family
- Convolutional neural network, Support vector machine, Gradient-boosted trees, Random forest, Clustering
- Checked by
- Held-out
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Mitochondria are the compartments inside our cells that release energy from food. They are not simple sacs. A mitochondrion has an outer membrane, an inner membrane folded into pleats, and a watery interior called the matrix. Which of these three places a protein ends up in largely decides what it can do, so knowing the location is part of knowing the protein. Finding out by experiment means isolating the compartments and detecting the protein in each, which is slow work. The alternative is to predict the location from the protein's sequence, the string of amino acids written in its gene. The difficulty is that the clues in that string are subtle and spread out.
The researchers set out to build such a predictor and test it on standard collections of mitochondrial protein sequences, each already labelled as inner membrane, matrix or outer membrane. Their stated framing was the protein Lon protease and the stressed heart muscle, but the result they report is the accuracy of the classifier itself.
Where AI came in
Sequences were first turned into numbers in several ways at once, counting amino acid pairs, scoring positions against related proteins and breaking the signal into layers in the manner of a wavelet analysis. That left far more numbers than were useful, so a gradient boosting method, which builds many small decision trees in sequence, picked out which ones to keep. A convolutional neural network then learned from those features. Each protein was cut into overlapping stretches of fixed length, and each stretch was fed in as a separate channel. An attention step weighted the combined features before the network produced a score for each of the three compartments.
Ninety percent of each collection trained the network and ten percent was held back for testing, with the test split redrawn ten times and the results averaged. On the M495 collection the authors report an overall accuracy of 0.9553, which they state is 0.0913 and 0.0763 above two methods using a single kind of feature. A support vector machine and a gradient boosting classifier were run as comparisons. The network stands in for laboratory separation of the compartments. A separate clustering step was applied to a computer simulation of Lon protease moving, grouping its amino acids by shape and measuring how strongly pairs of them varied together.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study builds a sequence-based classifier that assigns mitochondrial proteins, discussed in the context of Lon protease under myocardial ischemia, to the inner membrane, matrix or outer membrane. Protein sequences were cut into overlapping fixed-length subsequences used as channels of a two-layer convolutional network, with features drawn from position-specific score matrices, pseudoamino acid composition, dipeptide composition and wavelet decomposition, and random oversampling applied to balance the classes. On the M495 data set the authors report an overall accuracy of 0.9553, which they state is 0.0913 and 0.0763 higher than two single-feature methods, and they report a prediction accuracy of 94.28% for the integrated residual CNN on protein-protein interaction prediction. They also report that ROC values rose by 17.6%-21.3% when oversampling was used, and that the fourth wavelet decomposition layer gave an overall accuracy of 93.323%.
How AI was used
Protein sequences from the M317, M495 and M983 benchmark sets were converted into feature vectors combining a position correlation score matrix, pseudoamino acid composition, dipeptide composition, a position-weighted amino acid component and wavelet decomposition coefficients, with self-cross covariance transformation used in the feature extraction. Classes were balanced by random oversampling, and sliding windows centred on arginine and lysine were used to build positive and negative samples. Gradient boosting (described as the limit gradient hoist) was used to select features and drop redundant ones. Each overlapping fixed-length subsequence of a protein served as a channel in the convolution layer of a multichannel two-layer convolutional neural network, whose fused features were weighted by a scaled dot-product attention mechanism, followed by pooling and a fully connected layer producing class scores. Training used PyTorch on an Nvidia GeForce 2080Ti GPU with the Adam optimiser, batch size 36, 200 iterations, a fixed learning rate of 0.002 and dropout of 0.6. Ninety percent of the data was used for training and ten percent for independent testing, with the test set resampled ten times and results averaged. SVM and XGBoost classifiers were run as comparisons. Separately, a machine learning correlation algorithm with residue clustering was applied to a molecular dynamics trajectory of Lon protease, with mutual information computed between residue pairs.
The shape of the work
Structural · the record, drawn
no AI
Assemble benchmark sequence data sets
Obtaining raw data, whether by measurement, download or retrieval.
The data sets used in this study are M317, M495, and M983, and each data set is broken down into three subregionswhere the paper describes this · verbatim
no AI
Build samples, balance classes and split train/test
Cleaning, filtering, normalising or labelling data already obtained.
Random oversampling method is used to process the data set to ensure the balance among all kinds of submitochondrial proteins.where the paper describes this · verbatim
no AI
Encode sequences as features
Encoding data into features, descriptors, embeddings or graphs.
combines position correlation score matrix, pseudoamino acid composition, and dipeptide composition to address the issue of restricted single featurewhere the paper describes this · verbatim
AI
Select features with gradient boosting
Cleaning, filtering, normalising or labelling data already obtained. The AI stood in for expert judgement.
Utilize the limit gradient hoist to screen out crucial characteristics and eliminate superfluous and pointless featureswhere the paper describes this · verbatim
AI
Train multichannel CNN classifier
Fitting model parameters, including fine-tuning an existing model. The AI stood in for conventional algorithm.
In this paper, Pytorch backend is used to conduct all experiments on Nvidia Ge Force 2080TiGPU. Adam algorithm is used to optimize.where the paper describes this · verbatim
AI
Predict submitochondrial localization
Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.
Multichannel CNN is used to extract features from protein sequences and predict the results.where the paper describes this · verbatim
no AI
Evaluate against baseline classifiers
Testing outputs against ground truth.
this approach performs predictions better than SVM and XGBoost classifierswhere the paper describes this · verbatim
AI
Cluster simulation trajectory residues and compute mutual information
Extracting understanding from model behaviour. The AI stood in for expert judgement.
ML correlation algorithm is used to analyze the trajectory of molecular dynamics simulation of Lon proteasewhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's reported result is the accuracy of its own classifier for submitochondrial localization; no non-AI finding is reported
train a multichannel two-layer CNN (convective neural network) to learn advanced features in the sequencewhere the paper describes this · verbatim
we randomly select 90% of the data sets to build training sets and 10% to build independent test setswhere the paper describes this · verbatim
The data used to support the findings of this study are available from the corresponding author upon request.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of Multichannel two-layer CNN (integrated residual CNN)Which version of the model was used is not stated.
- Version of SVMWhich version of the model was used is not stated.
- Version of XGBoostWhich version of the model was used is not stated.
About this article
Record aix-00121, version 2, checked by a person on 2026-10-08. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error