structural-biology/ai produced the result/Bioinformatics 2025 · v2
AI-labelled protein pairs used to test how well function predictors handle unknown proteins
Researchers assembled thousands of microbial proteins with no close known relatives, then used a neural network and structure comparison to mark which pairs probably share a job. Thirteen annotation tools were scored against those labels.
spectrum · one line per step, placed by what the step does · bright lines used AI
Functional profiling of the sequence stockpile: a protein pair-based assessment of in silico prediction tools
Bioinformatics, 2025
doi:10.1093/bioinformatics/btaf035 · record aix-00138 v2 · checked 2026-10-09
- AI was for
- Classification
- Model family
- Recurrent neural network, Transformer, Protein language model, Convolutional neural network
- Checked by
- Held-out1927 tested, 1693 worked
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Proteins are chains of amino acids that fold into shapes and do jobs in a cell. Databases now hold enormous numbers of protein sequences read straight from environmental DNA, from soil, seawater or the gut, without anyone ever studying the organism they came from. For most of these, nobody knows what the protein does. The usual way to guess is to find a similar sequence that has been studied in the laboratory and borrow its description. That trick fails for so-called orphan proteins, which have no close match anywhere. Testing whether prediction software works on such proteins is awkward, because there is no laboratory answer to compare against.
The researchers gathered 11 444 orphan proteins with less than 30% sequence similarity to a large reference database, each with a computationally predicted three-dimensional structure. Rather than asking what each protein does, they asked a narrower question: do two orphan proteins share a function? They built a set of pairs judged likely to share one, then checked whether existing annotation tools gave both members of a pair matching descriptions.
Where AI came in
Because no experimental labels existed, the labels themselves came from software. A previously trained Siamese neural network, a model that compares two inputs and reports how alike they are, scored functional similarity for each pair of orphan proteins. This was combined with structural alignment, which measures how closely two folded shapes superimpose. Of 309 549 pairs that could be fully aligned, 6219 passed both cut-offs and were labelled as sharing a function. The network stood in for the manual curation that would otherwise assign function by hand.
The same procedure was then run on enzymes that do carry experimentally evidenced function codes, to estimate how often it was right. The tools under assessment were also largely learned models, including protein language models and convolutional networks, run off the shelf. The reported comparison therefore rests on model output rather than on independent laboratory ground truth.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
The study assembled 11 444 metagenome-derived 'orphan' proteins with under 30% sequence identity to UniRef100 and, lacking ground-truth function labels, used a Siamese neural network score together with Foldseek structural alignment to label 6219 of 309 549 structurally alignable protein pairs as likely sharing molecular function. Thirteen function annotation tools, several protein and DNA embedding models, and the ortholog finder SwiftOrtho were then scored on whether they assigned matching annotations to both proteins of a labelled pair. The paper reports ECPred, GOPredSim, Pfam HMMER and GhostKOALA as the top performers on the ΔS and F1max metrics, with all methods scoring closer to the random baselines than to the estimated performance of experimental annotation, and reports that embedding distances alone did no better than random on the ΔS and RecallmaxPPf50 metrics. Method predictions were enriched for sequence-similar pairs, which the authors read as a dependence on homology.
How AI was used
Protein sequences and pre-computed ESMFold structures were taken from the ESM Metagenomic Atlas, filtered by mmseqs2 alignment against UniRef100 and by length and CD-HIT redundancy reduction to an orphan set. Foldseek aligned all orphan structures against each other, and a previously trained Siamese neural network — a pretrained LookingGlass embedding layer, an LSTM layer and an embedding-distance computation — scored functional similarity for every alignable pair; pairs above fixed TM-score and SNN-score cut-offs were labelled siblings, giving a positive-versus-unlabelled test set. The same SNN plus TM procedure was applied to enzymes carrying single experimentally evidenced EC numbers to estimate its precision and recall and an ideal-predictor reference. Thirteen annotation tools predicting GO molecular function terms, EC numbers, Pfam domains or ortholog groups, plus seven sequence embedding models and SwiftOrtho, were run over the orphan proteins; their per-protein outputs were converted into per-pair similarity using information-accretion-weighted Jaccard similarity for GO, plain Jaccard for other vocabularies, and Euclidean or cosine distance for embeddings. Methods were scored by sweeping prediction-score and similarity-score thresholds, with unlabelled pairs under-sampled 100 times, against random-classifier and random-annotator baselines and Wilcoxon and t-test significance testing.
The shape of the work
Structural · the record, drawn
no AI
Collect metagenome proteins with confident predicted structures
Obtaining raw data, whether by measurement, download or retrieval.
we collected 53 501 759 protein sequences, translated from metagenome-assembled genes, and having high-confidence predicted 3D structureswhere the paper describes this · verbatim
no AI
Filter to orphan proteins
Cleaning, filtering, normalising or labelling data already obtained.
The final dataset of orphan proteins contained 11 444 sequences with ESM predicted structures and corresponding MGnify cDNA sequenceswhere the paper describes this · verbatim
no AI
All-against-all structural alignment of orphans
Reducing a candidate set by filtering or ranking, in a single pass.
Only 309 549 of these protein pairs (0.5% of ∼65M possible ones) were structurally similar enough for a complete alignment.where the paper describes this · verbatim
AI
Score functional similarity of pairs with the SNN
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
We then annotated functional similarity (SNN) scores for these 309K pairswhere the paper describes this · verbatim
no AI
Threshold scores to label orphan sibling pairs
Reducing a candidate set by filtering or ranking, in a single pass.
Only 6219 (2% of 309K) pairs attained the pre-set cutoffs (TM score ≥ 0.7, SNN score ≥ 0.98) for shared functionwhere the paper describes this · verbatim
AI
Check sibling labelling against experimental EC annotations
Testing outputs against ground truth. The AI stood in for manual curation.
We identified siblings in this set of enzymes and compared sibling annotations to EC pairings.where the paper describes this · verbatim
AI
Run annotation tools and embedding models over the orphans
Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.
We selected 13 protein annotation tools for our assessment based on the availability of a standalone version or a web serverwhere the paper describes this · verbatim
no AI
Convert predictions to pair similarity and score methods
Testing outputs against ground truth.
The performance of a given method in identifying orphan siblings was measured first by computing the Area under the Precision–Recall curvewhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The assessment's positive labels are produced by a trained Siamese neural network combined with structural alignment, and the methods being assessed are themselves largely learned models, so the reported findings rest on model output rather than on independent ground truth
ESM-2 and ProtT5 are transformer-based protein language models whereas LookingGlass is a bi-directional LSTM modelwhere the paper describes this · verbatim
our approach eliminated any overlap between the training data of the prediction methods and our test setwhere the paper describes this · verbatim
The code used to compute siblings is available openly at https://bitbucket.org/bromberglab/siblings-detector/.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of SNN (Siamese Neural Network)Which version of the model was used is not stated.
- Version of LookingGlassWhich version of the model was used is not stated.
- Version of ESMFold / ESM Metagenomic Atlas structuresWhich version of the model was used is not stated.
- Version of ESM-2Which version of the model was used is not stated.
- Version of ProtTrans (ProtT5)Which version of the model was used is not stated.
- Version of SeqVecWhich version of the model was used is not stated.
- Version of BeplerWhich version of the model was used is not stated.
- Version of CPCProtWhich version of the model was used is not stated.
- Version of Word2VecWhich version of the model was used is not stated.
- Version of DeepFriWhich version of the model was used is not stated.
- Version of DeepGOPlusWhich version of the model was used is not stated.
- Version of GoPredSimWhich version of the model was used is not stated.
- Version of GOProFormerWhich version of the model was used is not stated.
- Version of NetGOWhich version of the model was used is not stated.
- Version of ProtENNWhich version of the model was used is not stated.
- Version of ProtCNNWhich version of the model was used is not stated.
- Version of ProteInferWhich version of the model was used is not stated.
- Version of ECPredWhich version of the model was used is not stated.
- Version of MantisWhich version of the model was used is not stated.
About this article
Record aix-00138, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error