~/aixsci
200 records · all checked

structural-biology/ai produced the result/Database 2023 · v2

AlphaFold models used to map membrane-crossing segments in human proteins

Researchers built a database of the stretches of human proteins that pass through cell membranes. They took predicted three-dimensional structures from AlphaFold and used a physics-based program to settle each model into a membrane.

1. Compile candidate human transmembrane protein set2. Obtain AlphaFold structural models with pLDDT and PAE3. Strip peptides and low-confidence regions from models4. Partition models into ordered and disordered domains5. Position models in membrane with PPM36. Correct and label segments into final AFTM annotations7. Benchmark against experimental and curated annotations8. Add sequence-feature predictions to protein web pages

spectrum · one line per step, placed by what the step does · bright lines used AI

AFTM: a database of transmembrane regions in the human proteome predicted by AlphaFold
Database, 2023

doi:10.1093/database/baad008 · record aix-00154 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Structure determination, Property prediction
Model family
Transformer
Checked by
Benchmark601 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Cells are wrapped in membranes, thin sheets of fat that keep the inside separate from the outside. Many proteins sit in those sheets, with parts of the chain threading right through them. Knowing exactly which stretches of amino acids cross the membrane matters for understanding how a protein works, yet those stretches are hard to pin down. Membrane proteins are notoriously awkward to study by the usual experimental methods, which is why many human entries carry annotations based on indirect reasoning rather than a measured structure. Different databases often disagree about where a crossing starts and ends, or whether one is there at all.

The researchers set out to annotate these segments across the human proteome by working from predicted structures rather than sequence alone. They gathered 5491 candidate human membrane proteins using existing annotations from the UniProt and HTP databases, then worked out, protein by protein, which parts of each chain lie inside the membrane.

Where AI came in

The AI's part was supplying the shapes. AlphaFold predicts a protein's folded three-dimensional structure from its sequence, a job that otherwise needs laboratory structure determination, and the authors used its models as the raw material for everything that followed. AlphaFold also reports how sure it is: a per-residue confidence score and a matrix estimating the error between pairs of positions. Low-confidence positions were stripped out, and the error matrix guided an in-house script that split each model into separate ordered domains.

The membrane placement itself was not done by a learned model. A program called PPM3, which works from energy calculations rather than training data, settled each model into a simulated membrane and reported which segments crossed it; fixed geometric rules then tidied the output. Separately, the website shows off-the-shelf sequence predictions of local structure and of floppy, disordered regions. No model was trained or fine-tuned for this work.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors assembled 5491 candidate human transmembrane proteins from UniProt and HTP annotations, then located transmembrane segments by placing AlphaFold structural models in a membrane with the PPM3 program, after trimming low-confidence and peptide regions and splitting models into domains. Rule-based corrections separated re-entrant regions, merged broken helices and added segments PPM3 had skipped. Against transmembrane regions mapped from PDBTM experimental structures for 601 proteins, 66 PDBTM-defined segments were not predicted by AFTM, compared with 107 for TmAlphaFold, 197 for UniProt and 177 for HTP; on 2322 Membranome single-pass proteins, AFTM reported no segment for 309, compared with 69 for UniProt and 141 for HTP. Results from AFTM, UniProt, HTP, TmAlphaFold, PDBTM and Membranome are published together in an online database.

How AI was used

AlphaFold structural models of human proteins, together with their per-residue pLDDT scores and predicted aligned error matrices, were the substrate for the whole annotation procedure: pLDDT was used to strip low-confidence positions from full-length models, and PAE densities were used by an in-house splitting script to partition models into ordered and disordered domains. Both the trimmed full-length models and the individual domains were then passed to PPM3, a non-learned free-energy-based membrane placement program, whose reported segments were post-processed by fixed geometric rules into transmembrane versus re-entrant regions, with merging of segments of the same orientation separated by fewer than ten residues and insertion of segments PPM3 failed to report. Segment sets were compared with UniProt, HTP, TmAlphaFold, PDBTM-derived and Membranome annotations, using DIAMOND BLAST to map external segments onto human sequences. The web pages additionally display off-the-shelf per-residue predictions of secondary structure (PSIPRED, SPIDER3) and disorder (SPOT-Disorder, IUPRED2A). No model was trained or fine-tuned in this study.

The shape of the work

Structural · the record, drawn

ACQUISITIONINFERENCEPREPARATIONPREPARATIONSIMULATIONPREPARATIONVALIDATIONINFERENCE12345678AIAICompile candidatehumantransmembrane pr…Obtain AlphaFoldstructural modelswith pLDDT and P…Strip peptidesandlow-confidence r…Partition modelsinto ordered anddisordered domai…Position modelsin membrane withPPM3Correct and labelsegments intofinal AFTM annot…Benchmark againstexperimental andcurated annotati…Addsequence-featurepredictions to p…↤ physical experiment
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Compile candidate human transmembrane protein set

Obtaining raw data, whether by measurement, download or retrieval.

This dataset includes 5491 potential human TMPs compiled from a combination of reviewed UniProt entrieswhere the paper describes this · verbatim
in the paper
2Inference
AI

Obtain AlphaFold structural models with pLDDT and PAE

Running a trained model over new data to predict, classify or score. The AI stood in for physical experiment.

In addition to the predicted 3D structure, AlphaFold provides the predicted aligned error (PAE) for each residue pair in a proteinwhere the paper describes this · verbatim
our reading
3Preparation
no AI

Strip peptides and low-confidence regions from models

Cleaning, filtering, normalising or labelling data already obtained.

positions with low pLDDT scores (<0.5) were removed from the full-length AlphaFold modelwhere the paper describes this · verbatim
in the paper
4Preparation
no AI

Partition models into ordered and disordered domains

Cleaning, filtering, normalising or labelling data already obtained.

We wrote an in-house script to iterate the following procedure to split any AlphaFold model into segments (domains)where the paper describes this · verbatim
in the paper
5Simulation
no AI

Position models in membrane with PPM3

Numerical or physics simulation, including where a learned surrogate replaces it.

We used the positioning of proteins in membranes, version 3 (PPM3) method to predict the localizations of TMSswhere the paper describes this · verbatim
in the paper
6Preparation
no AI

Correct and label segments into final AFTM annotations

Cleaning, filtering, normalising or labelling data already obtained.

we merged the consecutive PPM3 segments if they have the same orientation and are separated by <10 residueswhere the paper describes this · verbatim
in the paper
7Validation
no AI

Benchmark against experimental and curated annotations

Testing outputs against ground truth.

we compared their TMS predictions to those derived from experimental structures in the PDBTM databasewhere the paper describes this · verbatim
in the paper
8Inference
AI

Add sequence-feature predictions to protein web pages

Running a trained model over new data to predict, classify or score.

secondary structure predictions by PSI-blast based secondary structure PREDiction (PSIPRED) and SPIDER3where the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

Every transmembrane segment reported by AFTM is derived from AlphaFold structural models; without the predicted structures there is no result, although the segment assignment itself is made by the non-learned PPM3 program

~What the AI was for
The PPM3 program was used to predict the TMSs in AlphaFold models of human proteinswhere the paper describes this · verbatim
~Model families
Transformerour reading
~How it was taught
Supervisedour reading
~Models named
AlphaFold · Off the shelfPSIPRED · Off the shelfSPIDER3 · Off the shelfSPOT-Disorder · Off the shelfour reading
+How results were checked
Benchmark601 testedin the paper
A total of 601 human proteins in our dataset were mapped to at least one entrywhere the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata availablein the paper
The AFTM database is available at http://conglab.swmed.edu/AFTMwhere the paper describes this · verbatim
+Compute
not reportedin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 8 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • ComputeThe hardware or time used is not stated.
  • Version of AlphaFoldWhich version of the model was used is not stated.
  • Version of PSIPREDWhich version of the model was used is not stated.
  • Version of SPIDER3Which version of the model was used is not stated.
  • Version of SPOT-DisorderWhich version of the model was used is not stated.
  • What step 8 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00154, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error