structural-biology/ai produced the result/Bioinformatics Advances 2024 · v2
Software splits coiled-coil proteins into short windows for AlphaFold to model
Researchers built a Python tool, CCfrag, that cuts a protein sequence into overlapping pieces and has AlphaFold predict each one. The program then merges the predictions into a position-by-position picture of the whole chain.
spectrum · one line per step, placed by what the step does · bright lines used AI
CCfrag: scanning folding potential of coiled-coil fragments with AlphaFold
Bioinformatics Advances, 2024
doi:10.1093/bioadv/vbae195 · record aix-00162 v2 · checked 2026-10-09
- AI was for
- Structure determination, Classification
- Model family
- Transformer, Protein language model
- Checked by
- Replication
- Code
- available
The finding the paper is about came from the AI.
What this research was about
Many proteins contain long rope-like stretches called coiled coils, in which two or more helices wind around each other. Helices are the corkscrew segments of a protein chain; in a coiled coil they pack together by slotting bumps on one helix into gaps on its neighbour, an arrangement known as knobs-into-holes. These structures can run for hundreds of residues, the individual building blocks of a protein chain, and they can be parallel, with the chains running the same way, or antiparallel, with one reversed. Working out where such a stretch starts and stops, how many chains it involves and in which direction they run is awkward, because the behaviour is spread over a long, repetitive sequence.
The authors set out to approach this piece by piece rather than all at once. Instead of asking a structure predictor to handle an entire fibrous protein, they wanted to scan along the sequence in short overlapping windows and record what each window looks like on its own. They applied this to long fibrous and non-canonical coiled coils, including EEA1, myosin, a MACH-family protein, tetrabrachion and the SARS coronavirus spike protein.
Where AI came in
AlphaFold, a program that predicts a protein's three-dimensional shape from its sequence, supplies every structural observation here. It was used as released, through ColabFold version 1.5.2 with the alphafold2_multimer_v3 model, default sequence-alignment generation and no templates. CCfrag itself does no predicting. A divider module chops the sequence according to a chosen window length, overlap and number of chains, and writes the query files; the predictions run separately. An integrator module then reads the resulting models and logs, for each position, the per-residue confidence score pLDDT, the averaged error estimate PAE, whether the helices run parallel or antiparallel, and whether the SOCKET program finds knobs-into-holes packing, averaging where windows overlap.
ESMFold is supported as an alternative predictor. For comparison the authors also ran a full-length AlphaFold model of the same protein and two sequence-based coiled-coil predictors, COILS and DeepCoil 2, and superposed selected fragment models on published crystal structures. For EEA1 the fragment models scored better on pLDDT than the full-length prediction, with 70-residue windows giving the best scores, while shorter windows often failed to encode a parallel orientation. In the spike protein, fragment 925-975 came out as a parallel trimer 0.4 Angstroms from the solved structure 1ZVB.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
CCfrag is a Python module that splits a protein sequence into overlapping windows of user-chosen size, overlap and oligomeric state, writes the FASTA files needed to run AlphaFold on each fragment, and then merges the resulting models into a per-residue table of pLDDT, mean PAE, parallel or antiparallel orientation and SOCKET-detected knobs-into-holes interactions. The authors applied it to long fibrous and non-canonical coiled coils, including EEA1, myosin, a MACH-family protein, tetrabrachion and the SARS coronavirus spike protein. For EEA1 the fragment models scored better on pLDDT than the full-length prediction, with the 70-residue window giving the best scores, while shorter windows often did not encode a parallel orientation. For the spike protein, modelled as dimers, trimers and tetramers in 50-residue windows with 25-residue overlap, fragment 925-975 was predicted as a parallel trimer with 0.4 Angstroms RMSD to the solved structure 1ZVB, and fragment 1150-1200 as an antiparallel tetramer with a register differing from PDB 1ZV7.
How AI was used
AlphaFold was used as released, through ColabFold 1.5.2 with the alphafold2_multimer_v3 model, a maximum of five recycling rounds, default MSA generation and no templates, to predict the structure of short overlapping fragments of coiled-coil sequences rather than the full-length chain. A divider module partitions the input sequence according to a specification of window length, overlap length and oligomeric state, optionally attaching a flanking sequence to each window, and emits FASTA queries; the predictions themselves are run outside the program. An integrator module then reads the fragment models and records, at each position of the full-length sequence, the per-residue pLDDT, the averaged PAE, whether the helices are parallel or antiparallel (from inter-chain C-alpha distances compared with the reversed residue index) and whether SOCKET finds knobs-into-holes packing, averaging the values where windows overlap. ESMFold is also supported as an alternative predictor. A full-length AlphaFold model of the same protein, and the sequence-based coiled-coil predictors COILS and DeepCoil 2, were run for comparison, and selected fragment models were superposed on published crystal structures.
The shape of the work
Structural · the record, drawn
no AI
Assemble target protein sequences
Obtaining raw data, whether by measurement, download or retrieval.
A summary of the set of proteins used in the examples is provided in Supplementary Table S1.where the paper describes this · verbatim
AI
Predict full-length models as reference
Running a trained model over new data to predict, classify or score.
A full-length model of Homo sapiens EEA1 is shown, colored by pLDDTwhere the paper describes this · verbatim
no AI
Divide sequence into overlapping windows
Cleaning, filtering, normalising or labelling data already obtained.
The divider is used to partition a full-length sequence according to a user-defined specificationwhere the paper describes this · verbatim
AI
Predict structure of each fragment
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
The examples shown in this article were run in ColabFold version 1.5.2, using the alphafold2_multimer_v3 modelwhere the paper describes this · verbatim
no AI
Integrate fragment features per residue
Encoding data into features, descriptors, embeddings or graphs.
the integrator module extracts a number of features from the models, and incorporates them into a rich per-residue representationwhere the paper describes this · verbatim
AI
Run sequence-based coiled-coil predictors
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
coiled-coil prediction probabilities for COILS (window size = 21) and DeepCoil 2 are shownwhere the paper describes this · verbatim
no AI
Compare models to full-length and solved structures
Testing outputs against ground truth.
Fragment 925–975 is predicted as a parallel trimer, nearly identical to the experimentally solved structure (1ZVB, in grey; RMSD: 0.4 Angstroms)where the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
Every reported observation is a property of AlphaFold models of sequence fragments; the per-residue representation the paper presents is built entirely from model outputs (pLDDT, PAE, orientation, knobs-into-holes detected on predicted coordinates)
The examples shown in this article were run in ColabFold version 1.5.2, using the alphafold2_multimer_v3 modelwhere the paper describes this · verbatim
the predicted fragments confirm previous analyses of the sequence, in terms of the annotation of different coiled-coil periodicitieswhere the paper describes this · verbatim
CCfrag, together with its documentation and additional examples, is available at https://github.com/Mikel-MG/CCfrag.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- Trained model weightsWhether the trained model is available is not stated.
- DataWhether the data are available is not stated.
- ComputeThe hardware or time used is not stated.
- How many were testedThe paper gives no count of what was tested.
- Version of ESMFoldWhich version of the model was used is not stated.
- What step 2 replacedThe paper gives no basis for what the AI stood in for.
About this article
Record aix-00162, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error