~/aixsci
200 records · all checked

structural-biology/ai produced the result/Knowledge-Based Systems 2025 · v2

A transformer model segments cryo-electron tomograms from three viewing directions at once

Researchers built MVGFormer, a neural network that labels the contents of three-dimensional cryo-electron tomography volumes. The model does the labelling itself, reading each volume from three perpendicular views and picking out particles that people would otherwise mark by hand.

1. Assemble simulated and real cryo-ET datasets2. Patch, resize and binarise volumes3. Encode multi-view token sequences4. Build context visual graph for attention guidance5. Train encoder and decoders with view-masked self-supervision6. Segment volumes and pick particles7. Score against ground truth and baselines

spectrum · one line per step, placed by what the step does · bright lines used AI

MVGFormer: Multi-view perspective with graph-guided transformer for cryo-ET segmentation
Knowledge-Based Systems, 2025

doi:10.1016/j.knosys.2025.114810 · record aix-00190 v2 · checked 2026-10-09

ai-resultrole of AI
AI was for
Segmentation, Detection
Model family
Transformer, Convolutional neural network, Clustering
Checked by
Benchmark4096 tested
Code
not reported

The finding the paper is about came from the AI.

read as

The science is explained before the AI appears. Switch to field specialist to go straight to the method.

Assumes the discipline and goes straight to the method.

The diagram, the record and what the paper did not report are identical in both modes. Only the framing changes — never the evidence.

Introduction by AIxSci · plain language

What this research was about

Cryo-electron tomography is a way of photographing frozen biological samples from many angles and combining the images into a three-dimensional map. The maps are murky. Because the sample cannot be tilted all the way round, a slice of angles is always missing, leaving smeared and stretched features. The microscope's optics also distort contrast. So the job of saying which voxels, the three-dimensional equivalent of pixels, belong to a protein and which are noise is slow and uncertain work. Researchers also want to find the individual copies of a molecule scattered through the volume, a step known as particle picking.

The authors set out to automate that labelling. They assembled simulated volumes and several real datasets, cut them into small cubes, and built a model to assign a class to every voxel, then compared it against existing methods for three-dimensional segmentation and particle picking.

Where AI came in

The artificial intelligence is the result here. MVGFormer is a transformer, a network that weighs how parts of its input relate to one another. Each cube is read three times, once from each of the three perpendicular directions, with the network told which view it is looking at. A separate convolutional branch groups the cube's features into sixteen representative nodes by k-means clustering, and those nodes steer the network's attention. Two alternative output modules turn the result into voxel labels, and the per-view predictions are added back together.

Training used labelled masks, plus a self-supervised trick: one of the three views was hidden and the model had to reconstruct it from the other two, so it learned from the data without extra human labels. Some versions were pre-trained on simulated cubes and then fine-tuned on small real datasets. In use, the model stands in for the manual curation of volumes, tiling each tomogram into overlapping subvolumes and averaging its voxel-by-voxel predictions. Scoring against known masks and particle positions was done without AI.

Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.

The work

Technical · from the record

The authors built MVGFormer, a 3D transformer for segmenting cryo-electron tomography volumes. The model reads each volume from the XY, XZ and YZ orthographic views with separate position embeddings, uses a convolutional context encoder whose k-means-derived graph nodes act as attention queries, and decodes voxel-level masks with either a multi-level feature fusion decoder or a parallel 3D atrous convolution decoder, with a view-masked reconstruction objective added during training. It was evaluated on six cryo-ET datasets across tomogram segmentation, subtomogram segmentation and particle picking, and compared against CNN- and transformer-based 3D segmentation baselines on mIoU, Dice, precision, recall and F1.

How AI was used

Tomograms were cut into non-overlapping voxel patches (size 32) and subtomograms were resized, with grey-scale simulated masks thresholded to binary. Each patch was transposed into three orthographic views, divided into 4x4x4 patches, linearly projected to 256-dimensional tokens and given view-specific learnable position embeddings before entering a 12-layer multi-head self-attention encoder; a parallel convolutional context encoder clustered its feature map into 16 graph nodes that served as attention queries. Two decoders were trained: a multi-level feature fusion segmentor aggregating features from several encoder layers, and a parallel 3D atrous convolution segmentor with dilation rates 1, 6, 12 and 18. Per-view predicted masks were re-aligned to the canonical grid and summed. Training used cross-entropy segmentation losses on per-view and fused masks plus a mean-squared-error reconstruction loss for a view-masked self-supervised objective at a 50% mask rate, with Adam at learning rate 1e-3 for 200 epochs at batch size 72. One configuration was pre-trained for 100 epochs on the simulated subtomogram dataset and then fine-tuned on the tomogram dataset or on small real subtomogram datasets. At inference, tomograms were tiled into subvolumes with 50% overlap and per-voxel class probabilities were fused by weighted averaging with a smooth window; for particle picking, predicted centres were matched to known centres within the particle radius. Baseline CNN and transformer segmentation models were trained by the authors on the same data for comparison.

The shape of the work

Structural · the record, drawn

ACQUISITIONPREPARATIONREPRESENTATIONREPRESENTATIONTRAININGINFERENCEVALIDATION1234567AIAIAIAIAssemblesimulated andreal cryo-ET dat…Patch, resize andbinarise volumesEncode multi-viewtoken sequencesBuild contextvisual graph forattention guidan…Train encoder anddecoders withview-masked self…Segment volumesand pickparticlesScore againstground truth andbaselines↤ manual curation
AI stepNo AI↤ what the AI stood in for
1Acquisition
no AI

Assemble simulated and real cryo-ET datasets

Obtaining raw data, whether by measurement, download or retrieval.

We chose the tomogram dataset used in SHREC2021 as the tomogram dataset.where the paper describes this · verbatim
in the paper
2Preparation
no AI

Patch, resize and binarise volumes

Cleaning, filtering, normalising or labelling data already obtained.

we cut each tomogram into multiple non-overlapping patches with size 32, and use each patch as the input for the modelwhere the paper describes this · verbatim
in the paper
3Representation
AI

Encode multi-view token sequences

Encoding data into features, descriptors, embeddings or graphs.

we perform high-dimensional transpose to obtain the transposed inputs from ‘YZ’ viewwhere the paper describes this · verbatim
in the paper
4Representation
AI

Build context visual graph for attention guidance

Encoding data into features, descriptors, embeddings or graphs.

we construct a visual graph via k-means clustering on fc, aiming to select more informative and representative graph nodeswhere the paper describes this · verbatim
in the paper
5Training
AI

Train encoder and decoders with view-masked self-supervision

Fitting model parameters, including fine-tuning an existing model.

we randomly select one input view as the masked view, and reconstruct the masked view from the remaining two viewswhere the paper describes this · verbatim
in the paper
6Inference
AI

Segment volumes and pick particles

Running a trained model over new data to predict, classify or score. The AI stood in for manual curation.

during inference the tomogram is divided into 32 subvolumes with a 50% overlap in each dimensionwhere the paper describes this · verbatim
in the paper
7Validation
no AI

Score against ground truth and baselines

Testing outputs against ground truth.

we choose the mean intersection of union (mIoU) and dice similarity coefficient (Dice) as the core evaluation metricswhere the paper describes this · verbatim
in the paper

What the record says

Technical · every part carries its own basis

+ in the paper~ our reading− not reported

How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.

~Role of AI
AI produced the resultour reading

The paper's result is the trained segmentation model itself; every reported finding is a model output on cryo-ET tomograms and subtomograms

+What the AI was for
propose a novel transformer-based framework for cryo-ET segmentation, named MVGFormerwhere the paper describes this · verbatim
+How it was taught
SupervisedSelf-supervisedTransfer / fine-tuningZero-shotin the paper
+Models named
MVGFormer (MF decoder) · Trained from scratchMVGFormer (P3DA decoder) · Trained from scratchVoxResNet · Trained from scratchMedNeXt · Trained from scratchSwin UNETR · Trained from scratchSwiFT · Trained from scratchDeepFinder · Trained from scratchcrYOLO · Trained from scratchEMAN2 · Trained from scratchin the paper
+How results were checked
Benchmark4096 testedin the paper
there are total 40,960 samples in the SHREC dataset (36,864 samples in training set and 4096 samples in test set)where the paper describes this · verbatim
+Code · weights · data
code not reportedweights not reporteddata not reportedin the paper
We train our model on two NVIDIA A100 Tensor Core GPUs with a 80GB memory per card.where the paper describes this · verbatim
+Compute
Two NVIDIA A100 Tensor Core GPUs with 80GB memory per card; 200 training epochs at batch size 72, and 100 pre-training epochs for the subtomogram pre-training experimentin the paper

What this paper did not report

Technical · absence is published deliberately

Reported as not stated — 15 items
  • CodeWhether the code is available is not stated.
  • Trained model weightsWhether the trained model is available is not stated.
  • DataWhether the data are available is not stated.
  • Version of MVGFormer (MF decoder)Which version of the model was used is not stated.
  • Version of MVGFormer (P3DA decoder)Which version of the model was used is not stated.
  • Version of VoxResNetWhich version of the model was used is not stated.
  • Version of MedNeXtWhich version of the model was used is not stated.
  • Version of Swin UNETRWhich version of the model was used is not stated.
  • Version of SwiFTWhich version of the model was used is not stated.
  • Version of DeepFinderWhich version of the model was used is not stated.
  • Version of crYOLOWhich version of the model was used is not stated.
  • Version of EMAN2Which version of the model was used is not stated.
  • What step 3 replacedThe paper gives no basis for what the AI stood in for.
  • What step 4 replacedThe paper gives no basis for what the AI stood in for.
  • What step 5 replacedThe paper gives no basis for what the AI stood in for.

About this article

Record aix-00190, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error