structural-biology/ai produced the result/Frontiers in Molecular Biosciences 2024 · v2
Study tests whether AI clean-up of cryo-EM maps harms ligand density
Researchers ran three released machine-learning tools for tidying up cryo-electron microscopy density maps across a panel of deposited structures, then compared the results with conventional methods. The AI tools were the objects under test rather than the researchers' own creations.
spectrum · one line per step, placed by what the step does · bright lines used AI
Machine learning approaches to cryoEM density modification differentially affect biomacromolecule and ligand density quality
Frontiers in Molecular Biosciences, 2024
doi:10.3389/fmolb.2024.1404885 · record aix-00192 v2 · checked 2026-10-09
- AI was for
- Denoising
- Model family
- Convolutional neural network, Generative adversarial network
- Checked by
- Held-out7 tested
- Code
- not reported
The finding the paper is about came from the AI.
What this research was about
Cryo-electron microscopy, or cryo-EM, works out the shape of molecules by freezing many copies of them and imaging them with electrons. The result is not a picture of atoms but a three-dimensional map of density: a cloud showing where matter is more or less likely to be. Scientists then fit a model of the molecule into that cloud. The maps are noisy, so they are usually sharpened or otherwise processed first. The difficulty is that processing which makes the main body of a protein easier to read may not treat everything else equally. Small attachments matter too: drug molecules, metal ions, bound oxygen, fats and stretches of DNA or RNA.
Machine-learning tools now exist to do this clean-up step, learning from training data what good density tends to look like and reshaping a map accordingly. The researchers set out to check how such tools treat the non-protein parts of a map. They gathered high-resolution maps from the public Electron Microscopy Data Bank, chosen for quality and for containing cofactors, ligands and ions, plus one unpublished map of haemoglobin bound to oxygen at 2.4 ångströms.
Where AI came in
The machine learning here was not built by the authors. Three released tools, DeepEMHancer, EMReady and EM-GAN, were taken off the shelf and run on each map. They work directly on the volume, adjusting density point by point using features learned during training. Nothing was trained or fine-tuned in the study. For comparison, the same maps also went through plain B-factor sharpening and through PHENIX RESOLVE, a map-modification algorithm that does not learn from data.
Atomic models were then refined into every version of every map with PHENIX real-space refinement, so each could be judged on the same footing, and Q-scores were calculated separately for backbone, sidechain, nucleic acid, ligand and water. Individual ligand, metal, lipid and nucleic acid sites were also compared by eye. The learned tools stand in for the hand-designed sharpening and error-correction steps a researcher would otherwise apply.
Written by AIxSci from the checked record below, to give context for readers outside the field. It is not part of the record.
The work
Technical · from the record
Three machine-learning cryoEM map modification tools — DeepEMHancer, EMReady and EM-GAN — were run on a panel of deposited high-resolution maps containing ligands, metals, nucleic acids and lipids, alongside B-factor sharpening and the non-ML PHENIX RESOLVE algorithm. Models were refined into each modified map with PHENIX real-space refinement and compared by Q-score, then inspected visually around individual ligand sites. For protein regions, PHENIX RESOLVE, EMReady and EM-GAN generally improved Q-scores over the deposited maps, while all three DeepEMHancer models slightly reduced them; for ligands and ordered waters, Q-score distributions for all ML methods were worse than for the deposited sharpened maps and the PHENIX RESOLVE maps. Visual inspection found cases where ligand density was removed entirely, including the O2 ligand in hemoglobin by DeepEMHancer and EMReady, and bupivacaine in Nav1.7 by all three ML methods.
How AI was used
The study used three released machine-learning map modification tools as the objects of evaluation rather than building a model of its own. A panel of cryoEM maps was assembled from the EMDB together with one unpublished hemoglobin map, chosen for resolution and for containing non-protein components. Each map was passed through DeepEMHancer (in its highRes, tightTarget and wideTarget variants), EMReady and EM-GAN, which operate directly on volumetric data to modify density voxel-wise using features encoded during training, and in parallel through B-factor sharpening and the non-learned PHENIX resolve_cryo_em maximum-likelihood error-correction algorithm. Atomic models were then refined into each resulting map using PHENIX real-space refinement so that model–map fit could be compared on equal footing, and Q-scores were computed per atom class — backbone, sidechain, nucleic acid, ligand and water. Densities at individual cofactor, ion, small-molecule, nucleic acid and lipid sites were additionally compared by eye across the modified and unmodified maps. No model was trained or fine-tuned in this study.
The shape of the work
Structural · the record, drawn
no AI
Assemble evaluation set of deposited maps
Obtaining raw data, whether by measurement, download or retrieval.
These maps were selected due to their quality, high resolution, and the variety of cofactors/ligands/ions that they containwhere the paper describes this · verbatim
no AI
Non-ML map modification
Cleaning, filtering, normalising or labelling data already obtained.
We also evaluated PHENIX RESOLVE to provide a point of comparison as a map modification approach that is more fully-featured than B-factor sharpeningwhere the paper describes this · verbatim
AI
ML-based map modification
Running a trained model over new data to predict, classify or score. The AI stood in for conventional algorithm.
we compared the performance of three ML-based map modification methods–DeepEMHancer, EMReady, and EM-GAN–with traditional B-factor-based map sharpeningwhere the paper describes this · verbatim
no AI
Real-space refinement of models into modified maps
Cleaning, filtering, normalising or labelling data already obtained.
Q-scores were calculated using models that had been subjected to PHENIX real-space refinement into their respective modified mapwhere the paper describes this · verbatim
no AI
Quantitative map quality scoring
Testing outputs against ground truth.
To provide a quantitative overview of the performance of the map modification techniques in our panel, we calculated the Q-scores for each mapwhere the paper describes this · verbatim
no AI
Visual inspection of ligand and biomacromolecule densities
Extracting understanding from model behaviour.
we sought to further evaluate affected regions of the maps in our evaluation set on an individual basis to gauge map quality near ligand siteswhere the paper describes this · verbatim
What the record says
Technical · every part carries its own basis
+ in the paper~ our reading− not reported
How to read the quotations. A quotation shows where the paper describes something. It does not quote every value beside it: one passage locates a part of the work, and values without their own quotation are our reading of that passage.
The paper's subject is the output of ML map-modification tools; the modified densities produced by DeepEMHancer, EMReady and EM-GAN are the data the conclusions rest on
Examples of ML tools in this class are DeepEMHancer, EMReady, and EM-GAN.where the paper describes this · verbatim
The data presented in the study are deposited in Zenodo, accession number 10934245, doi: 10.5281/zenodo.10934245.where the paper describes this · verbatim
What this paper did not report
Technical · absence is published deliberately
- CodeWhether the code is available is not stated.
- Trained model weightsWhether the trained model is available is not stated.
- ComputeThe hardware or time used is not stated.
- Version of DeepEMHancerWhich version of the model was used is not stated.
- Version of EMReadyWhich version of the model was used is not stated.
- Version of EM-GANWhich version of the model was used is not stated.
About this article
Record aix-00192, version 2, checked by a person on 2026-10-09. The record describes the paper; it does not assess whether the paper's findings are right. The paper is published under CC-BY; quotations are at most 25 words. How we work · Report an error