In brief
- The next generation of science experiments will produce vast amounts of data, challenging capacity for storage and analysis.
- SLAC researchers developed a new method that uses AI to compress data without erasing subtle details that are valuable for experiments but lost in conventional file-compression methods.
- The method separates features of a dataset by size, then shrinks and maps these features onto a neural network to preserve them.
Next-generation science experiments will collect more data than ever – so much so that they’ll surpass the capabilities of current data storage and analysis methods. To help, researchers at the Department of Energy’s SLAC National Accelerator Laboratory developed a method to use artificial intelligence to compress large amounts of raw data without losing subtle details critical to scientific discovery. They published the work in Nature Machine Intelligence.
“There is going to be such a flood of data that there’s really no way to handle it in the way we’ve done before,” said Joshua Turner, a lead scientist at SLAC and the Stanford Institute for Materials and Energy Sciences, a joint institute between SLAC and Stanford, and principal investigator of this work. “There are many applications in science now where data storage and analysis speed are really important problems, and I think this method is a clever way to solve them.”
There is going to be such a flood of data that there’s really no way to handle it in the way we’ve done before.Joshua TurnerLead scientist, SLAC and Stanford Institute for Materials and Energy Sciences
The method could be useful for data coming from SLAC’s Linac Coherent Light Source (LCLS), an ultrafast X-ray free-electron laser that will eventually generate up to a million X-ray pulses per second to take “snapshots” of atoms and molecules. Nearly one terabyte of data per second requires novel types of processing needed for instruments that draw on the full capabilities of the LCLS, such as the X-ray photon fluctuation spectroscopy (XPFS) instrument, which will study the movements of particles in exotic topological and quantum materials.
Saving the science-rich speckles
Conventional data-compression methods can erase some of the fine details in measurements that correspond to valuable scientific information. For example, the tiny speckles in X-ray images of molecules can contain important information about how materials transform.
“Those speckles often reflect the underlying arrangement, disorder, or dynamics of a material,” said Yuan Ni, research associate at SLAC and lead author of the work. “If we lose them, we would lose unique scientific insights, like how a material is structured and how that changes over time.”
The AI-based method uses neural networks to compress data, reducing the overall file size and making it easier to store, move, and manage, while allowing users to control what information is kept, like the finer details. “Depending on the underlying data and the desired quality/fidelity, we can typically achieve 10- to 100-fold reductions in file size,” Ni said.

The AI-based method uses neural networks to reduce the overall file size while allowing users to control what information is kept. The neural network encodes features of the measurement at different scales in a compact form. At decoding, users can select a small region of interest to decode, and this can be done at different scales and resolutions, allowing finer features to be recovered when needed. | Yuan Ni and Greg Stewart / SLAC National Laboratory
The team tested the method on various types of data, including measurements of molecules and materials from several experimental techniques, solar magnetic field measurements, and photographs. They found the neural network was able to adapt to different types of data, learning what features matter for different measurements.
A new tool to compress, remember, and retrieve data
Unlike traditional compression, this AI-based method doesn’t compress an entire dataset equally. First, it uses a mathematical tool, known as wavelet analysis, to separate features of the data by scale. Then, the neural network compresses the different-scaled features separately, ensuring the finer features aren’t lost by generalized compression. The neural network learns a compact representation of those features to preserve them.
In addition to lowering the cost of storing and managing massive datasets, the method makes retrieving data easier. Sometimes, researchers want to revisit a tiny slice of compressed data. With this method, they can decompress only the data they want, saving time and cost.
“If you are using a conventional compressor, you would need to decompress the entire file, which could take you minutes, hours, or days,” said Zhantao Chen, assistant professor at the University of Texas at Austin, who worked on this method while a SLAC research associate. “This method can decompress only the region of interest rather than the entire dataset, so it’s much more efficient.”
This new data-compression approach can operate alongside broader data-reduction techniques such as selecting only the events and features of interest, the researchers noted. “Rather than replacing existing compression methods,” Ni said, “our work provides an additional AI-based approach.”
For more information
Other contributors include the University of California, Davis, and Carnegie Mellon University. To train the neural networks, the researchers used Perlmutter, a computational resource of the National Energy Research Scientific Computing Center (NERSC), a U.S. Department of Energy Office of Science User Facility located at Lawrence Berkeley National Laboratory. This work was supported by the DOE Office of Science and the Laboratory Directed Research and Development program at SLAC National Accelerator Laboratory. LCLS is an Office of Science user facility.
Citation: Y. Ni et al., Nature Machine Intelligence, 24 August 2026 (10.1038/s42256-026-01287-9)
This story was originally published by SLAC National Accelerator Laboratory.
Media contact
Media inquiries: media@slac.stanford.edu
Other questions or comments: SLAC Strategic Communications & External Affairs, communications@slac.stanford.edu
Writer
Chris Patrick
