Different laboratories sometimes report different values for the same observable even when each experiment looks persuasive. A Los Alamos, NIST and Brookhaven team therefore built a workflow that does not hand judgement to an algorithm. Bayesian machine learning instead highlights a few suspicious features within extensive metadata describing historical procedures.
The demonstration concerned the energy spectrum of neutrons released within less than a nanosecond after spontaneous fission of californium-252, an important nuclear-physics standard. The model related discrepancies to measurement design and procedure; physicists then tested plausible causes by simulating historical apparatus or designing modern experiments.
Only this human and experimental follow-up allowed historical data to be corrected or rejected on a physical basis. The resulting spread among spectra fell by as much as a factor of six. That figure belongs to this specific case, not to a general ability of AI to repair scientific literature automatically.
If the workflow transfers, it could help metrology, materials science, chemistry or medical measurement where legacy data are expensive or impossible to repeat. Its realistic value is not choosing a winning result but ranking potential systematic errors so scarce expert time can be spent on the most informative checks.
Translation requires rich comparable metadata, physically credible simulations, independent replication and safeguards against fitting explanations after the fact. Optimistically, demonstrations in several other domains could appear within 2–4 years and a practical specialist toolkit within 5–8 years. Automatically rewriting databases without expert review would not be trustworthy.

Be the first to open the discussion.