Apple Machine Learning Research presented a study on how consistently large language models update their probabilistic beliefs when new evidence appears. The authors describe this question through divergence from Bayesian updating.

The work proposes an approach that treats language models as information-processing rules and measures the gap between their updates and Bayesian updating. In the description of the study medicine, science and law are named as fields where, given the available data, there is often not a single correct answer.

The practical value of the approach relates to checking model behavior in tasks with uncertainty. But the provided package contains only a publisher's summary: it contains no data about tested models, experiments, numerical results, or comparisons of systems.