
Apple Machine Learning Research presented a study on how consistently large language models update their probabilistic beliefs when new evidence appears. The authors describe this question through divergence from Bayesian updating.
The work proposes an approach that treats language models as information-processing rules and measures the gap between their updates and Bayesian updating. In the description of the study medicine, science and law are named as fields where, given the available data, there is often not a single correct answer.
The practical value of the approach relates to checking model behavior in tasks with uncertainty. But the provided package contains only a publisher's summary: it contains no data about tested models, experiments, numerical results, or comparisons of systems.
editorial commentary
Why it matters
A likely consequence is increased attention to checking the process of updating a model's beliefs in tasks with uncertainty, not only to its final answers. The nearest observable signal is the publication of a full description of the experiments with specific models and numerical measurements. Substantial uncertainty remains due to the absence of such data in the provided source.