Apple Machine Learning Research published a study on the alignment of multimodal large language models, or MLLM. In the provided description, it is stated that the impact of preferred alignment on such systems has been studied less than for ordinary language models.

According to the description of the study, multimodal models that work with images face hallucinations. An error can manifest not only as a wrong fact, but also as an answer that does not correspond to the content of the image.

The authors consider one of the main goals of alignment to be bringing the model's responses closer to the information contained in the image. The source is presented as an abstract of metadata, so details of methods, experiments, and results are missing from the provided materials.