Apple Machine Learning Research has published material on STARFlow2, an approach to unified multimodal generation that bridges language models with normalizing flows.

The description states that existing systems for alternating sequences of text and images face various limitations: discrete tokenization may reduce visual accuracy, combining autoregressive text generation with iterative diffusion denoising creates structural asymmetry, and adapting vision-language models for generation can degrade their original understanding.

The practical significance of STARFlow2 and its comparative results remain unclear from the available material: the source contains only a synopsis, without details of experiments or quality assessments.