Apple Machine Learning Research unveiled the study "Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning." It examines how multimodal large language models reason about spatial, temporal, and action-related environments.

At the core of the work is whether models can learn visual thinking during training and argue directly at inference. This approach contrasts with Visual CoT, where the model creates intermediate images for visual foresight.

The authors link the use of intermediate images to substantial additional inference costs, especially for proactive video analysis. The source provides no information about experimental results, specific architecture, or performance measurements.