
DeepSeek released an experimental multimodal model V4-Flash-Vision-Exp. According to The Decoder, it adds image understanding to the text capabilities of V4-Flash.
In its own multimodal tests for agent tasks, the model approached Opus 4.8 and, in some cases, surpassed it. These results are presented via the metadata of The Decoder publication, so they cannot be considered independent confirmation or a complete assessment of quality.
editorial commentary
Why it matters
A likely consequence is increased competition between multimodal models, especially in tasks that require working with both text and images. The next observable signals will be independent tests, details of the methodology, and access conditions. Significant uncertainty is tied to the fact that currently only a synopsis from one source is available and internal tests of DeepSeek.