ByteDance's Seed team introduced SeedRealtime, a multimodal language model with native audio and video processing in full-duplex mode. According to MarkTechPost, the model integrates audio, video, and text within a single architecture.

Instead of exchanging individual queries and responses, SeedRealtime interacts with continuous multimodal streams in real time. This means the model is designed for simultaneous perception and generation of audiovisual responses.

The practical significance of the development is currently limited to the claimed architectural approach: the available material contains no data on testing, latency, public access, integrations, or superiority over other models.