27 August 2026

WeChat releases models that convert multiple media types into unified format

First reported

TLDR AI ran this on .

  • WeChat released WeMM-Embedding, models that convert text, images, videos, and documents into a single comparable format.
  • The models can process interleaved inputs, meaning text and images mixed together, not just separate files.
  • This allows developers to search or compare across different media types as if they were the same kind of data.

How it was covered

TLDR AITLDR editorial team

WeChat released WeMM-Embedding, multimodal embedding models that map text, images, videos, visual documents, and interleaved inputs into unified representation space. The release expands multimodal AI capabilities.