Skip to content
GenAI BR
Archive · Theme2 entries in this theme

Multimodal

Image, audio, video and text in the same model — and what it costs.

  1. GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

    It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.

    Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026
  2. CLIP: learning image and text in the same space, from 400 million web pairs, is what opened modern multimodal AI

    Almost every current multimodal model, including this week's “natively multimodal” ones, descends from this cross-modal alignment.

    Multimodalintroductorytechnical, beginnerpaper · OpenAI · ICML 2021 · 26 Feb 2021