Archive · Theme
Multimodal
Image, audio, video and text in the same model — and what it costs.
GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens
It is the clearest datapoint yet that a competitive model can be trained and served at scale without NVIDIA hardware — and it prices Opus-class agentic work at cents.
CLIP: learning image and text in the same space, from 400 million web pairs, is what opened modern multimodal AI
Almost every current multimodal model, including this week's “natively multimodal” ones, descends from this cross-modal alignment.