Skip to content
GenAI BR
MultimodalIn the archive since 06 Sept 2026

CLIP: learning image and text in the same space, from 400 million web pairs, is what opened modern multimodal AI

Multimodalintroductorytechnical, beginnerpaper · OpenAI · ICML 2021 · 26 Feb 2021
Download card
What it is
Radford and colleagues (OpenAI) trained an image encoder and a text encoder to pull correct pairs together and push wrong pairs apart, using 400 million image-caption pairs collected from the web, without manual labels.
What was demonstrated
Zero-shot classification: 76.2% ImageNet accuracy without using any of the 1.28 million training examples, matching a supervised ResNet-50; transfer to over 30 datasets.
What was not
Weak on fine-grained and counting tasks; inherits web data biases. It is from 2021 — the conceptual base, not the state of the art.
Why it matters
Almost every current multimodal model, including this week's “natively multimodal” ones, descends from this cross-modal alignment.

Same theme

Multimodal
  1. GLM-5.3-Flash: an MIT-licensed 320B model served entirely on Chinese chips, at US$ 0.15 per million tokens

    Infra and costtechnicaltechnical, decision-makerannouncement · Z.ai · 26 Aug 2026

Get the next Cinco

Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.

Just the e-mail. No third-party trackers on this page. Privacy policy.