CLIP: learning image and text in the same space, from 400 million web pairs, is what opened modern multimodal AI
- What it is
- Radford and colleagues (OpenAI) trained an image encoder and a text encoder to pull correct pairs together and push wrong pairs apart, using 400 million image-caption pairs collected from the web, without manual labels.
- What was demonstrated
- Zero-shot classification: 76.2% ImageNet accuracy without using any of the 1.28 million training examples, matching a supervised ResNet-50; transfer to over 30 datasets.
- What was not
- Weak on fine-grained and counting tasks; inherits web data biases. It is from 2021 — the conceptual base, not the state of the art.
- Why it matters
- Almost every current multimodal model, including this week's “natively multimodal” ones, descends from this cross-modal alignment.
Same theme
Multimodal→Get the next Cinco
Tuesday, 7am BRT, in your inbox. Five items, with sources. No daily newsletter, no promotions, and unsubscribing is one click.