Multimodal models: connecting images and language

Give a model an image and ask what the person on the left is holding. It must connect visual information to the question. Captioning, image–text retrieval, and video understanding involve related connections.

Multimodal learning connects information expressed as images, text, sound, or video. The same cat can appear in a photograph and in a sentence; a model must learn what corresponds and what complements other information. This page starts with images and text to explain retrieval, visual question answering, and visual conversation.

First distinguish matching from generation. CLIP converts images and text into comparable vectors, making it useful for selecting a matching description. A vision-language model must also pass visual information to a language model that produces an answer. Teaching it to follow questions and checking whether answers depend on the image are separate research problems.

Research tasks
  • Cross-modal retrieval: find images from a sentence or match an image with a description.
  • Visual question answering: answer questions about photos, charts, or documents using visible information.
  • Combining evidence: connect video, audio, and text to interpret events or inform later decisions.

Fluent answers are only an observation. Changing the image and checking whether the answer changes begins to test whether visual evidence matters.

Connecting images and language
  1. 01Image and text
  2. 02Two encoders
  3. 03Image–text alignment
  4. 04Visual tokens
  5. 05Response

Research milestones

  1. 2021

    Making images and text comparable

    CLIP learns image–text pairings and maps both inputs into representations with comparable similarities. Matching scores support retrieval or candidate-label selection.

  2. 2023

    Connecting existing vision and language models

    BLIP-2 freezes both large models and trains an intermediate module to extract and translate visual information across their representation gap.

  3. 2023

    Learning to answer visual instructions

    LLaVA uses visual instruction data to train the connection and language responses. Its connector differs from BLIP-2.

Key concepts

Modality
A form of information, such as text, images, or sound. Modalities differ in structure and information density.
Encoder
A model that transforms raw input into numerical features. An image encoder produces representations that still need a task such as matching or question answering.
Contrastive learning
Increase similarity for matched image–text pairs relative to mismatched pairs in a group of examples. Learning to match does not by itself generate an answer.
Connector and instruction tuning
A connector turns visual features into inputs a language model can receive. Instruction tuning uses examples such as questions and answers to teach responses. One concerns architecture, the other training.

CLIP and LLaVA: matching and question answering

CLIP introduces aligning representations through image–text matching. LLaVA shows how images connect to a language model and instruction data teaches answering.

Follow image encoding, representation conversion, and language output with Hugging Face's VLM explanation. Revisit vision, LLMs, or foundations where needed.

Change an image and inspect the answer

Choose an image with clear answers. Ask about objects, counts, positions, and text. Change one detail and see whether the answer changes. Save questions, images, and raw responses.

This helps reveal whether the model uses the image or guesses from language. Turning an observation into research requires control of samples, tasks, and comparisons; see experiments.

Multimodal papers and applications

The multimodal LLM paper index helps locate particular tasks. For making images, see generation; for connecting vision and language to action, see embodied AI.

Encoder design and system costs

Purshow's Encoder-Free analysis explains design through workload and parallel scheduling. More courses, papers, and projects are in the multimodal catalog.

Representative papers

Use CLIP to understand alignment, then compare BLIP-2 with LLaVA: the connection, the trained parameters, and the data.

2021 · ICML 2021

Learning Transferable Visual Models From Natural Language Supervision

Learning Transferable Visual Models From Natural Language Supervision · Radford et al. · Fig. 1
Figure excerpt from the paper · Radford et al. · Fig. 1 · Original paper

The problem

Can captions teach visual concepts beyond fixed category labels?

The key idea

Train image and text encoders so matched pairs are more similar than mismatches. Compare an image with candidate descriptions at inference.

Why this paper

Retrieval and zero-shot classification become an observable matching problem.

Where to start

Follow pretraining and zero-shot prediction, then change the official example's candidate text.

A question to keep asking

CLIP produces representations or scores, not conversation. Correct matching need not imply relational understanding.

2023 · ICML 2023

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models · Li et al., BLIP-2, Figure 1
Figure excerpt from the paper · Li et al., BLIP-2, Figure 1 · Original paper

The problem

Can existing vision and language models be reused with less retraining?

The key idea

Freeze both models and train a Q-Former to extract a compact set of visual features and connect them to the language model.

Why this paper

It turns the idea of giving a language model vision into modules and training stages.

Where to start

Mark frozen and trainable parts in the two-stage framework, then trace visual features into language inputs.

A question to keep asking

Frozen components retain their limitations, and visual errors can accompany fluent but inaccurate answers.

2023 · NeurIPS 2023

Visual Instruction Tuning

Visual Instruction Tuning · Liu et al., Visual Instruction Tuning, Figure 1
Figure excerpt from the paper · Liu et al., Visual Instruction Tuning, Figure 1 · Original paper

The problem

How can a model answer natural-language instructions about an image?

The key idea

Connect an image encoder to a language model with a projection and train on visual instructions. The original uses captions and box information with text-only GPT-4 to create part of its dialogue data.

Why this paper

Data organization can change interaction as much as the connection itself.

Where to start

Start with instruction-data construction, then the connector and staged training.

A question to keep asking

Answers can invent visual details.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗