Multimodal models: connecting images and language
Give a model an image and ask what the person on the left is holding. It must connect visual information to the question. Captioning, image–text retrieval, and video understanding involve related connections.
Multimodal learning connects information expressed as images, text, sound, or video. The same cat can appear in a photograph and in a sentence; a model must learn what corresponds and what complements other information. This page starts with images and text to explain retrieval, visual question answering, and visual conversation.
First distinguish matching from generation. CLIP converts images and text into comparable vectors, making it useful for selecting a matching description. A vision-language model must also pass visual information to a language model that produces an answer. Teaching it to follow questions and checking whether answers depend on the image are separate research problems.
- Cross-modal retrieval: find images from a sentence or match an image with a description.
- Visual question answering: answer questions about photos, charts, or documents using visible information.
- Combining evidence: connect video, audio, and text to interpret events or inform later decisions.
Fluent answers are only an observation. Changing the image and checking whether the answer changes begins to test whether visual evidence matters.
- 01Image and text
- 02Two encoders
- 03Image–text alignment
- 04Visual tokens
- 05Response
Research milestones
- 2021
Making images and text comparable
CLIP learns image–text pairings and maps both inputs into representations with comparable similarities. Matching scores support retrieval or candidate-label selection.
- 2023
Connecting existing vision and language models
BLIP-2 freezes both large models and trains an intermediate module to extract and translate visual information across their representation gap.
- 2023
Learning to answer visual instructions
LLaVA uses visual instruction data to train the connection and language responses. Its connector differs from BLIP-2.
Key concepts
- Modality
- A form of information, such as text, images, or sound. Modalities differ in structure and information density.
- Encoder
- A model that transforms raw input into numerical features. An image encoder produces representations that still need a task such as matching or question answering.
- Contrastive learning
- Increase similarity for matched image–text pairs relative to mismatched pairs in a group of examples. Learning to match does not by itself generate an answer.
- Connector and instruction tuning
- A connector turns visual features into inputs a language model can receive. Instruction tuning uses examples such as questions and answers to teach responses. One concerns architecture, the other training.
CLIP and LLaVA: matching and question answering
CLIP introduces aligning representations through image–text matching. LLaVA shows how images connect to a language model and instruction data teaches answering.
Follow image encoding, representation conversion, and language output with Hugging Face's VLM explanation. Revisit vision, LLMs, or foundations where needed.
Change an image and inspect the answer
Choose an image with clear answers. Ask about objects, counts, positions, and text. Change one detail and see whether the answer changes. Save questions, images, and raw responses.
This helps reveal whether the model uses the image or guesses from language. Turning an observation into research requires control of samples, tasks, and comparisons; see experiments.
Multimodal papers and applications
The multimodal LLM paper index helps locate particular tasks. For making images, see generation; for connecting vision and language to action, see embodied AI.
Encoder design and system costs
Purshow's Encoder-Free analysis explains design through workload and parallel scheduling. More courses, papers, and projects are in the multimodal catalog.
Representative papers
Use CLIP to understand alignment, then compare BLIP-2 with LLaVA: the connection, the trained parameters, and the data.