Generative models: how an image is made
Generating an image from text or progressively turning noise into a sample requires learning patterns in data and a practical sampling process. Research can concern objectives, generation paths, conditioning, or speed.
Generative models learn recurring patterns in data and use them to produce new samples. An image model can start from random numbers or fill in content using text, a sketch, or another image. Research concerns visual quality, but also whether conditions are followed, how diverse repeated samples are, and how much computation is required.
Image generation offers a way into several ideas: GANs train a generator against a discriminator, diffusion models learn denoising at different noise levels, and latent methods move much of the computation into a compressed representation. These ideas overlap and can be combined.
- Content generation: produce images from text or other conditions for design, creation, and data research.
- Editing and restoration: fill missing regions, modify content, or restore degraded images.
- Modeling complex outputs: generate audio, video, or action sequences while representing multiple possible outcomes.
Ask what objective is learned, how many computations sampling requires, and how results are evaluated. One attractive image cannot answer those questions.
- 01Random noise
- 02Denoising network
- 03Iterative sampling
- 04Conditioning
Research milestones
- 2014
Learning through an adversary
GANs train a generator using feedback from a discriminator, offering an adversarial route to learning a data distribution.
- 2020
Generation as repeated denoising
DDPM learns predictions at different noise levels and samples progressively from random noise. It is an important route alongside GANs.
- 2022
Doing the expensive work in a smaller space
LDM compresses images, runs diffusion on compact latent representations, and decodes them into pixels. Text and other conditions can guide denoising.
- 2023
A Transformer can be the denoising network
DiT processes latent patches with a Transformer and studies compute versus generation quality.
Key concepts
- Noise
- Random perturbations drawn from a chosen distribution. Knowing the noise added during training lets us construct a task with a known target.
- Denoising objective
- Predict noise, a clean sample, or a related quantity from a noisy input and its noise level. A common DDPM formulation predicts added noise, which is used to compute sampling updates.
- Latent space
- An encoder compresses an image into a smaller numerical representation, and a decoder reconstructs it. Lower compute comes with a tradeoff in retained detail.
- Sampler
- The rule that uses model predictions to move from an initial random state to a sample. Steps, update equations, and randomness all affect speed and quality.
Start with diffusion sampling
Begin with unit one of the Hugging Face Diffusion Course. Observe the relationship between noisy samples and denoising predictions, following each step's inputs and outputs in a small example.
Then choose an explanation from Lilian Weng's Diffusion Models or Yang Song's score-based generative modeling.
The mathematics of diffusion and flow matching
Distributions, noise, gradients, and differential equations gradually enter the picture. Use foundations for the current gap, then the lectures and exercises in MIT Flow Matching and Diffusion, 2025.
First distinguish what the model learns during training from how it is used during generation. Then compare objectives and sampling procedures.
Sampling steps, speed and output quality
Hold the model and input conditions fixed, change sampling steps, and record time and generated outputs. Keep a set of results under matching conditions; selecting only the prettiest image makes the overall effect hard to judge.
These methods can also predict robot actions: see Diffusion Policy and embodied AI. For generation as environment prediction and decision-making, see world models.
Diffusion derivations and model comparisons
For Chinese derivations, see Jianlin Su's diffusion series. For comparative experiments, see LoopDiT. More courses and projects are in the generation catalog.
Representative papers
Compare GAN and DDPM training first, examine LDM's compute tradeoff next, and use DiT to connect Transformers with diffusion.