Generative models: how an image is made

Generating an image from text or progressively turning noise into a sample requires learning patterns in data and a practical sampling process. Research can concern objectives, generation paths, conditioning, or speed.

Generative models learn recurring patterns in data and use them to produce new samples. An image model can start from random numbers or fill in content using text, a sketch, or another image. Research concerns visual quality, but also whether conditions are followed, how diverse repeated samples are, and how much computation is required.

Image generation offers a way into several ideas: GANs train a generator against a discriminator, diffusion models learn denoising at different noise levels, and latent methods move much of the computation into a compressed representation. These ideas overlap and can be combined.

Research tasks
  • Content generation: produce images from text or other conditions for design, creation, and data research.
  • Editing and restoration: fill missing regions, modify content, or restore degraded images.
  • Modeling complex outputs: generate audio, video, or action sequences while representing multiple possible outcomes.

Ask what objective is learned, how many computations sampling requires, and how results are evaluated. One attractive image cannot answer those questions.

From noise to an image
  1. 01Random noise
  2. 02Denoising network
  3. 03Iterative sampling
  4. 04Conditioning

Research milestones

  1. 2014

    Learning through an adversary

    GANs train a generator using feedback from a discriminator, offering an adversarial route to learning a data distribution.

  2. 2020

    Generation as repeated denoising

    DDPM learns predictions at different noise levels and samples progressively from random noise. It is an important route alongside GANs.

  3. 2022

    Doing the expensive work in a smaller space

    LDM compresses images, runs diffusion on compact latent representations, and decodes them into pixels. Text and other conditions can guide denoising.

  4. 2023

    A Transformer can be the denoising network

    DiT processes latent patches with a Transformer and studies compute versus generation quality.

Key concepts

Noise
Random perturbations drawn from a chosen distribution. Knowing the noise added during training lets us construct a task with a known target.
Denoising objective
Predict noise, a clean sample, or a related quantity from a noisy input and its noise level. A common DDPM formulation predicts added noise, which is used to compute sampling updates.
Latent space
An encoder compresses an image into a smaller numerical representation, and a decoder reconstructs it. Lower compute comes with a tradeoff in retained detail.
Sampler
The rule that uses model predictions to move from an initial random state to a sample. Steps, update equations, and randomness all affect speed and quality.

Start with diffusion sampling

Begin with unit one of the Hugging Face Diffusion Course. Observe the relationship between noisy samples and denoising predictions, following each step's inputs and outputs in a small example.

Then choose an explanation from Lilian Weng's Diffusion Models or Yang Song's score-based generative modeling.

The mathematics of diffusion and flow matching

Distributions, noise, gradients, and differential equations gradually enter the picture. Use foundations for the current gap, then the lectures and exercises in MIT Flow Matching and Diffusion, 2025.

First distinguish what the model learns during training from how it is used during generation. Then compare objectives and sampling procedures.

Sampling steps, speed and output quality

Hold the model and input conditions fixed, change sampling steps, and record time and generated outputs. Keep a set of results under matching conditions; selecting only the prettiest image makes the overall effect hard to judge.

These methods can also predict robot actions: see Diffusion Policy and embodied AI. For generation as environment prediction and decision-making, see world models.

Diffusion derivations and model comparisons

For Chinese derivations, see Jianlin Su's diffusion series. For comparative experiments, see LoopDiT. More courses and projects are in the generation catalog.

Representative papers

Compare GAN and DDPM training first, examine LDM's compute tradeoff next, and use DiT to connect Transformers with diffusion.

2014 · NIPS 2014

Generative Adversarial Nets

Generative Adversarial Nets · Goodfellow et al., 2014 · Figure 1
Figure excerpt from the paper · Goodfellow et al., 2014 · Figure 1 · Original paper

The problem

How can a model learn to generate without directly specifying a complex image probability?

The key idea

A generator turns random inputs into samples; a discriminator learns to distinguish them from training examples and supplies feedback.

Why this paper

The two-model game gives a clear alternative training principle.

Where to start

Start with both objectives and alternating updates, then the samples.

A question to keep asking

Training can be unstable or cover only some patterns in the data.

2020 · NeurIPS 2020

Denoising Diffusion Probabilistic Models

Denoising Diffusion Probabilistic Models · Ho et al. · Fig. 1
Figure excerpt from the paper · Ho et al. · Fig. 1 · Original paper

The problem

Can learning denoising tasks produce clear images?

The key idea

Train on known noise added to real images; repeatedly use the learned predictions to sample from an initial random state.

Why this paper

Its training and sampling algorithms provide a useful foundation for later diffusion work.

Where to start

Compare the two algorithms, tracing model inputs and predictions.

A question to keep asking

Original sampling uses many sequential model calls.

2022 · CVPR 2022

High-Resolution Image Synthesis with Latent Diffusion Models

High-Resolution Image Synthesis with Latent Diffusion Models · Rombach et al. · Fig. 3
Figure excerpt from the paper · Rombach et al. · Fig. 3 · Original paper

The problem

Can diffusion process less information while retaining useful image detail?

The key idea

Compress images with an autoencoder, diffuse in latent space, and decode. Cross-attention introduces conditions such as text.

Why this paper

It links compression, quality, and compute, and underpins Stable Diffusion.

Where to start

Trace encoding, diffusion, and decoding; then compare compression levels.

A question to keep asking

Compression loses information. Text conditioning does not guarantee accurate counts or spatial relations.

2023 · ICCV 2023

Scalable Diffusion Models with Transformers

Scalable Diffusion Models with Transformers · Peebles & Xie · Fig. 3
Figure excerpt from the paper · Peebles & Xie · Fig. 3 · Original paper

The problem

Can a scalable Transformer serve as the diffusion network?

The key idea

Process latent patches with a Transformer and compare depth, width, and patch size.

Why this paper

It separates the generation objective from architecture and examines scaling through compute.

Where to start

Read the architecture, then compare compute and generation metrics across sizes.

A question to keep asking

The original studies class-conditioned ImageNet images, not text-to-video generation.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗