Efficiency and systems: making models run

Why does changing batch size speed up a model? Why can a program struggle before memory is full? Why can throughput improve while users wait longer? These questions appear when models meet software and hardware.

Efficiency and systems research asks how to train and serve models within hardware, time, and cost limits. An unchanged formula can run at very different speeds: data movement, kernel implementation, memory allocation, request scheduling, and communication determine whether GPUs stay busy and how long users wait.

This field is closely tied to model design. Longer contexts increase cache requirements, concurrent requests change memory pressure, and lower precision may affect quality. Read an optimization together with its workload: which memory does it save, which wait does it reduce, and how would another task change the result?

Research tasks
  • Distribute model and training states that do not fit on one device.
  • Reduce data movement and redundant work to use existing hardware more effectively.
  • Measure tradeoffs among model quality, response time, concurrency, and service cost.

Measure the bottleneck before choosing an optimization. Lower memory use, higher throughput, and shorter user waits are different outcomes.

How models use less compute and memory
  1. 01Weight quantization
  2. 02KV cache
  3. 03Continuous batching
  4. 04Pipeline parallelism

Research milestones

  1. 2020

    ZeRO: partition redundant training state

    Shard optimizer states, gradients, and parameters that would otherwise be replicated in data-parallel training.

  2. 2022

    FlashAttention: change data movement

    Preserve exact attention while tiling computation to reduce transfers between GPU memory levels.

  3. 2023

    SmoothQuant: make low precision practical

    Handle activation outliers to support INT8 weights and activations during inference.

  4. 2023

    PagedAttention: manage a growing KV cache

    Organize caches in blocks to reduce fragmentation and duplication and support more concurrent requests.

Key concepts

Throughput and latency
Throughput measures work completed per unit time; latency measures how long one request waits. Larger batches can improve throughput while increasing queues.
Prefill / decode
Prefill processes the input context; decode generates output step by step. Their compute and memory-access patterns differ, so one speed number is insufficient.
KV cache
Stored attention keys and values for processed tokens avoid repeated computation during generation. Context length and concurrent requests consume cache memory.
Quantization
Lower-precision representations reduce storage and some computation costs. Speedups depend on hardware and kernel support.

Compute, memory access and execution overhead

Read Horace He's Making Deep Learning go Brrrr for compute, memory, and execution overhead. Record inputs, hardware, execution mode, and time for a small program, then investigate its bottleneck.

Use Machine Learning Systems for a broader view. For quantization, compression, and efficient computing, see MIT 6.5940, 2024.

Latency, throughput, memory and quality

Interactive services may prioritize time to first token and tail latency; offline batches may prioritize throughput. Memory and cost also constrain deployment. Write these conditions down and retain an evaluation of task quality.

Measure different input lengths or batch sizes, then consult the relevant PyTorch Performance Tuning Guide advice. Remeasure under the same workload after a change.

Distributed training and inference

How To Scale Your Model introduces distributed computation and hardware constraints. With a concrete bottleneck, follow communication, parallelism, memory, and compute.

This also tests your use of AI. It can write benchmark scripts, but you should explain what the test measures and whom the optimization affects. See working with AI.

System case studies, courses and research teams

Compare Purshow's systems notes, the systems catalog, and teams such as MIT HAN Lab to connect model design and deployment constraints.

Representative papers

These papers cover training state, kernel memory traffic, numerical precision, and serving caches. Identify each bottleneck before reading performance plots.

2020 · SC 2020

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models · Rajbhandari et al. (2020), Fig. 1
Figure excerpt from the paper · Rajbhandari et al. (2020), Fig. 1 · Original paper

The problem

How much memory is wasted by replicating training state on every GPU?

The key idea

Progressively partition optimizer states, gradients, and parameters, communicating required pieces when needed.

Why this paper

It explains training memory beyond parameter count.

Where to start

Identify what each of the three stages partitions, then compare memory and communication costs.

A question to keep asking

Sharding changes communication. Benefits depend on bandwidth and device count.

2022 · NeurIPS 2022

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness · Dao et al. · Fig. 1
Figure excerpt from the paper · Dao et al. · Fig. 1 · Original paper

The problem

Can attention be limited by data movement rather than multiplication count?

The key idea

Tile work into fast on-chip memory instead of repeatedly materializing the full attention matrix in device memory.

Why this paper

It shows how the same mathematical operation can have very different execution costs.

Where to start

Start with GPU memory levels and tiling, then IO analysis and runtime.

A question to keep asking

Exact dense attention still has quadratic arithmetic complexity. Speedup depends on shapes, hardware, and implementation.

2023 · ICML 2023

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models

SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models · Xiao et al. (2023), Fig. 2
Figure excerpt from the paper · Xiao et al. (2023), Fig. 2 · Original paper

The problem

Why do a few very large activations make INT8 quantization difficult?

The key idea

Redistribute scale between weights and activations with an equivalent transformation, then quantize.

Why this paper

It connects outliers, numerical error, and hardware execution in one example.

Where to start

Start with activation outliers and scale migration, then accuracy and speed comparisons.

A question to keep asking

The pre-quantization transformation is equivalent; quantization still introduces error. Representative calibration data and suitable INT8 kernels matter.

2023 · SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

Efficient Memory Management for Large Language Model Serving with PagedAttention · Kwon et al. (2023), Fig. 6
Figure excerpt from the paper · Kwon et al. (2023), Fig. 6 · Original paper

The problem

How can changing request lengths avoid wasting KV-cache memory?

The key idea

Split the cache into blocks, map logical sequences to noncontiguous storage, and support cache sharing.

Why this paper

It connects operating-system paging to the capacity of a model-serving system.

Where to start

Start with fragmentation and block mapping, then throughput under matched latency conditions.

A question to keep asking

It optimizes cache management. Concurrency gains must be assessed with latency targets and request lengths.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗