Large language models: understanding, generation, and reasoning

Behind the chat models you use daily are connected questions: how text becomes model input, what data teaches the model, why different training changes its answers, and how to tell whether it improved.

Research can begin with any one of those steps.

A large language model turns text into tokens, learns relationships from large text collections, and generates continuations from context. One model can support questions, summaries, code, and translation. Tasks are often expressed through instructions and examples instead of a separately trained classifier.

The field studies both how capabilities emerge from data and training and how to use them for particular tasks. Pretraining, instruction tuning, preference learning, and search at inference time affect different stages. Ask what a paper changes and how it tests that change before interpreting a claim of better answers.

Research tasks
  • Turn scattered text into summaries, structured information, or answers that can be queried further.
  • Generate, revise, and explain text and code from natural-language requests.
  • Express different tasks through one language interface and study transfer to new requirements.

Separate architecture, training objectives, and usage: long answers, instruction following, and correct reasoning require different tests.

How text becomes a response
  1. 01Text → tokens
  2. 02Representations and attention
  3. 03Next-token distribution
  4. 04Select and append a token

Research milestones

BERT 2018

Bidirectional encoding and representations

GPT 2018

Autoregressive generation and scaling

Explore model families and technical reports
  1. 2017

    Architecture: exchange information across positions

    Transformer organizes sequence computation through attention. Originally studied for translation, it became an important building block for later pretrained models.

  2. 2018

    Transfer: learn from text, then adapt to tasks

    GPT connects generative language pretraining with supervised task adaptation, reusing learned representations.

  3. 2020

    Prompting: put task examples in context

    GPT-3 studies scale and few-shot prompting without updating parameters for each evaluation task.

  4. 2022

    Behavior: train answers toward human intent

    InstructGPT combines demonstrations, preference comparisons, and reinforcement learning. How a model answers becomes a training problem of its own.

Key concepts

Token
A discrete piece of input: a word, character, or part of a word. Different tokenizers can assign different lengths to the same text.
Autoregressive generation
Predict a token from preceding context, append it, and continue. During training, predictions at multiple positions can be computed in parallel.
Pretraining and post-training
Pretraining establishes broad capabilities; post-training adjusts task skills or behavior. Instruction tuning and preference learning are common post-training methods.
In-context learning
Instructions and examples in the input alter performance without updating weights during that use. This differs from fine-tuning on the examples.

Tokenization, representations and next-token prediction

Use the Hugging Face LLM Course to meet models, tokenizers, and datasets. Inspect the tokenization of a sentence and the model's inputs and outputs. For embeddings, matrix operations, or gradients, revisit foundations.

Then follow text, tokens, attention, and outputs through LLMs from Scratch. The repository accompanies the book; begin with the module you want to understand.

Pre-training, post-training and inference efficiency

How capabilities are learned: study data, objectives, and optimization. Once you can explain training code, enter Stanford CS336, 2025 through a topic in data, training, or systems.

How a model becomes better for a task: study fine-tuning, feedback, and evaluation. Fix the task and evaluation before comparing base and adapted models. Rewards and policy updates connect to RL.

Why models are slow and expensive: see efficiency and systems. Scale, context, precision, caching, and hardware jointly determine feasible experiments.

Language representations, semantics, and tasks also deserve study. CS224N connects them with model methods.

Transformer, BERT and LoRA

Read Transformer for architecture, BERT for pretraining objectives and representations, and LoRA for reducing trainable adaptation parameters. Pair a paper with Mu Li's readings.

Then read Jiaxuan Zou's pretraining and scaling as methodology and scientific perspective to see how a researcher connects scale, data, stability, and problem selection.

Inspect a small model’s tokens and outputs

Choose a small model you can run and a small text set. Record tokenization, input length, outputs, and errors. Change one setting and observe its effect. Check the selected tutorial's model and device requirements.

For tool use, continue to agents; for image understanding, see multimodal models.

Position encoding, optimizers and scaling

For positional encodings and optimizers, see Jianlin Su; for scaling experiments and extrapolation, see Jiaxuan Zou. Courses and projects are in the LLM catalog.

Model families and technical reports

Compare reports through a concrete question: how data is chosen, how training objectives change, how reasoning is measured, or where compute is spent.

Tencent · Hunyuan

  • Hunyuan-Large ↗2024 · Technical report

    MoE architecture, synthetic data and training at scale.

Research and further releases ↗

Xiaomi · MiMo

  • MiMo ↗2025 · Technical report

    How pre-training and post-training jointly improve reasoning.

  • MiMo-Audio ↗2025 · Technical report

    Speech representations, audio language modeling and speech-task evaluation.

  • MiMo-V2-Flash ↗2026 · Technical report

    Model architecture, attention design and inference efficiency.

Research and further releases ↗

ByteDance · Seed

  • Seed1.5-Thinking ↗2025 · Technical report

    Reasoning RL, training stability and capability evaluation.

Research and further releases ↗

Moonshot AI · Kimi

  • Kimi K2 ↗2025 · Technical report

    MoE, optimization and large-scale training for tool interaction.

  • Kimi K2.5 ↗2026 · Technical report

    Vision-language training, tool use and parallel agents.

Research and further releases ↗

DeepSeek

  • DeepSeek-V3 ↗2024 · Technical report

    MoE, MLA, load balancing and FP8 training.

  • DeepSeek-R1 ↗2025 · Technical report

    Reasoning through RL, cold-start data and distillation.

  • DeepSeek-V3.2 ↗2025 · Technical report

    Sparse attention, reasoning RL and synthetic agent tasks.

  • DeepSeek-V4 ↗2026 · Technical report

    Hybrid compressed attention, long context and the post-training pipeline.

Research and further releases ↗

Alibaba · Qwen

  • Qwen2.5 ↗2024 · Technical report

    Pre-training data, instruction tuning and capabilities across model sizes.

  • Qwen2.5-VL ↗2025 · Technical report

    Dynamic resolution, video understanding and visual localization.

  • Qwen3 ↗2025 · Technical report

    Dense and MoE models, thinking modes, multilingual training and distillation.

  • Qwen3-Omni ↗2025 · Technical report

    Unified processing of text, images, audio and video.

Research and further releases ↗

MiniMax

  • MiniMax-01 ↗2025 · Technical report

    Lightning Attention, MoE and long context.

  • MiniMax-M1 ↗2025 · Technical report

    Long-context reasoning and test-time computation.

Research and further releases ↗

Z.ai · GLM

  • GLM-4.5 ↗2025 · Technical report

    Hybrid reasoning modes, staged post-training, coding and agents.

  • GLM-5 ↗2026 · Technical report

    Long-horizon agents, asynchronous RL and software-engineering evaluations.

Research and further releases ↗

Google DeepMind

  • Gemini 1.5 ↗2024 · Technical report

    Long-context retrieval, multimodal understanding and million-token evaluations.

  • Gemini 2.5 ↗2025 · Technical report

    Training and evaluation of reasoning, multimodal and agent capabilities.

Research and further releases ↗

OpenAI

  • GPT-4 ↗2023 · Technical report

    Scaling predictions, capability evaluation and post-training.

  • GPT-4o ↗2024 · System card

    Multimodal input/output and evaluations of voice interaction.

Research and further releases ↗

Anthropic

  • Claude Sonnet 4.5 ↗2025 · System card

    Coding and agent evaluations, alignment and model behavior tests.

Research and further releases ↗

Meta AI

  • Llama 3 ↗2024 · Technical report

    Data mixtures, pre-training, post-training and large-scale infrastructure.

Research and further releases ↗

Representative papers

Read these four papers for architecture, pretraining, in-context learning, and post-training. Identify the stage each changes.

2017 · NIPS 2017

Attention Is All You Need

Attention Is All You Need · Vaswani et al. · Fig. 1
Figure excerpt from the paper · Vaswani et al. · Fig. 1 · Original paper

The problem

Can a sequence model handle dependencies without recurrent computation?

The key idea

Each position reads other positions through attention, combined with feed-forward layers and positional information.

Why this paper

Later models reuse these components. It is a useful first map of information flow.

Where to start

Start with the architecture and attention explanation, then the translation experiments.

A question to keep asking

The original is an encoder–decoder translation model.

2018 · OpenAI technical report · 2018

Improving Language Understanding by Generative Pre-Training

Improving Language Understanding by Generative Pre-Training · Radford et al., 2018 · Figure 1
Figure excerpt from the paper · Radford et al., 2018 · Figure 1 · Original paper

The problem

How can unlabeled text help supervised language-understanding tasks?

The key idea

Pretrain by predicting text, then fine-tune on task data arranged as input sequences.

Why this paper

It makes the transfer from general representations to particular tasks concrete.

Where to start

Figure 1 shows the training stages and task-specific input transformations.

A question to keep asking

Downstream task fine-tuning is still required.

2020 · NeurIPS 2020

Language Models are Few-Shot Learners

Language Models are Few-Shot Learners · Brown et al., 2020 · Figure 2.1
Figure excerpt from the paper · Brown et al., 2020 · Figure 2.1 · Original paper

The problem

Can a few examples specify a new task without weight updates?

The key idea

Scale an autoregressive model and evaluate instructions and demonstrations under zero-, one-, and few-shot settings.

Why this paper

It distinguishes in-context learning from fine-tuning and connects to questions about scale.

Where to start

Start with the three evaluation settings, then inspect a familiar task.

A question to keep asking

Results depend on the task and prompt; overlap between web training data and tests complicates evaluation.

2022 · NeurIPS 2022

Training language models to follow instructions with human feedback

Training language models to follow instructions with human feedback · Ouyang et al. · Fig. 2
Figure excerpt from the paper · Ouyang et al. · Fig. 2 · Original paper

The problem

How can a text-completion model better follow a user's request?

The key idea

Fine-tune on demonstrations, learn a reward model from ranked answers, and use reinforcement learning to adjust behavior.

Why this paper

It separates pretrained capability from the way answers are produced.

Where to start

Follow the training pipeline and identify the data needed for demonstrations, ranking, and policy updates.

A question to keep asking

Preferences depend on labelers and prompt distributions; preferred answers can still be factually wrong.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗