Computer vision: understanding images
What is in a photo, where is it, which pixels belong to an object, and how can images reveal a 3D scene? Vision makes an inviting entry point because you can see what a model gets right and wrong.
Computer vision extracts useful information from images and video. Organizing photos by person, counting road traffic, and outlining a lesion all turn pixels into decisions. Classification asks what is present, detection adds locations, and segmentation labels pixels. 3D vision also studies distance, shape, and spatial relationships.
Choose an output first, then follow how the model produces it. Convolutional networks reuse small filters across locations; vision Transformers split images into patches and let them exchange information. Architecture is only part of the story: annotation quality and differences between training and test images often determine whether a model is useful.
- Recognition and retrieval: classify an image or find similar content in a collection.
- Localization and measurement: find positions and outlines, count objects, or estimate size.
- Scene understanding: use video or multiple views to estimate motion, depth, and 3D structure.
Clarify the input, labels, and evaluation metric before comparing networks.
- 01Input image
- 02Feature extraction
- 03Object detection
- 04Segmentation
- 05Depth estimation
Research milestones
- 2012
Learning visual features
AlexNet demonstrated the strength of convolutional networks on large-scale image classification. Data, GPU training, and architecture belong in the same explanation.
- 2016
Making deeper networks trainable
ResNet uses additive shortcut connections to improve optimization in deep networks. The resulting features can also support tasks beyond classification.
- 2016
A separate task: recognition plus location
YOLO predicts boxes and classes from the full image in a unified prediction problem. Detection is a different task, not a replacement for classification.
- 2021
Images as sequences of patches
ViT feeds image patches into a Transformer and studies the effect of large-scale pretraining. Convolution and attention offer design choices with different data and compute tradeoffs.
Key concepts
- Convolution / CNN
- A small filter moves across an image, reusing its parameters at different positions. Stacked layers combine local patterns into more complex features.
- Feature / representation
- Intermediate numbers that retain information useful for a task. They can feed a classifier or support detection and retrieval.
- Patch and attention
- A patch is a small image region. Attention uses the input to determine how much information patches exchange. A patch need not correspond to an object.
- IoU and mAP
- IoU measures overlap between predicted and reference boxes. mAP summarizes precision–recall performance across classes; its overlap thresholds and evaluation protocol must match when comparing scores.
Start with classification, detection and segmentation
Start with MNIST digit recognition or YOLO object detection. The former takes you through training, prediction, and a Kaggle submission; the latter begins with detections on your own photos, then training. Follow with the vision-task, training, and convolutional-network sections of CS231n.
Classification asks what, detection also asks where, and segmentation assigns meaning to individual pixels. Choose one task and examine its data and metrics.
Features, receptive fields and generalization
ResNet shows how an architectural change can be supported experimentally. For receptive fields, read Distill's visual explanation.
Inspect a prediction: what cases go wrong, does hiding the background change it, and what happens with another data source? Such observations lead to data, representations, and generalization.
Vision connects to language, generation and action
Connect images and language through multimodal models, make images through generation, and connect vision to action through embodied AI. For runtime and memory, see systems.
Vision courses and research groups
Courses and projects are in the vision catalog. For connections to robot tasks, start with SVL, RAIL, and other labs.
Representative papers
Use AlexNet and ResNet to understand feature learning and training, then choose YOLO or ViT. Follow each paper's inputs, outputs, and experiments.