Embodied AI: from seeing to acting
Putting a cup on a shelf requires locating both objects, choosing actions, controlling a body, and handling slipping, new angles, and unfamiliar environments. A few seconds of video connect data, learning, control, and systems.
Embodied AI studies how systems with bodies understand and act in their environments. Recognizing a cup is only the beginning: a robot must choose a grasp, move its body, and adjust after contact. Perception, planning, control, and learning meet in physical interaction.
Start here with manipulation, then branch into navigation, locomotion, and other bodies. Imitation learns from demonstrated actions; VLA models connect vision, language, and action; reinforcement learning improves behavior through feedback. These approaches can be combined. Data coverage, action interfaces, and evaluation conditions often matter as much as the model.
- Execute manipulation tasks from camera observations and body state.
- Turn human instructions into actions a particular robot can execute.
- Continue acting under changed objects or scenes, using feedback to correct errors.
For any robot demo, identify observations, action outputs, and evaluated starting conditions before discussing generalization.
- 01Image and joint state
- 02Encode inputs
- 03Action sequence
- 04Pick and place
- 05Observe the result
Research milestones
- 2022
Learn actions from multitask demonstrations
RT-1 connects language, camera observations, and robot actions while studying the scale and diversity of real robot data.
- 2023
Generate action sequences
Diffusion Policy denoises continuous action sequences, offering a way to represent several valid action choices.
- 2023
Bring vision-language knowledge into control
RT-2 encodes actions as tokens and co-trains with vision-language tasks. Semantic transfer and action-generation methods are intersecting directions.
- 2024
Generate continuous actions with a VLM
π0 combines a pretrained VLM with a flow-matching action expert, learning from cross-robot data before task-specific fine-tuning.
- 2025
Take tasks into new homes
π0.5 co-trains on heterogeneous data and tasks to study multi-stage manipulation in previously unseen homes.
Key concepts
- Imitation learning
- Learning a policy from paired observations and demonstrated actions. Gaps in demonstration coverage often become deployment failures.
- Vision–Language–Action (VLA)
- Models conditioned on vision and language that output actions, represented as discrete tokens or continuous values.
- Action chunk
- Several future actions predicted together. A system can execute part of the sequence, observe again, and replan.
- Closed-loop control
- Observe after acting and adjust to the updated environment. Feedback rates and control interfaces differ across systems.
ACT and Diffusion Policy: where actions come from
Watch the ACT / ALOHA videos and method diagram. Pick a bimanual task and follow demonstration data, camera inputs, and action prediction. Then compare another way of modeling actions in Diffusion Policy.
For a first reading, being able to explain what the model sees, what it outputs, and how success is judged gives you a basis for further questions.
Imitation learning, VLA and control
Manipulation and imitation learning. Learn actions from human demonstrations. Investigate data collection, action representations, and differences between training and deployment. Use LeRobot to inspect project structure and data.
VLA and generalization. How do language, vision, and action connect? Does the system work with new instructions, objects, or environments? Pair the model with multimodal learning, and use Open X-Embodiment to explore robot data from different sources.
Reinforcement learning and control. For improvement through feedback, return to RL. For contact, geometry, and control, consult MIT Robotic Manipulation and Underactuated Robotics.
From robot data to simulation
Find a real task in a dataset and connect images to actions. Then try inference or simulation in the selected project. RoboTwin is an entry to simulation tasks, data, and evaluation.
Real robots, simulation, and reading offline data have different requirements. Check hardware, GPU memory, operating system, and data requirements before deciding which level to attempt.
Embodied learning paths and paper lists
Lumina's Embodied AI Guide provides a wider map. Once you have a question, consult paper lists on RL-VLA, humanoid robot learning, or efficient VLA.
After a successful demo, ask where failures concentrate and how initial conditions were set. That leads back to experiments and evaluation.
Embodied communities, teams and scaling
The catalog preserves Lumina, courses, and projects and specialized embodied paper lists. Meet researchers through labs; for data and scaling, read Jiaxuan Zou.
Representative papers
Watch the task videos before reading the method diagrams. Compare inputs, action outputs, and training data across the projects.