Embodied AI: from seeing to acting

Putting a cup on a shelf requires locating both objects, choosing actions, controlling a body, and handling slipping, new angles, and unfamiliar environments. A few seconds of video connect data, learning, control, and systems.

Embodied AI studies how systems with bodies understand and act in their environments. Recognizing a cup is only the beginning: a robot must choose a grasp, move its body, and adjust after contact. Perception, planning, control, and learning meet in physical interaction.

Start here with manipulation, then branch into navigation, locomotion, and other bodies. Imitation learns from demonstrated actions; VLA models connect vision, language, and action; reinforcement learning improves behavior through feedback. These approaches can be combined. Data coverage, action interfaces, and evaluation conditions often matter as much as the model.

Research tasks
  • Execute manipulation tasks from camera observations and body state.
  • Turn human instructions into actions a particular robot can execute.
  • Continue acting under changed objects or scenes, using feedback to correct errors.

For any robot demo, identify observations, action outputs, and evaluated starting conditions before discussing generalization.

Observation, action and feedback
  1. 01Image and joint state
  2. 02Encode inputs
  3. 03Action sequence
  4. 04Pick and place
  5. 05Observe the result

Research milestones

  1. 2022

    Learn actions from multitask demonstrations

    RT-1 connects language, camera observations, and robot actions while studying the scale and diversity of real robot data.

  2. 2023

    Generate action sequences

    Diffusion Policy denoises continuous action sequences, offering a way to represent several valid action choices.

  3. 2023

    Bring vision-language knowledge into control

    RT-2 encodes actions as tokens and co-trains with vision-language tasks. Semantic transfer and action-generation methods are intersecting directions.

  4. 2024

    Generate continuous actions with a VLM

    π0 combines a pretrained VLM with a flow-matching action expert, learning from cross-robot data before task-specific fine-tuning.

  5. 2025

    Take tasks into new homes

    π0.5 co-trains on heterogeneous data and tasks to study multi-stage manipulation in previously unseen homes.

Key concepts

Imitation learning
Learning a policy from paired observations and demonstrated actions. Gaps in demonstration coverage often become deployment failures.
Vision–Language–Action (VLA)
Models conditioned on vision and language that output actions, represented as discrete tokens or continuous values.
Action chunk
Several future actions predicted together. A system can execute part of the sequence, observe again, and replan.
Closed-loop control
Observe after acting and adjust to the updated environment. Feedback rates and control interfaces differ across systems.

ACT and Diffusion Policy: where actions come from

Watch the ACT / ALOHA videos and method diagram. Pick a bimanual task and follow demonstration data, camera inputs, and action prediction. Then compare another way of modeling actions in Diffusion Policy.

For a first reading, being able to explain what the model sees, what it outputs, and how success is judged gives you a basis for further questions.

Imitation learning, VLA and control

Manipulation and imitation learning. Learn actions from human demonstrations. Investigate data collection, action representations, and differences between training and deployment. Use LeRobot to inspect project structure and data.

VLA and generalization. How do language, vision, and action connect? Does the system work with new instructions, objects, or environments? Pair the model with multimodal learning, and use Open X-Embodiment to explore robot data from different sources.

Reinforcement learning and control. For improvement through feedback, return to RL. For contact, geometry, and control, consult MIT Robotic Manipulation and Underactuated Robotics.

From robot data to simulation

Find a real task in a dataset and connect images to actions. Then try inference or simulation in the selected project. RoboTwin is an entry to simulation tasks, data, and evaluation.

Real robots, simulation, and reading offline data have different requirements. Check hardware, GPU memory, operating system, and data requirements before deciding which level to attempt.

Embodied learning paths and paper lists

Lumina's Embodied AI Guide provides a wider map. Once you have a question, consult paper lists on RL-VLA, humanoid robot learning, or efficient VLA.

After a successful demo, ask where failures concentrate and how initial conditions were set. That leads back to experiments and evaluation.

Embodied communities, teams and scaling

The catalog preserves Lumina, courses, and projects and specialized embodied paper lists. Meet researchers through labs; for data and scaling, read Jiaxuan Zou.

Representative papers

Watch the task videos before reading the method diagrams. Compare inputs, action outputs, and training data across the projects.

2022 · arXiv

RT-1: Robotics Transformer for Real-World Control at Scale

RT-1: Robotics Transformer for Real-World Control at Scale · Brohan et al., RT-1, Figure 1(a)
Figure excerpt from the paper · Brohan et al., RT-1, Figure 1(a) · Original paper

The problem

Can one model learn instruction-following from multitask robot data?

The key idea

Condition a Transformer on images and language to predict discretized actions.

Why this paper

It makes data scale and diversity concrete in robotics.

Where to start

Inspect inputs, outputs, data collection, and generalization evaluations.

A question to keep asking

Its results depend on a particular robot body and action interface.

2023 · Robotics: Science and Systems

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Diffusion Policy: Visuomotor Policy Learning via Action Diffusion · Chi et al. · Fig. 3
Figure excerpt from the paper · Chi et al. · Fig. 3 · Original paper

The problem

How can a policy represent several valid, coherent action sequences?

The key idea

Denoise an observation-conditioned action sequence, execute part, then observe and predict again.

Why this paper

A tangible bridge from generative models to control.

Where to start

Follow denoising and receding-horizon execution, then Push-T videos.

A question to keep asking

Iterative denoising costs time; demonstration coverage and control frequency constrain deployment.

2023 · CoRL

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control · Zitkovich et al. · Fig. 1
Figure excerpt from the paper · Zitkovich et al. · Fig. 1 · Original paper

The problem

Can web-trained knowledge help robots interpret new instructions?

The key idea

Encode actions as tokens and co-fine-tune with vision-language data.

Why this paper

It connects semantic knowledge to existing action capabilities.

Where to start

Read action encoding, then semantic-generalization tasks.

A question to keep asking

Semantic transfer does not automatically create new low-level motor skills; robot demonstrations remain important.

2024 · arXiv

π0: A Vision-Language-Action Flow Model for General Robot Control

π0: A Vision-Language-Action Flow Model for General Robot Control · Physical Intelligence · Fig. 1
Figure excerpt from the paper · Physical Intelligence · Fig. 1 · Original paper

The problem

How can a vision-language model produce continuous actions across robot tasks and platforms?

The key idea

An action expert uses flow matching to generate continuous action sequences alongside a pretrained VLM. The policy is pretrained on demonstrations from multiple robots, then fine-tuned on task data.

Why this paper

It brings VLM pretraining, action generation, and cross-robot data together, connecting RT-2 and Diffusion Policy.

Where to start

Start with the first-page overview, then the action-expert architecture. Compare direct use after pretraining with task-specific fine-tuning.

2025 · arXiv

π0.5: a Vision-Language-Action Model with Open-World Generalization

π0.5: a Vision-Language-Action Model with Open-World Generalization · Physical Intelligence · Fig. 1
Figure excerpt from the paper · Physical Intelligence · Fig. 1 · Original paper

The problem

Can a robot carry out long tasks in homes it did not visit during training?

The key idea

Building on π0, the model co-trains on robot actions, subtask prediction, object detection, and web vision-language data, connecting high-level subtasks to continuous actions.

Why this paper

It studies multi-stage manipulation in new homes and how the mix of training data affects generalization.

Where to start

Start with the first-page data and deployment overview, then inspect the new-home evaluation and data-source ablations.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗