RoboDojo-Sim mountain of dual-system and VLA models

Notes on Embodied IntelligenceThoughts in Progress

Working notes on VLA/WAM, coding agents as world builders, embodied harnesses, task-relevant physics, and semantic grounding.

What kind of embodied model do we need?

The most concise definition of an embodied model today is:

(observation, language prompt, history) → action

But the final goal is clearly larger than regressing robot-arm actions from images. We want robots to develop six kinds of capability, much like mature coding agents:

  1. A scaling law: generalization should improve steadily and predictably as models, data, and compute grow.
  2. High-frequency control: the system should combine high-level semantic understanding with real-time, continuous, precise manipulation.
  3. Continual learning and rapid transfer: it should absorb new tasks from a few demonstrations, much as an agent uses prompts, memory, and skills.
  4. Agentic reasoning: it should plan, supervise execution, locate errors in a trajectory, and correct them proactively.
  5. World-modeling ability: it should understand 3D space, viewpoint, dynamics, and how actions change the world.
  6. Autonomous capability growth: it should explore online, collect experience, form skills, and gradually solve tasks outside its training distribution.

True embodied intelligence should therefore be a closed-loop system that can perceive, predict, act, verify, adapt, and accumulate experience.

The current stack for embodied foundation models

Pre-training

A typical VLA can be summarized as:

VLM backbone + action expert

The vision-language model handles semantic and scene understanding, while the action expert predicts action chunks through diffusion, flow matching, or another generative method.

The action expert itself learns demonstrations through imitation learning and SFT. It can produce smooth actions and complete tasks, but that does not mean it understands the task, knows the success condition, or can tell when it has made a mistake.

WAMs, world-action models, and latent action models try to model state transitions alongside action prediction, for example by learning both:

P(st+1 | st, at)P(at | st, st+1)

This means learning both “what consequence will this action cause?” and “what action should produce this change?” It makes sense: we should try to learn how the world changes and how its dynamics work.

Some recent representative systems include:

  • Atlas: World Labs released this multimodal autoregressive diffusion world model in 2026. It organizes text, images, camera poses, and depth maps into a 3D spatial context for controllable viewpoint generation, 3D reconstruction, spatiotemporal simulation, and robotic Real-to-Sim.
  • Genie: Google DeepMind’s Genie series learns controllable latent actions and environment evolution from videos without action labels. Genie 3 can generate a visual world from text that supports sustained real-time interaction and navigation, targeting open-ended environment simulation and agent training.
  • GEN-1.5: Generalist’s embodied foundation model places a 3–12 second sensor-action demonstration inside an approximately 30-second context as a physical prompt. Without parameter updates, it generates 100 Hz closed-loop actions and demonstrates one-demonstration learning for short-horizon tasks, skill composition, and transfer from simulation demonstrations to real robots.
  • Zero-WAM: Zero-WAM uses human manipulation video as a visual prompt for a new task. Through causal video-action modeling, it jointly predicts future robot observations and executable actions, allowing one policy to follow video and complete tasks unseen during training without fine-tuning.
  • Fast-WAM: Fast-WAM learns action-relevant world representations through future-video prediction during training, then skips explicit video generation at deployment and directly outputs actions. It tests whether the main benefit of WAM can be retained through training-time world modeling while substantially reducing control latency.

Post-training and inference

Possible methods for post-training and deployment adaptation include: SFT for learning from demonstrations and corrective trajectories; RL for optimizing success from environment feedback; on-policy distillation, in which a teacher labels the states actually visited by the student; ICL for conditioning behavior on demonstrations in context; TTT for updating a latent, adapter, or local parameters at deployment; continual learning for absorbing experience over time while limiting forgetting; teacher–student distillation for compressing slow reasoning, world prediction, or repeated sampling into a fast policy; and one-shot/few-shot adaptation for transferring across tasks, scenes, or embodiments from very few demonstrations. These methods can form a loop of attempt, verification, update, distillation, and skill consolidation.

Astra and the MLLM Harness

Beyond scaling the VLA/WAM itself, another mainstream route is to place a frontier MLLM such as Astra outside the policy as a Harness. These models may first be tools: they read images, video, and execution traces; call perception systems, VLAs, WAMs, conventional controllers, and code tools; decompose tasks; and compose atomic skills into long-horizon behavior. The system’s capability then does not live entirely inside one end-to-end model. It also depends on whether the model can select the right tools, organize skills, inspect outcomes, and recompose the plan after failure.

The difficult questions are control granularity and time scale. Asking Astra to choose every action makes latency, cost, and motion stability unacceptable; asking it to plan only once at the beginning leaves it unable to handle contact errors, state drift, and mid-course failure. A more plausible hierarchy lets the controller provide the highest-frequency stability, a VLA or atomic skill handle local closed-loop behavior, and Astra intervene at task decomposition, subgoal transitions, high-uncertainty states, and repeated failures. When to trigger it, how long a horizon each decision should cover, which observations should return upward, and how a successful episode becomes a faster skill will determine whether the Harness is genuinely useful.

Even if Astra is not yet suitable as a high-frequency system 1, it still represents the intelligence ceiling we can currently reach. Its role need not be direct robot control alone. It can serve as a probe, teacher, task constructor, verifier, and skill compiler: first use the strongest intelligence to discover what the system can do, then distill recurring reasoning, corrections, and tool calls into a faster and cheaper policy. The question is therefore not simply whether Astra is a robot model, but how to connect this intelligence ceiling to an embodied system without letting latency and usage consume the gain.

ICL, TTT, and embodied skill adaptation

Systems such as HOST, Zero-WAM, and WAM-TTT emphasize ICL, TTT, and one-shot learning. The focus is moving from training a fixed policy toward letting a pre-trained foundation model absorb information about the current task, scene, and robot at deployment, reducing the cost of acquiring a new skill. ICL keeps parameters mostly fixed and conditions behavior on demonstrations and interaction history. TTT updates an adapter, latent, memory, normalization state, or local parameters from test-time data. Retrieval finds prior trajectories and atomic skills. System identification estimates dynamics and control characteristics from a small number of interactions. Online fine-tuning continuously changes the policy.

Demonstration trajectories, interaction history, and skill libraries are taking on the roles of prompt, context, and skill in embodied systems. A demonstration simultaneously contains:

(task, scene, camera, embodiment, dynamics, action convention, success criterion)

One-shot learning can therefore provide a task prompt while also adapting implicitly to the camera, robot morphology, control frequency, and dynamics. A reusable skill should also contain a precondition, closed-loop controller, termination condition, success verifier, and failure recovery. Generated video can provide a visual target or motion prior, but it must still pass through correspondence, motion intent, robot-compatible action, and closed-loop correction before becoming an executable skill. The ideal loop is:

demonstration → ICL attempt → execution and verification → TTT/OPD → skill distillation

Why are ICL and TTT rising together?

The optimistic reading is that embodied foundation models have developed composable visual, action, and dynamics priors that can be activated by a small amount of on-site information. The cautious reading is that a large domain gap remains between offline policies and real deployment, so performance depends on local adaptation. Their value cannot be judged from a single one-shot success. We should test transfer across tasks, layouts, and embodiments; confirm that the model uses action consequences rather than visual matching; check whether adaptation damages old capabilities; and see whether failure experience can settle into a stable skill.

TTT must also handle catastrophic forgetting and incorrect updates. Options include episodic adapters, embodiment-specific LoRA, latent adaptation, replay, trust regions, update gating, or an episode-level reset. Short-term experience should first enter temporary memory. It should be written into the long-term model or skill library only after a success verifier confirms it, the behavior succeeds repeatedly, and old skills do not regress substantially.

Harness and CaP: shortcut and necessary path

VLA/WAM provides continuous perception and action, while a Harness organizes those capabilities into an embodied Agent that can be inspected, composed, and continually updated. Recent systems such as OpenETA, CaP-X, ASPIRE, RoboHarness, and RoboRSI are converging on a shared pattern: an MLLM handles task decomposition, policy routing, visual attribution, and program revision; VLAs, conventional controllers, and code skills execute; and videos, actions, tool calls, logs, and success criteria form a replayable execution trace.

RoboRSI offers a relatively mature example of an embodied Harness. It organizes Task, Compound Skill, Atomic Task, and Base Skill in a top-down skill tree, then returns execution results to the responsible node along the call chain. This structure makes failures traceable, skills replaceable, and stable procedures easy to consolidate. It also gives subsequent learning a clear place to attach its data.

A coding agent can also become a world builder

There is another natural role here. Astra does not have to sit only above a rollout as a supervisor. Like a coding agent, it can directly construct scenes, environments, and tasks; write rewards, success verifiers, and data pipelines; then revise the scene, task difficulty, and curriculum around the policy’s actual failures. The idea in systems such as AgentWorld—that the world can grow with the agent—suggests that an embodied environment need not remain a benchmark built once and frozen. It can become a programmable, testable, and extensible part of the training loop.

world / env / task → VLA or RL rollout → failure and data → code, curriculum and policy update → harder world

This goes one layer beyond “LLM supervises VLA”: the upper-level model organizes both the policy and the world in which that policy learns. It can turn a real scene into an interactive simulation, generate neighboring tasks and counterexamples at scale, and convert capability gaps exposed by rollouts into new training settings. A verified policy then returns to the real robot and produces new discrepancies and failures. Growing and evolving should happen across the environment, task, data, verifier, and policy—not merely by running a fixed policy in a larger collection of static scenes.

For me, the more important criterion for embodied RSI is whether its capability boundary keeps expanding under OOD conditions. Compositional generalization means rearranging the same atomic skills for new objects, scenes, and task orders. Through explicit task structure, skill interfaces, and execution history, an MLLM Harness has a clear structural advantage here. Systems such as Zero-WAM further show that a new task can enter the model’s context through human video. RSI must go further and ask whether this temporary success can be diagnosed, verified, and consolidated into a capability that is directly reusable next time.

A more direct experiment would give the robot an OOD task while deliberately introducing different kinds of capability gaps: completely remove a required atomic skill; keep the skill but prevent the Planner from composing it correctly; or provide a skill that fails repeatedly on the current embodiment. From observations, the call chain, and physical outcomes, the system should determine whether the missing element is task planning, perception, action data, contact control, dynamics experience, or embodiment adaptation. It should then choose among replanning, retrieving experience, requesting a human demonstration, using ICL, fine-tuning on corrective trajectories, or practicing autonomously under safety constraints.

For example, remove the insert skill from a plug-in task and provide only localization, grasping, motion, and basic force control. In another setting, provide an insert skill transferred from a different robot that repeatedly fails on the current embodiment. The key question is whether the model can recognize that its present capability cannot complete contact alignment, then request one demonstration, collect a few corrective trajectories, or learn a new insertion strategy through constrained exploration. It should ultimately consolidate that strategy into an atomic skill reusable across sockets, positions, and objects. Merely becoming better at the same task still falls short of capability growth.

Evaluation should focus on new-task coverage, capability-gap attribution, time-to-competence, the cost of human demonstrations, the number of autonomous exploration attempts, regression on old skills, and reuse across settings. These measures can tell us whether the system has truly acquired a new capability.

Speed remains a basic constraint for this route. High-frequency actions should be executed by a VLA, controller, or consolidated code, while the MLLM is triggered at task transitions, skill failures, and high-uncertainty states. Beyond per-step inference latency, reports should measure the total wall-clock time from detecting a capability gap to acquiring a stable skill. The efficiency of embodied RSI ultimately appears in a downward trend: when the robot encounters the next OOD task, it should require less human time, fewer real rollouts, and less online inference.

How we think about embodied scaling laws

An embodied scaling law probably exists, but its core independent variable should not be the raw number of trajectories. It should be effective interaction experience that is physically aligned, weighted by task relevance, covers new states, and contains closed-loop feedback:

Neffective(T) = Σi qi ai ri(T) ni fi Imixture

Here, qi represents demonstration and control quality; ai represents alignment in action coordinates, control frequency, sensors, and embodiment; ri(T) represents relevance to the target task; ni measures whether the data covers new states; fi measures whether it includes failures, corrections, rewards, and action consequences; and Imixture represents conflicts and negative transfer across data sources. Embodied data cannot be treated like language tokens: the same numeric action can carry completely different physical meanings on different robots, coordinate systems, or control frequencies.

Diversity is valuable only when the model conditions on it correctly. Diversity in tasks, objects, scenes, geometry, dynamics, and failure recovery can provide the supervision needed for compositional generalization. Differences in operator style, speed, action convention, and embodiment can instead make similar observations map to conflicting action labels when they are not identified explicitly. Modeled diversity expands capability coverage; unconditioned diversity often becomes noise.

The criticism that “another ten years of data collected under the current VLA paradigm still would not be enough” therefore has some force. Success-only offline imitation mainly increases sampling density around expert trajectories. It says little about the consequences of wrong actions, recovery after leaving the distribution, where a long-horizon task fails, or which interaction would reduce current uncertainty most effectively.

Embodied scaling must proceed along four axes: semantic scaling from VLMs, internet data, and human video; physical scaling from aligned real or simulated interaction; rapid-adaptation scaling from ICL, TTT, and system identification; and experience scaling from on-policy rollout, RL, corrective data, and autonomous exploration. The real inflection point will arrive after capability and data form a loop. Stronger models will identify their own uncertainty and capability gaps more effectively and actively collect high-information experience. That experience will improve the models in turn, allowing data production itself to scale with model capability.

The robot world lacks its own tokenizer and ISA

A deep obstacle to cross-embodiment learning is the lack of shared definitions across vendors:

  • joint definitions;
  • coordinate systems;
  • position/velocity/torque interfaces;
  • control frequencies;
  • timestamps;
  • latency;
  • gripper semantics;
  • proprioception;
  • camera calibration;
  • safety control.

Large models are therefore often forced to spend billions of parameters guessing the driver protocols of different vendors.

An ideal system would include a Robot Action IR:

VLA/WAM → Robot Action IR → Embodiment Adapter → Vendor Controller

This IR should be more than a seven-dimensional action vector. It could contain:

atIR = (reference frame, pose/force objective, constraints, horizon, termination, uncertainty)

Vendors could still use different low-level motors while the model shares an intermediate layer of physical action semantics.

This could be the tokenizer, ISA, or USB-C of the embodied world. It cannot eliminate differences in friction, load, latency, and dynamics, but it can separate task learning from driver-protocol adaptation. The remaining differences can then be handled through calibration, system identification, and embodiment adapters.

A few insights about current embodied-model narratives

The gains from diffusion cannot be attributed directly to a “multimodal action distribution”

Diffusion Policy explains the advantage of diffusion as the ability to express a multimodal action distribution, but later work suggests that multimodal modeling may not be the main source of the performance gain. In Much Ado About Noising, Pan et al. control for network architecture, training procedure, and inference compute, and find that ordinary regression policies can approach generative policies. The important factors may instead include iterative supervised computation, regularization from random perturbations, and avoiding the averaged actions produced by direct regression. The two-step regression method MIP can also match flow-based policies. To validate the “multimodality” of diffusion, different samples should correspond to semantically distinct, stable, and successful strategies—such as approaching a grasp from different directions or taking different collision-free paths—rather than trajectory jitter or random variation around one mode.

References: Diffusion Policy; Much Ado About Noising, ICLR 2026.

A module’s contribution must be established against a fully tuned baseline

A complex structure may benefit from extra parameters, auxiliary supervision, training budget, or more extensive tuning rather than from its claimed mechanism. When studying recurrent policies in POMDPs, Luo et al. find that simply assigning a separate, smaller learning rate to the context encoder substantially stabilizes training and outperforms several methods with auxiliary prediction, explicit history features, or specialized memory mechanisms. Similarly, by systematically optimizing architecture, history representation, and hyperparameters, Ni et al. make a simple recurrent model-free baseline match or outperform more complex methods in 18 of 21 environments. A semantic head, future predictor, world-guidance component, or memory module should therefore be compared under matched parameter counts, compute, data, and search budgets. Ablations that remove inference inputs, replace supervision targets, or freeze modules should identify the true causal source of the gain.

References: Luo et al., NeurIPS 2024; Ni et al., ICML 2022.

The value of a world model should be measured through decision-making

Sharp video generation captures only part of visual modeling ability. A robot needs future consistency in object identity, geometry, occlusion, contact state, and action-conditioned change. A blurry prediction may still contain low-frequency geometry and dynamics useful for control. Object drift, penetration, incorrect contact relations, or inconsistency across viewpoints indicate a failure in world-state modeling itself. A more direct evaluation would ask the model to compare the consequences of candidate actions, predict success, collision, and irreversible failure, and test whether world prediction raises OOD success rates, reduces real rollouts, improves recovery, or can be distilled into a better action policy. A world model moves beyond video generation only when its future representation produces a verifiable improvement in action selection.

Simulation does not need one universal “reality”; it needs task-relevant physics

Once a coding agent begins constructing environments automatically, one question becomes unavoidable: how accurate must the physics be? I find it more useful to treat fidelity as a task-dependent budget than to pursue one scalar measure of photorealism or physics fidelity. Semantic search, object counting, and task ordering depend mainly on vision and spatial relations, so lighting, texture, background, and camera can be randomized aggressively. Ordinary pick-and-place requires reasonable geometry, collision, reachability, and friction. Insertion, twisting, folding, liquids, and bimanual coordination quickly depend on contact, compliance, force, friction, delay, and controller response. If an approximation does not change the optimal action, success boundary, or dominant failure mode, it can remain approximate or even randomized. If it changes any of them, visual plausibility is not enough.

A human noticing that light, wind, or an object’s motion looks wrong does not mean that this is the error the robot cares about most. Conversely, gripper delay, stiffness, center of mass, sensor synchronization, and contact-model errors may be invisible to a viewer and still break a policy. The more useful test is not whether a simulation fools a person, but whether action rankings, state transitions, success judgments, and recovery strategies remain consistent between sim and real. The real world exposes what the simulator does not yet know; simulation scales the causal factors that have already been identified.

“The model does not know what it is doing” should become a testable capability

A current VLA/WAM may complete short-horizon manipulation without explicitly representing the task stage, success condition, failure cause, goal reachability, or the capability required next. This contrasts with systems such as SayCan: their low-level actions were limited, but they used language models for task decomposition and affordance ranking. Today’s end-to-end policies are stronger at action, yet task state and semantic judgment are often compressed into a latent representation that is hard to inspect. A language explanation is not a sufficient criterion. Better tests ask whether the system can reject or clarify an impossible instruction; replan when a critical object is missing; locate the failed step after execution drifts; update its plan when the environment is disturbed; and, when an atomic skill is missing, identify the capability gap and choose retrieval, a requested demonstration, safe exploration, or fine-tuning. What matters is intervenable task understanding, counterfactual judgment, and autonomous recovery.

Likewise, possessing a semantic label, an instruction, or a phase index does not mean that the model knows what it is doing. It may predict “step three” or repeat “put the cup in the tray” without knowing why that phase applies, which state change constitutes completion, or when an action has become unrecoverable. Grounding can be tested by removing a critical object, swapping spatial relations, renumbering phases, starting from a partially completed state, introducing a mid-course disturbance, constructing an impossible task, or making visual appearance disagree with the success detector. Genuine semantics should connect instruction, task predicate, world state, action consequence, and success in a loop that can survive intervention.

The research directions that interest us most

Deployment-time post-training for VLA/WAM

Study how a robot can quickly acquire new capabilities from a small number of demonstrations, failure trajectories, and environment feedback generated at deployment. The system could first attempt the task through ICL, then use TTT, on-policy distillation, local SFT, or RL based on the execution result. Verified experience could be consolidated into an adapter, policy update, or long-term skill. The focus is a complete path from temporary adaptation to continual learning while controlling incorrect updates, distribution drift, and catastrophic forgetting. Experiments should distinguish task reproduction, scene adaptation, embodiment adaptation, and genuinely new capabilities reusable across objects and layouts.

Harnesses for embodied foundation models

Build an execution system around a VLA/WAM executor with an MLLM supervisor, memory, planner, skill library, verifier, and recovery mechanism. Beyond interpreting open-world instructions, decomposing tasks, and selecting skills, the MLLM can write or revise the environment, task, reward, success condition, and data-generation code. The VLA handles high-frequency closed-loop control; the WAM predicts action consequences; and the verifier judges intermediate success and triggers recovery. The value of a Harness is that it turns a long-horizon task into traceable state transitions and allows failures to be located in perception, planning, skill selection, environment mismatch, or control. Research should record the full trace, state clearly whether self-improvement changed the prompt, memory, plan, environment, code, skill, adapter, or model weights, and measure whether those changes improve OOD behavior, recovery, and long-horizon capability.

Let environments, tasks, and policies grow together

Let a coding agent generate neighboring tasks, capability-gap probes, and progressively harder curricula from rollout evidence, then improve the policy through RL, OPD, or corrective data. Once the policy crosses the current success boundary, the agent can continue revising scenes, physical parameters, goal combinations, and failure conditions. The important question is not whether many tasks can be generated. It is whether each new task reliably reveals the current policy boundary, whether its reward and verifier are trustworthy, whether the resulting data produces a reusable capability, and whether gaps found in real deployment flow back into the next environment. When all of these parts update together, the process begins to look like environment–policy co-evolution rather than repeated rollout on a static benchmark.

Cross-teaching among MLLM, CaP, VLA, and WAM

Different models can form a complementary teacher–student system. MLLM/CaP can provide task decomposition, semantic subgoals, program structure, success judgment, data filtering, and trajectory annotation. A WAM can predict the future consequences of candidate actions and provide counterfactual supervision. A VLA can handle low-latency closed-loop execution. During training, a slow model can label the student’s on-policy states and correct errors on the distribution the student actually visits through OPD, SFT, or preference learning. Planning, video prediction, and repeated sampling can then be compressed into a fast action model.

But teaching should not run only from MLLM to VLA. A VLA and a real robot may carry low-level physical priors that a language model lacks: which poses are reachable, which contacts will slip, when a motion becomes unstable, and how much delay a controller can tolerate. This knowledge does not first need to become a polished language explanation. Candidate-action rankings, subgoal reachability, action uncertainty, contact outcomes, residuals between predicted and observed states, and clusters of repeated failures can all be returned to Astra, a planner, or a WAM. If the upper-level model changes its plan, task, or environment in response, that becomes VLA → MLLM/WAM teaching rather than a stronger LLM supervising an action model in only one direction.

Embodied Action IR

Study a cross-embodiment intermediate action representation between semantic tasks and native robot control. It could describe end-effector pose changes, contact modes, force intent, keypoint constraints, phase termination, and permitted regions of motion, then use a small embodiment adapter to map these descriptions to the joints, grippers, and control commands of different robots. The main model and high-level skills could then be shared across robots, while a new platform would require only a small amount of calibration, system identification, and action alignment. Evaluation should cover cross-robot zero/few-shot transfer, adapter size, calibration-data cost, control precision, regression on old robots, and whether adding a new embodiment produces positive transfer rather than negative transfer.

A hierarchy between high-frequency control and expensive reasoning

Let a fast VLA or reactive policy handle ordinary closed-loop control, and invoke a WAM, MLLM, or CaP only for OOD states, high uncertainty, repeated failure, task transitions, or long-horizon planning. The slow model can generate candidate futures, compare action consequences, replan, and supply corrective trajectories. After verification, distillation or OPD can compress the new strategy back into the fast action model. The system can also use diffusion cache, agent KV cache, latent reuse, and asynchronous prediction to reduce repeated computation. Final evaluation should combine task success, trigger accuracy, control frequency, end-to-end latency, the number of expensive model calls, and the ability of the distilled student to execute independently of its teacher.

Among the models we can currently access, Astra is a useful highest-intelligence tier for probing the capability ceiling, but the questions are always at what cost and how to solve it. Beyond success rate, our recent exploration focuses on speed, usage, physical intuition, agentic ability, and generalization across settings. Cost should include model calls per episode, token/API spend, simulation and GPU wall time, human repairs, time from a description to a validated task, and the total cost of acquiring one reusable skill. A sensible system is not “Astra everywhere”: use it for first builds, unknown states, and high-uncertainty decisions; use cheaper models for batch checks; leave high-frequency execution to the VLA and controller; then compile repeated reasoning into environment code, verifiers, skills, or student weights.

How to develop real understanding from here

Research can begin with judgments, but in the end we have to run the systems ourselves. Many current benchmarks use short tasks and fixed scenes, where different paradigms may reach similar success rates. Those results make it hard to tell what VLA, WAM, ICL, TTT, and CaP have actually learned. We do not need to rush to dismiss CaP as a shortcut, declare MLLMs too slow, reduce world models to video generation, or say that end-to-end VLAs understand nothing. Reproduce systems step by step, design probes, and test each question directly. That is how real understanding is built.

END / 2026.09.07Back to top ↑