Skip to main content
Back to News Hub
🐻Berkeley BAIR
July 1, 2025
General AI

Whole-Body Conditioned Egocentric Video Prediction

Overview

Berkeley BAIR researchers present PEVA, short for Predicting Ego-centric Video from human Actions, a world model for embodied agents. Given past video frames and an action that specifies a desired change in 3D pose, PEVA predicts the next video frame. The work targets the gap that few world models are designed for truly embodied agents acting in the real world.

Key Takeaways

  • Given past video frames and an action specifying a desired change in 3D pose, PEVA predicts the next video frame.

    Our results show that, given the first frame and a sequence of actions, our model can generate videos of atomic actions (a), simulate counterfactuals (b), and support long video generation (c).

  • But few are designed for truly embodied agents.

    In order to create a World Model for Embodied Agents, we need a real embodied agent that acts in the real world.

  • Why It's Hard Action and vision are heavily context-dependent.

    The same view can lead to different movements and vice versa.

  • Egocentric view reveals intention but hides the body.

    First-person vision reflects goals, but not motion execution, models must infer consequences from invisible physical actions.

  • At every moment, our egocentric view both serves as input from the environment and reflects the intention/goal behind the next movement.
Whole-Body Conditioned Egocentric Video Prediction

Given past video frames and an action specifying a desired change in 3D pose, PEVA predicts the next video frame. Our results show that, given the first frame and a sequence of actions, our model can generate videos of atomic actions (a), simulate counterfactuals (b), and support long video generation (c). Recent years have brought significant advances in world models that learn to simulate future outcomes for planning and control.

From intuitive physics to multi-step video prediction, these models have grown increasingly powerful and expressive. But few are designed for truly embodied agents. In order to create a World Model for Embodied Agents, we need a real embodied agent that acts in the real world.

A real embodied agent has a physically grounded complex action space as opposed to abstract control signals. They also must act in diverse real-life scenarios and feature an egocentric view as opposed to aesthetic scenes and stationary cameras. 💡 Tip: Click on any image to view it in full resolution.

For more details please read the original article at Berkeley BAIR.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by Berkeley BAIR
Read the original