Learning Transferable World-Action Models
from Task-Paired Human Videos

WATCH: World-Action model Training from Cross-Human videos

▶1 NAVER AI Lab   ▶2 Georgia Institute of Technology
*This work was done while at NAVER AI Lab.
NAVER AI Lab
Teaser video goes here: assets/teaser.mp4

Abstract

World-action models inherit broad visual and physical priors from video pre-training. However, standard within-video next-frame prediction allows them to rely on appearance shared between the context and future, rather than learning interactions that transfer across scenes and embodiments. We introduce WATCH, a cross-video pre-training that learns from task-paired human demonstrations collected across diverse environments. Conditioned on the video and actions of one demonstration, the model jointly predicts another demonstration’s actions and future frames from its current observation. Predicting one demonstration from another encourages the model to learn task-relevant interactions that transfer across scenes and demonstrators.

Our pre-training shows that these capabilities can be learned from readily scalable human–human pairs alone, without requiring matched human–robot demonstrations. Analyses before robot post-training show improved motion alignment across embodiments and better action and video prediction on unseen robots, with representations closer to those learned from robot pairs. After robot post-training, WATCH improves MimicDroid in-context success from 40.9% to 68.3% with Cosmos Policy and from 59.8% to 77.7% with Cosmos3, with the largest gains on the unseen levels. Gains persist under context-free post-training across robot embodiments. WATCH outperforms within-video pre-training and co-training baselines, supporting cross-video prediction as a scalable approach to learning transferable world-action representations.

Overview

(a) Human videos of the same task are paired across collection sources and scenes. (b) Given the reference video and its actions, the model predicts the target’s actions and future frames from the target’s initial frame. (c) After robot post-training, WATCH improves success rates over the base model in both in-context and context-free settings.

Within-video pre-training

p(at:t+C, It+1:t+C+1 | It, st, ℓ)

Context and target come from the same demonstration and share its background, viewpoint and demonstrator.

Cross-video pre-training (WATCH)

p(aqt:t+C, Iqt+1:t+C+1 | Iqt, sqt, ℓ, Ir, ar)

The reference comes from a different demonstration of the same task, and the query’s current frame defines the scene to predict in.


Data

Human demonstrations collected by different people, with different rigs and in different environments, vary along the same axes on which a robot differs from a human. We pair demonstrations of the same task across these collection sources, without any robot data. Query–reference pairs are built from different episodes of EgoVerse, with human actions expressed in a shared human–robot task space.

  1. Task taxonomy. A semantic task taxonomy restricts candidate references to compatible activities.
  2. Motion matching. Candidates are ranked by the dynamic-time-warping distance between motion trajectories, standardized per window.
  3. Reference set. Up to 8 nearest matches from the query’s source and up to 7 from other sources. One reference is sampled at each training step.
Query–reference pairsPaired 3-s windowsReferences per queryQueries with a cross-source reference
10.99M10.0K h10.52 (6.35 cross-source)90.7%

Examples of Paired Demonstrations

Query windows and their retrieved references, ranked by motion distance. Hand trajectories are overlaid.

Query

Retrieved references

Different collection source Same collection source

Model

WATCH follows the joint video–action formulation of Cosmos Policy. The reference video and actions, the query’s current frame and the proprioceptive state are clean conditioning inputs; the query’s action chunk and future frames are noised and denoised by the model. The same layout is used in robot post-training, either with a human demonstration as the reference or with the reference slots removed.

task ℓ
aqt:t+C
Action chunk
Future video
WATCH World-Action ModelCosmos-Predict2.5
aru:u+C
Reference action
removed
Reference video
removed
sqt
Current proprio.
Current frame
Current frame
noise
Noisy action chunk
Noisy future video
Reference (clean) Query context (clean) Generated (noised, loss applied)

Robot Post-Training

After pre-training on human pairs, the model is post-trained on the GR1 humanoid and conditioned on a human demonstration at test time. MimicDroid has three levels: L1 (seen object, seen environment), L2 (unseen object, seen environment) and L3 (unseen object, unseen environment). All rollouts use Cosmos Policy (Cosmos-Predict2.5 backbone).

MimicDroid Rollouts

WATCH rollouts and the human reference, from three cameras.

Human reference

Rear

Left

Right

WATCH rollout

Rear

Left

Right

Comparison

Cosmos Policy without pre-training, with within-video pre-training and with WATCH, on the same evaluation episode and reference. One episode per task.

Camera

Human reference

→

Without pre-training

Within-video

WATCH (Ours)

2× Speed Reference

The human reference is played at 2× speed in post-training and evaluation. One episode per task.

Camera

Human reference (2×)

→

Without pre-training

Within-video

WATCH (Ours)

TableTop-Sim (Bimanual ALOHA)

Context-free post-training on a bimanual ALOHA robot. One episode per task.

Without pre-training

Within-video

WATCH (Ours)

RoboCasa Kitchen-H50 (Franka)

Context-free post-training on a single-arm Franka with 50 demonstrations per task. One episode per task.

Camera

Without pre-training

Within-video

WATCH (Ours)

Context-Free MimicDroid

MimicDroid (GR1 humanoid) with the reference slots removed in post-training and evaluation. One episode per task.

Camera

Without pre-training

Within-video

WATCH (Ours)

Scaling Pre-training Data

Post-training
Human pairs (WATCH) Robot pairs, 8× collection cost Robot pairs, 4× to 16× cost Within-video, all human data

MimicDroid success (%) against the collection budget, in units of the full human corpus. Each panel has its own y range.

  • Success increases with the amount of paired human video used for pre-training. With 25% of the data, WATCH matches within-video pre-training on all human data on in-context L2 and L3, and it exceeds that baseline at 50% on L3 and 75% on L2.
  • With robot video counted at 8× the collection cost of human video, in-context L2 reaches 72.5% with the full human data, and the largest robot-pair setting reaches 49.0% at about twice that cost.

Analysis

We analyze the pre-trained model before robot post-training, on held-out EgoVerse pairs, RoboMIND and MimicDroid observations.

Sensitivity to the Reference

Action MSE by reference

137 held-out episodes · mean ± SEM · lower is better

Holm-corrected p, episode-paired Wilcoxon test.

Attention for the same query under two references

WATCH: action-slot attention · Original video model: video-slot attention

Model

Ground truth

Prediction

Attention, current frame

Attention, reference

Same
Attention on the current frame, same-source reference
Different
Attention on the current frame, different-source reference

The same query is conditioned on a matched reference, a similar-motion reference from a different collection source, and a different-task reference from the same source. Action MSE is 0.0322, 0.0335 and 0.0367, respectively; only the task change is significant. With a reference from another source, attention stays on the hands and the manipulated object.

Embodiment-Invariant Features

Legend: original video model, within-video, WATCH
Motion retrieval across human sources

(a) Retrieval, human

Motion retrieval across robots

(b) Retrieval, robot

Held-out embodiment regression

(c) Held-out embodiment regression

On 359 held-out EgoVerse windows and 396 RoboMIND windows from three robots, WATCH retrieves the matching motion from a different human source (a) and a different robot (b) at higher rank than the original video model and within-video pre-training. A regressor trained on two embodiments and tested on the third (c) reaches a Spearman correlation of 0.18 (human) and 0.24 (robot) from WATCH features, compared with 0.06 and 0.01 for the original video model.

Effect of Joint Action–Video Supervision

Action slot

Action-only model, action-slot attention
Act only
WATCH, action-slot attention
Act+Video (Ours)

Video slot

Video-only model, video-slot attention
Video only
WATCH, video-slot attention
Act+Video (Ours)

All three models use cross-video pairs and differ only in the prediction target: actions, future frames, or both. Action-only attention peaks near the image boundaries, and video-only attention spreads over the scene. With joint supervision, both action-slot and video-slot attention concentrate on the hand and the object. Frames are from MimicDroid, which is not seen during pre-training.

Attention on Robot and Human Frames

Original video model attention
Original video model
Within-video attention
Within-video
WATCH attention
WATCH (Ours)

Current-frame attention of the original video model (Cosmos-Predict2.5), within-video pre-training and WATCH. Robot frames are from the side-right camera of GR1 MimicDroid rollouts; human frames are EgoVerse queries.

Prediction on Unseen Data

Action and video prediction before robot post-training on three datasets not used in pre-training: RoboMIND (robot), Humanoid Everyday (humanoid teleoperation) and EgoDex (human). 2,048 clips per dataset.

Held-out set

Generated future video

Predicted from the current frame. WATCH is also given the reference.

task ℓ

Current frame

Current frame

Reference video

Ground truth

Predicted future

WATCH (Ours)

Within-video

Co-training

Original video model

Attention on an unseen robot (RoboMIND)

Video-slot attention on the reference for a Franka query (bread in basket), with references from the same robot and from a different robot (UR).

Model

Ground truthFranka query

Same robotleft view

Same robotright view

Different robotUR

On RoboMIND, WATCH is best on all metrics. Compared with within-video pre-training, end-effector MSE decreases from 0.046 to 0.039 and accuracy increases from 55.6% to 68.5%. On EgoDex, WATCH is best on all action metrics and on FVD and LPIPS. On Humanoid Everyday, which co-training uses for training, WATCH matches or exceeds co-training on all metrics except FVD, with half its end-effector error.

Human Pairs vs. Robot Pairs

WATCH trained on about 1K hours of human pairs is compared with the same objective trained on about 1K hours of RoboMIND robot pairs. Only the robot-pair model sees RoboMIND in pre-training.

Agreement with the robot-trained model, layer by layer

Procrustes match@1 against the robot-trained model at every DiT block; chance is 1/N (0.0033 and 0.0028).

Match@1 per layer on RoboMIND and EgoVerse windows
Show feature space on robot windows (t-SNE)

Feature space on robot windows

Action-slot features of RoboMIND windows, aligned and embedded jointly with t-SNE, so every window appears three times. Hover a point to link its three copies.

Robot
Original video model Human pairs (WATCH) Robot pairs

Layer 12

Layer 27

After alignment, a window’s features from the human-pair and robot-pair models are 0.03 to 0.06 apart, compared with 0.12 to 0.31 for the original video model. Over layers 10–27, the mean match@1 with the robot-pair model is 0.50 on robot windows and 0.71 on human windows for WATCH, and 0.13 and 0.11 for the original video model.

Acknowledgements

We thank our colleagues at NAVER AI Lab for their valuable feedback and support throughout this work. We also thank Junha Song for helpful discussions.

BibTeX

@article{park2026watch,
  title={Learning Transferable World-Action Models from Task-Paired Human Videos},
  author={Park, Jeongeun and Kim, Taekyung and Han, Dongyoon and Yun, Sangdoo},
  year={2026}
}