World-action models inherit broad visual and physical priors from video pre-training. However, standard within-video next-frame prediction allows them to rely on appearance shared between the context and future, rather than learning interactions that transfer across scenes and embodiments. We introduce WATCH, a cross-video pre-training that learns from task-paired human demonstrations collected across diverse environments. Conditioned on the video and actions of one demonstration, the model jointly predicts another demonstration’s actions and future frames from its current observation. Predicting one demonstration from another encourages the model to learn task-relevant interactions that transfer across scenes and demonstrators.
Our pre-training shows that these capabilities can be learned from readily scalable human–human pairs alone, without requiring matched human–robot demonstrations. Analyses before robot post-training show improved motion alignment across embodiments and better action and video prediction on unseen robots, with representations closer to those learned from robot pairs. After robot post-training, WATCH improves MimicDroid in-context success from 40.9% to 68.3% with Cosmos Policy and from 59.8% to 77.7% with Cosmos3, with the largest gains on the unseen levels. Gains persist under context-free post-training across robot embodiments. WATCH outperforms within-video pre-training and co-training baselines, supporting cross-video prediction as a scalable approach to learning transferable world-action representations.
(a) Human videos of the same task are paired across collection sources and scenes. (b) Given the reference video and its actions, the model predicts the target’s actions and future frames from the target’s initial frame. (c) After robot post-training, WATCH improves success rates over the base model in both in-context and context-free settings.
Context and target come from the same demonstration and share its background, viewpoint and demonstrator.
The reference comes from a different demonstration of the same task, and the query’s current frame defines the scene to predict in.
Human demonstrations collected by different people, with different rigs and in different environments, vary along the same axes on which a robot differs from a human. We pair demonstrations of the same task across these collection sources, without any robot data. Query–reference pairs are built from different episodes of EgoVerse, with human actions expressed in a shared human–robot task space.
| Query–reference pairs | Paired 3-s windows | References per query | Queries with a cross-source reference |
|---|---|---|---|
| 10.99M | 10.0K h | 10.52 (6.35 cross-source) | 90.7% |
Query windows and their retrieved references, ranked by motion distance. Hand trajectories are overlaid.
Query
Retrieved references
WATCH follows the joint video–action formulation of Cosmos Policy. The reference video and actions, the query’s current frame and the proprioceptive state are clean conditioning inputs; the query’s action chunk and future frames are noised and denoised by the model. The same layout is used in robot post-training, either with a human demonstration as the reference or with the reference slots removed.
After pre-training on human pairs, the model is post-trained on the GR1 humanoid and conditioned on a human demonstration at test time. MimicDroid has three levels: L1 (seen object, seen environment), L2 (unseen object, seen environment) and L3 (unseen object, unseen environment). All rollouts use Cosmos Policy (Cosmos-Predict2.5 backbone).
WATCH rollouts and the human reference, from three cameras.
Human reference
Rear
Left
Right
WATCH rollout
Rear
Left
Right
Cosmos Policy without pre-training, with within-video pre-training and with WATCH, on the same evaluation episode and reference. One episode per task.
Human reference
Without pre-training
Within-video
WATCH (Ours)
The human reference is played at 2× speed in post-training and evaluation. One episode per task.
Human reference (2×)
Without pre-training
Within-video
WATCH (Ours)
Context-free post-training on a bimanual ALOHA robot. One episode per task.
Without pre-training
Within-video
WATCH (Ours)
Context-free post-training on a single-arm Franka with 50 demonstrations per task. One episode per task.
Without pre-training
Within-video
WATCH (Ours)
MimicDroid (GR1 humanoid) with the reference slots removed in post-training and evaluation. One episode per task.
Without pre-training
Within-video
WATCH (Ours)
MimicDroid success (%) against the collection budget, in units of the full human corpus. Each panel has its own y range.
We analyze the pre-trained model before robot post-training, on held-out EgoVerse pairs, RoboMIND and MimicDroid observations.
Action MSE by reference
137 held-out episodes · mean ± SEM · lower is better
Holm-corrected p, episode-paired Wilcoxon test.
Attention for the same query under two references
WATCH: action-slot attention · Original video model: video-slot attention
Ground truth
Prediction
Attention, current frame
Attention, reference
SameThe same query is conditioned on a matched reference, a similar-motion reference from a different collection source, and a different-task reference from the same source. Action MSE is 0.0322, 0.0335 and 0.0367, respectively; only the task change is significant. With a reference from another source, attention stays on the hands and the manipulated object.

(a) Retrieval, human

(b) Retrieval, robot

(c) Held-out embodiment regression
On 359 held-out EgoVerse windows and 396 RoboMIND windows from three robots, WATCH retrieves the matching motion from a different human source (a) and a different robot (b) at higher rank than the original video model and within-video pre-training. A regressor trained on two embodiments and tested on the third (c) reaches a Spearman correlation of 0.18 (human) and 0.24 (robot) from WATCH features, compared with 0.06 and 0.01 for the original video model.
Action slot
Video slot
All three models use cross-video pairs and differ only in the prediction target: actions, future frames, or both. Action-only attention peaks near the image boundaries, and video-only attention spreads over the scene. With joint supervision, both action-slot and video-slot attention concentrate on the hand and the object. Frames are from MimicDroid, which is not seen during pre-training.
Current-frame attention of the original video model (Cosmos-Predict2.5), within-video pre-training and WATCH. Robot frames are from the side-right camera of GR1 MimicDroid rollouts; human frames are EgoVerse queries.
Action and video prediction before robot post-training on three datasets not used in pre-training: RoboMIND (robot), Humanoid Everyday (humanoid teleoperation) and EgoDex (human). 2,048 clips per dataset.
Generated future video
Predicted from the current frame. WATCH is also given the reference.
task ℓ
Current frame
Reference video
Ground truth
Predicted future
WATCH (Ours)
Within-video
Co-training
Original video model
Attention on an unseen robot (RoboMIND)
Video-slot attention on the reference for a Franka query (bread in basket), with references from the same robot and from a different robot (UR).
Ground truthFranka query
Same robotleft view
Same robotright view
Different robotUR
On RoboMIND, WATCH is best on all metrics. Compared with within-video pre-training, end-effector MSE decreases from 0.046 to 0.039 and accuracy increases from 55.6% to 68.5%. On EgoDex, WATCH is best on all action metrics and on FVD and LPIPS. On Humanoid Everyday, which co-training uses for training, WATCH matches or exceeds co-training on all metrics except FVD, with half its end-effector error.
WATCH trained on about 1K hours of human pairs is compared with the same objective trained on about 1K hours of RoboMIND robot pairs. Only the robot-pair model sees RoboMIND in pre-training.
Agreement with the robot-trained model, layer by layer
Procrustes match@1 against the robot-trained model at every DiT block; chance is 1/N (0.0033 and 0.0028).
Feature space on robot windows
Action-slot features of RoboMIND windows, aligned and embedded jointly with t-SNE, so every window appears three times. Hover a point to link its three copies.
Layer 12
Layer 27
After alignment, a window’s features from the human-pair and robot-pair models are 0.03 to 0.06 apart, compared with 0.12 to 0.31 for the original video model. Over layers 10–27, the mean match@1 with the robot-pair model is 0.50 on robot windows and 0.71 on human windows for WATCH, and 0.13 and 0.11 for the original video model.
We thank our colleagues at NAVER AI Lab for their valuable feedback and support throughout this work. We also thank Junha Song for helpful discussions.
@article{park2026watch,
title={Learning Transferable World-Action Models from Task-Paired Human Videos},
author={Park, Jeongeun and Kim, Taekyung and Han, Dongyoon and Yun, Sangdoo},
year={2026}
}