从第三人称人类演示视频学习机器人执行
Learning Robot Execution from Third-Person Human Demonstration Videos
两个 demo 都来自 H&R v1 同一 HDF5 episode 的原生逐帧配对,并特意选择了 dataset report 未展示的任务与 episode。这里展示的是 ground-truth task contract,不是模型生成结果。
Demo 1 · Place both cubes
data/v1/grab_both_cubes_v1/episode_0.hdf5 · 418 frames · Class A native pair · not shown in report
③ Prompt
Observe the third-person human demonstration. Generate the synchronized robot execution video that places both cubes into the tray and predict the corresponding end-effector and gripper trajectory.
prompt.txt④ Metadata
pairing = source HDF5 path + frame index
human = cam_data/human_camera
robot = cam_data/robot_camera
action = mapped human pose in robot frame [T,7]
robot state = end_position + gripper + qpos/qvel
Demo 2 · Draw a circle
data/v1/writing/writing_circle/episode_0.hdf5 · 407 frames · Class A native pair · not shown in report
③ Prompt
Observe the third-person human demonstration. Generate the synchronized robot execution video that draws the demonstrated circular trajectory and predict the corresponding end-effector trajectory.
prompt.txt④ Metadata
pairing = source HDF5 path + frame index
human = cam_data/human_camera
robot = cam_data/robot_camera
action = mapped human pose in robot frame [T,7]
robot state = end_position + gripper + qpos/qvel
Demo boundary
派生部分:prompt、独立 MP4 文件与本页布局属于 demo layer;MP4 只做无 crop/resize 的 H.264 浏览器转码。
正确评价:视频阶段/物体状态 + end-effector trajectory + gripper accuracy;不能只看 PSNR/SSIM。
Source: dannyXSC/HumanAndRobot v1 @ 79c1b0e5f488567aa8fea87f98c1c8a286227cc1 · source HDF5 SHA-256 recorded in each metadata.json · raw report FULL VERIFIED.