StereoEngine: Scaling Stereo Bimanual Robot Data without Stereo Cameras

Anonymous Authors

SCROLL

Why Use Stereo Cameras?

1

Better Depth Estimation Compared to Regular Depth Cameras

2

Works Better on Transparent and Highly Reflective Objects

Scene Comparison of Transparent Bottles

RealSense

Conventional depth camera

Preparing point cloud…

ZED + FoundationStereo

Stereo reconstruction

Preparing point cloud…

Robot Training Data Is Overwhelmingly Monocular

Large Scale Monocular Datasets

Mostly single-view RGB

Open X-Embodiment BridgeData V2 DROID RoboNet

StereoEngine

Left and Right View RGB

StereoEngine

We need a way to scale stereo dataset by (1) leveraging existing monocular datasets or (2) synthetically generating both views.

StereoEngine-L: Stereo Generation with a Ground-Truth Left View

Left Video

Warping
Module

Warped Right Video

Language Instruction

Masked Right Video

Video
Diffusion

Generated Right Video

Warping Module

Moves every left pixel to its position in the right view, giving a warped right video and a mask marking the pixels no left pixel could reach.

Video Diffusion

Fills the masked region using the warped video, mask, and language instruction.

Stereo Pair

Produces a right video that pairs with the original ground-truth left video.

StereoEngine-L Data Examples

Example 1

Left View
Right View

Example 2

Left View
Right View

Example 3

Left View
Right View

Example 4

Left View
Right View

Example 5

Left View
Right View

Example 6

Left View
Right View

StereoEngine-S: Stereo Generation without a Ground-Truth Left View

Left Simulation Video

Canny Edge
Converter

Left Canny Video

Right Simulation Video

Right Canny Video

Left and right reference images

Reference Images

Left + Right
Canny Videos
Language Instruction
Video
Diffusion

Generated Left Video

Generated Right Video

Two-Camera Simulation

Render the episode in simulation from two cameras, then convert both views into Canny-edge videos.

Geometry + Appearance

Pass both edge videos to the model with two reference images and a language instruction. Edges define geometry; the references supply background, lighting, and surface texture.

Stereo Generation

A video generation model produces the synthetic left and right videos as a stereo pair.

StereoEngine-S Data Examples

Stack Two Bowls

Left View
Right View

Lift Two Bottles

Left View
Right View

Container Plate

Left View
Right View

Successful Real-World Rollouts

Stack Two Bowls

Lift Transparent Bottle

Place Plate in Container

Close Door

Lift Bottles

Real-World Policy Results

We run each task 20 times on the real robot. StereoEngine-L uses generated stereo data to mid-train ExStereo, while StereoEngine-S uses task-specific generated data to train ExStereo and StereoVLA.

Blue and purple: StereoEngine (Ours) Gray, orange, cyan, and pink: baselines

StereoEngine-L: Left → Right Generation

BibTeX