RoboDawn Explore the research
A NEW DAWN FOR EMBODIED INTELLIGENCE

Intelligence,
embodied.

VLMs are few-shot robotic learners.

The intelligence is already there.
Give it an interface. Show it an example.
Let it act in the physical world.

Frozen VLM+Intuitive interface+In-context learning
EMBODIED INTELLIGENCE LABFIG. 01
WORLD FRAME
x · y · z
GRIPPER INTERACTION POINT
01 / OBSERVEFROZEN VLM

“Pick up the object. Place it in the tray.”

VISUAL INPUT → SEMANTIC ACTION● ○ ○ ○
INTERACTIVE CONCEPT DEMO
73.6%
One demonstration.RoboTwin C2R success rate¹
0
Parameter updates.Pretrained intelligence, kept intact
+15.2pp
A single example changes things.0 → 1 shot, same Gemini backbone
01 / THE INSIGHTINTELLIGENCE TRANSFER

A new body.
The same intelligence.

What if the path to capable robots begins with the intelligence our models already have?

DIGITAL WORLDPerceive.
Understand.
Reason.
RoboDawn
AN INTERFACE TO ACT
PHYSICAL WORLDReach.
Grasp.
Manipulate.

Pretrained vision-language models bring rich visual understanding and reasoning to a new setting: the physical world.

RoboDawn makes robot control legible to a VLM. Simple spatial commands connect its decisions to motion; a few demonstrations show how to use them.

Transfer intelligence.
Teach the interface.

02 / WATCH THE AGENTRECORDED EPISODE · ROBODOJO

Watch intelligence
become action.

From simulated manipulation to real robot experiments. Choose an episode and follow the agent’s decisions, tool calls, and corrections.

GPT-6 Astra · RoboDojo · simulation recording

TASK / ALIGN BLOCKS

One ruler. Three blocks. A closed loop.

SESSION REPLAY
HEAD CAMERAUNANNOTATED FOOTAGE
Use the ruler to align the three blocks.
LAST OBSERVATIONAwaiting camerasClean camera snapshots
LEFT WRIST
RIGHT WRIST
RIGHT TCP — / — / — cmSTEPS 0 / 500
R↗
Agent workspaceGPT-6 Astra · public session

Loading recorded session…

0:00 / —

Guided replay: pauses are added for reading and recorded motion is slowed. Public decision notes and real tool calls are aligned to control steps; this is a replay, not a live model connection. Camera snapshots are unannotated versions of the recorded views.

Original 10-second clip
Explore all tool calls Exact arguments & execution feedback
03 / THE APPROACHSMALL INTERFACE. OPEN POSSIBILITIES.

See. Think. Act.
And see again.

A frozen VLM, a semantic action interface, and a feedback loop. Each action becomes the context for the next decision.

ROBODAWN / CONTROL LOOP● OBSERVATION
θ FROZEN THROUGHOUT THE EPISODE↻ CLOSED LOOP
01 — INTERFACE

Speak in actions.

Move, rotate, orient, and grip. Commands are grounded at the gripper interaction point, then executed by the robot’s motion stack.

02 — DEMONSTRATIONS

Show, then generalize.

A shared command primer explains the interface. Task demonstrations show how those commands compose into useful behavior.

03 — ADAPTATION

Learn in context.

Observations, execution feedback, and memory evolve. The model’s weights stay fixed. Adaptation happens inside the interaction.

04 / THE EVIDENCEROBOTWIN 2.0 · CLEAN → RANDOMIZED

A little context.
A remarkable leap.

One task demonstration helps turn pretrained multimodal capability into effective robotic manipulation.

SUCCESS RATE (%)

Few-shot, real gains.

Success rate (%) · RoboTwin C2R

Trained-policy baselines
TASK DEMONSTRATIONS

¹ Results reported in the research manuscript. C2R evaluates clean-demonstration transfer under domain randomization. Full-set baselines use 50 demonstrations per task across 50 tasks. “0-shot” denotes zero task demonstrations; a shared command primer provides interface examples. GPT-6 Astra is reported at 1-shot only.

05 / THE RESEARCHROBODAWN
VLMs ARE FEW-SHOT ROBOTIC LEARNERS

The next frontier
is physical.

An accessible action interface. A few in-context examples.
A complementary path toward general-purpose manipulation.

PaperComing soon
GitHub codeComing soon
PAPER & CODE · COMING SOON

Research framework

RoboDawn research framework