Speak in actions.
Move, rotate, orient, and grip. Commands are grounded at the gripper interaction point, then executed by the robot’s motion stack.
RoboDawn↗
Explore the research ↗
VLMs are few-shot robotic learners.
The intelligence is already there.
Give it an interface. Show it an example.
Let it act in the physical world.
“Pick up the object. Place it in the tray.”
What if the path to capable robots begins with the intelligence our models already have?
Pretrained vision-language models bring rich visual understanding and reasoning to a new setting: the physical world.
RoboDawn makes robot control legible to a VLM. Simple spatial commands connect its decisions to motion; a few demonstrations show how to use them.
Transfer intelligence.
Teach the interface.
From simulated manipulation to real robot experiments. Choose an episode and follow the agent’s decisions, tool calls, and corrections.
GPT-6 Astra · RoboDojo · simulation recording
Loading recorded session…
Guided replay: pauses are added for reading and recorded motion is slowed. Public decision notes and real tool calls are aligned to control steps; this is a replay, not a live model connection. Camera snapshots are unannotated versions of the recorded views.
Original 10-second clip ↗A frozen VLM, a semantic action interface, and a feedback loop. Each action becomes the context for the next decision.
Move, rotate, orient, and grip. Commands are grounded at the gripper interaction point, then executed by the robot’s motion stack.
A shared command primer explains the interface. Task demonstrations show how those commands compose into useful behavior.
Observations, execution feedback, and memory evolve. The model’s weights stay fixed. Adaptation happens inside the interaction.
One task demonstration helps turn pretrained multimodal capability into effective robotic manipulation.
Success rate (%) · RoboTwin C2R
¹ Results reported in the research manuscript. C2R evaluates clean-demonstration transfer under domain randomization. Full-set baselines use 50 demonstrations per task across 50 tasks. “0-shot” denotes zero task demonstrations; a shared command primer provides interface examples. GPT-6 Astra is reported at 1-shot only.
An accessible action interface. A few in-context examples.
A complementary path toward general-purpose manipulation.