Today I am reading GR00T N1: An Open Foundation Model for Generalist Humanoid Robots. I came to this after reading Diffusion Policy, so my first question was simple: what changes when an action-generation policy becomes part of a robot foundation model?
This is an English adaptation of my Korean reading notes. The scope is the N1 paper: its introduction, architecture, training data, evaluation, and limitations. Later GR00T releases deserve their own reading rather than being folded into the same model name.
1. The problem: data islands
A robot's body is part of its learning problem. Different robots have different joints, action dimensions, sensors, and camera placements. A demonstration collected on one platform cannot simply be treated as a demonstration on another.
The paper calls attention to these separated pools of experience. Robot data is already expensive; splitting it by embodiment makes each pool smaller still. I found this a useful starting point because it changes what we should look for in the model. The question is also how to make more experience usable together.
2. Two systems, one policy
GR00T N1 separates vision-language understanding, called System 2, from action generation, called System 1. System 2 processes the image and instruction. System 1 uses that representation to generate robot actions. The modules are coupled and trained together.
The names suggest slow understanding and fast execution. That is a helpful analogy, but it is not a claim that the network reproduces human cognition.
System 2: a representation for control
The paper uses Eagle-2 as its vision-language backbone. One choice caught my attention in the Korean reading: it takes an intermediate representation rather than assuming the last language-model layer must always be the best feature for control.
I still find that choice less intuitive than it first sounds. A feature useful for producing text need not be the feature that best supports a downstream motor policy. The practical question is the policy's behavior and computation time, not which layer feels most complete.
System 1: generating an action chunk
The action module uses a diffusion transformer and flow-matching training. It predicts a chunk of future actions, so the connection to Diffusion Policy is easy to see: the output is a short trajectory rather than a single isolated command.
The two systems can also operate at different rates. A robot does not necessarily need to reinterpret the full instruction every time it issues another motor command. I read this as a connection to the distinction between observing, predicting a horizon, and executing part of that horizon.
3. A pyramid of training data
This was the most interesting part of the paper for me. Its training material includes real robot demonstrations, human video, and synthetic trajectories. Each source contributes something different, and each comes with a different gap to bridge.
Human video and latent actions
Human video does not contain the robot's joint commands. The paper learns a latent representation of the change between frames and uses those inferred actions in training. The representation provides supervision from visual change; it should not be confused with measured robot controls.
That distinction matters. Calling both things “actions” makes the pipeline sound simpler than it is. One is an observed command in a particular body; the other is a learned description of what happened between images.
Generated video and simulated trajectories
The neural-trajectory pipeline expands available experience with generated video and filters the results. Simulation provides another route: demonstrations can be transformed and replayed under changed object arrangements.
These are different sources of synthetic experience. A generated video may look convincing without respecting all physical constraints. A simulation trajectory depends on the simulator and its task assumptions. More examples are useful only to the extent that those examples teach the policy something transferable.
Sharing a model across bodies
Embodiment-specific mappings convert states and actions into a shared representation. This lets the large model share learned structure while smaller components account for differences between robot bodies.
The training sequence then moves from broad pre-training to adaptation for a target robot. In my notes, I think of this as learning broadly and then narrowing the policy to the body and tasks where it will actually run.
4. What the evaluation tells me
The paper evaluates both simulated tasks and real manipulation on the Fourier GR-1. My Korean notes compare its results with Diffusion Policy and BC-Transformer, including experiments with reduced demonstration data.
The point that stayed with me was data efficiency. A strong result with less target-robot data is particularly relevant when collecting that data is the expensive step. I would still read each comparison together with its task set, training data, and evaluation conditions; a benchmark result is not a guarantee on an arbitrary robot.
See Section 4 of the paper for the numerical tables and experimental settings. This adaptation keeps the interpretation separate from those reported measurements.
5. Limitations
N1 focuses on relatively short tabletop manipulation. Walking while manipulating objects, extended tasks, and broader whole-body behavior require more work. Synthetic experience also inherits limitations from its generator or simulator.
I appreciated that the paper leaves these boundaries visible. “Generalist” is a direction for the model, not evidence that every kind of robot behavior has already been solved.
My takeaway
Reading this after Diffusion Policy helped me see the continuity. Generating a sequence of actions remains central. What expands is the conditioning information, the range of bodies, and the effort spent bringing different kinds of experience into the training process.
That is the question I want to carry into the next reading: when a robot foundation model improves, how much comes from the policy architecture, and how much comes from the experience it can now learn from?