Figure says Helix 2.5 generalizes across 30 unseen homes — what the test actually shows
Figure reports that a single Helix 2.5 foundation model was adapted to tidying, towel folding and bed making, then evaluated across 30 homes with no data collected or fine-tuning in those homes.
The strongest result is not a single household demo: Figure reports 56% zero-shot task success with Index pretraining versus 9% from scratch under matched downstream conditions. The evaluation is detailed, but it remains a company-run test.
What we know
- Evaluation scope
Figure evaluated three whole-body behaviors across 30 Bay Area homes that were not used for environment-specific data collection or adaptation. ↗
Company claim- Tasks
The evaluated behaviors were living-room tidying, towel folding and bed making. ↗
Verified- Pretraining result
Figure reports 56% zero-shot success with Index pretraining versus 9% for a policy trained from scratch with the same task-specification data. ↗
Company claim- Data efficiency
Figure says Helix 2.5 matched a representative Helix 02 behavior using half as much task-specific adaptation data while generalizing across 30 unseen homes. ↗
Company claim- Independence
The evaluation protocol and results were published by Figure; Tomorrow, Tested. has not independently replicated the 30-home study. ↗
Limitation
The headline is not “a humanoid can make a bed”
Figure has shown household behaviors before. The more important claim in Helix 2.5 is that the company took the same learned behaviors into 30 homes that were not used to collect environment-specific training data, without fine-tuning the model after arrival.
The three evaluated behaviors were living-room tidying, towel folding and bed making. Together they require the robot to move its whole body, manipulate rigid and deformable objects and reposition itself when its first approach is not enough.
That is a much harder question than performing one polished demo in one prepared room.
What Figure means by zero-shot
“Zero-shot” here applies to the evaluation homes and manipulated objects.
Figure says it collected no data in the evaluation homes, did not include the evaluation toys, towels or bedding in task-specification data, and did not update model weights using rollout data from those homes.
The tasks themselves were not invented at deployment time. Helix 2.5 was adapted to the three behaviors using task-specific data collected elsewhere.
That distinction matters. This is evidence about environment and object generalization, not a claim that the robot learned bed making from a text prompt with no task training.
The 56% vs 9% experiment is the most informative part
Figure trained two policies with the same task-specification data and held architecture, optimization and evaluation fixed.
One started from random weights. The other started from the Index-pretrained Helix 2.5 model.
Figure reports 9% zero-shot success for the scratch model and 56% for the Index-pretrained model. Success required completing the whole task; partial completion did not count.
Because pretraining was the experimental variable Figure changed in that comparison, the result is meant to isolate the value of broad human-behavior pretraining.
It is also a company-run experiment. The dataset, robot, evaluation environments and grading pipeline are controlled by Figure, so the result should be treated as strong first-party evidence rather than independent replication.
Why Index changes the story
Figure’s Index project is designed to collect large amounts of human-behavior video across many environments.
The thesis is familiar from foundation models: learn broadly first, then use smaller amounts of task-specific data to define a behavior.
Helix 2.5 is Figure’s attempt to show that this logic can transfer from language and vision into whole-body robot control.
The company says Index pretraining allowed a representative behavior to match the success rate of a previous Helix 02 policy while using half as much task-specific adaptation data, even though the new behavior was evaluated across 30 unseen homes.
Evaluation rules are unusually important here
Figure published fairly detailed pass/fail criteria.
For living-room tidying, all 13–15 scattered toys had to end up in the basket. For towel folding, towels had to be folded and placed in the basket. For bed making, the pillows and comforter had placement and alignment criteria.
Figure also states that a rollout was aborted and failed if human intervention was required for safety.
Publishing those rules does not make the study independent, but it makes the claim much more interpretable than a montage with no failure definition.
What Helix 2.5 still does not prove
A 56% success result is not solved home robotics. Figure itself says general humanoid robotics is not solved.
Thirty homes are substantially more diverse than one lab kitchen, but they are still a bounded evaluation. The tests cover three trained behavior families, not arbitrary household instructions.
The study also does not answer consumer questions such as service reliability, cost, safety certification or how often a person would need to rescue the robot over months of use.
Why this is worth tracking
Humanoid robotics has spent years producing visually impressive single-environment demos.
Helix 2.5 moves the evidence toward a more useful benchmark: can one policy survive changes in room layout, surfaces, objects and available space without collecting new data in every house?
If the effect of broad human pretraining continues to scale, the economics of deploying home robots could change significantly. Instead of teaching each location separately, more of the adaptation cost would move into pretraining.
That is the real significance of Figure’s announcement — and also why independent replications and broader task suites will matter next.
Sources & provenance
A first-party source records the maker’s claims; it is not independent verification.
- Helix 2.5: Zero-Shot 30-Home Generalization ↗Figure · First-party source · ENRetrieved · Published
- Introducing Index: Building The World’s Largest and Most Diverse Physical Dataset ↗Figure · First-party source · ENRetrieved · Published
- Figure 03 ↗Figure · First-party source · ENRetrieved