Videos

UME: how a robot learns to fetch a drink from the fridge

The can-fetching demo is the visible result of a larger idea: collect whole-arm motion and torque feedback with an exoskeleton, then train a mobile robot policy from that data.

TOMORROW, TESTED. / VIDEO

UME: how a robot learns to fetch a drink from the fridge

Loading the player connects to YouTube. You can also watch directly.
Watch on YouTube ↗

In UME’s fridge task, the research team says its learned policy checks whether a drink is present, turns to the fridge if needed, opens the door, retrieves a full can and places it on the table. The team reports training this task with 157 demonstrations.

What we know

Fridge task

The policy checks for a can, turns to the fridge when none is present, opens the door, retrieves a full unopened can and places it on the table.

Demonstration
Training data

The project reports 157 demonstrations for the fridge drink-retrieval policy.

Verified
UME hardware

UME is an upper-limb exoskeleton that records whole-arm configuration and joint torque signals while providing real-time haptic torque feedback.

Verified
Declared autonomy

The project labels the fridge policy autonomous; this is the authors’ research demonstration rather than independent verification.

Company claim
Reported hardware cost

The project reports about $1,900 for the UME system and $9,533 for its in-house mobile manipulator.

Cost

The Short is about more than opening a fridge

The visible sequence is simple enough to understand immediately. A drink can is removed from the table. The robot turns around, opens the fridge, takes another can from the door, closes the fridge, carries the drink back and places it on the table.

The UME research team labels this a whole-body fridge drink-retrieval autonomous policy. On the project page, the task is defined as maintaining a drink on the table: if no can is detected, the robot is supposed to retrieve one; if the can is removed again, the procedure repeats.

That is what our Short shows. But the more important story sits one level earlier: how did the researchers collect the kind of training data needed for a robot to learn a long, contact-heavy sequence like this?

UME is a data-collection interface

UME stands for Universal Manipulation Exoskeleton. It is an upper-limb exoskeleton designed to record a human operator’s arm configuration and joint torque signals while also giving that operator real-time torque feedback.

That force information matters because household manipulation is not only about where a hand should go. Opening a door, pushing an object against a surface or reaching through a constrained space all involve contact forces.

The authors argue that many robot-demonstration systems capture motion well but do not capture this interaction information well enough. UME is intended to collect both.

The system is also designed to retarget human motion to different robot arms. In the paper and project material, the researchers show teleoperation with OpenArm, Franka and X-ARM systems.

From teleoperation data to an autonomous policy

The fridge Short is not presented by the authors as live teleoperation. It appears in their section of autonomous policy demonstrations.

According to the project, all of those policies were trained using data collected with UME. The fridge task used 157 demonstrations, more than the other listed tasks: 26 demonstrations for visually occluded box pushing, 40 for box flipping and 42 for constrained GPU picking.

The research platform used for the autonomous demonstrations is an in-house dual-arm mobile manipulator.

For the fridge scenario, the robot has to coordinate much more than a single grasp. It must rotate its mobile body, approach the fridge, manipulate the door, pick up a full unopened can, support and transport it, return to the table and place it down.

That is why the authors describe the task as both whole-body mobile manipulation and long-horizon.

What “autonomous” means here

The project team calls the learned fridge behaviour autonomous. Tomorrow, Tested. preserves that wording as an attributed research claim.

What the public demonstration does not establish is that this robot can enter an arbitrary home, identify an unfamiliar fridge and perform the same task reliably without adaptation. It is a learned policy demonstrated on the researchers’ platform and setup.

That distinction is important because a polished video can compress away the difference between:

  • a behaviour working in the environment used for the experiment;
  • a policy being robust to specific disturbances tested by the researchers;
  • and a general household robot that can repeat the behaviour anywhere.

The UME project also includes a separate robustness-to-disturbance demonstration, which is useful evidence, but it is still part of the same research programme rather than an independent deployment study.

Why torque feedback is central to the research

The fridge retrieval is visually easy to appreciate, but some of the project’s other tasks explain why the team focused on torque data.

One demonstration asks the robot to retrieve a GPU from a narrow space between two desktop computers. Another pushes a box into a visually occluded space. Another flips a box using force against a fixed surface.

Those are cases where vision alone does not describe the entire interaction. The way forces change during contact can contain information the robot needs.

The researchers’ thesis is that collecting better physical interaction data makes it possible to learn policies that are more compliant and better suited to constrained manipulation.

The hardware is deliberately relatively low-cost

The project reports a cost of $1,900 for the complete UME system.

It also reports $9,533 for the in-house mobile manipulator used to evaluate the learned policies. The team says the robot includes an OpenArm 1.0 bimanual system, a powered caster base, a DJI Power 1000 and fisheye cameras on the head and wrists.

These are project-reported hardware costs, not retail quotes independently verified by Tomorrow, Tested. They are still useful because they reveal an explicit design goal: make the data-collection system portable and relatively accessible rather than dependent on a unique laboratory rig.

Why this matters for Physical AI

The popular version of the Physical AI story often focuses on the model: bigger policies, more data, more compute.

UME highlights the other side of that equation. Before a robot can learn from physical experience, someone needs a practical way to collect good physical demonstrations — including not just camera images and trajectories but contact information.

The fridge Short is therefore a useful endpoint to watch, but the exoskeleton behind the training data may be the more important technology.

If systems like UME make high-quality demonstrations easier to collect across different robot bodies, the bottleneck for useful robot learning shifts. Better data interfaces can become as important as better policies.

Our Short uses selected excerpts from the UME project demonstration and adds English narration, captions and a vertical edit. The research project and paper below are the primary references for the claims on this page.

Sources & provenance

A first-party source records the maker’s claims; it is not independent verification.

  1. Universal Manipulation Exoskeleton: Learning Compliant Whole-body Policies with Real-time Torque Feedback ↗
    UME research team — Ant Group / Stanford University · Research · EN
    Retrieved
  2. Universal Manipulation Exoskeleton: Learning Compliant Whole-body Policies with Real-time Torque Feedback ↗
    arXiv · Research · EN
    Retrieved · Published