Reward Anything

Real world RL does not work well for long horizon robot learning tasks.

I think the problem with most robotics RL approaches is that straightforward, end-of-episode reward models provide insufficiently dense signal – the signal becomes so diluted when stretched over long time horizons that modeling it with a network becomes difficult. This is why today, to get RL to work well we have to either hand-craft heuristic reward functions or select an artificially short-horizon task.

To try to address this horizon issue in general, I propose a three-part approach:

Early evidence from similar approaches (without scaling up) like ReWiND, REDS, SARM, and SARM2 indicate that good subtask reward densification can lead to better, more robust policies.

Premises and priors

Execution Plan

Step 1: Breaking long-horizon demonstrations down into subtasks

What do we want?

We want to break long, multi-minute demonstrations down into short subtasks.

The form of this annotation is a time span plus a language description — (start time, end time, language description) — across enormous quantities of long-horizon data. Notably, these spans do not need to be non-overlapping.

Importantly, we want demonstrations to contain both successes and failures for given subtasks to ensure failures are in-distribution.

How do we get it?

I think this is mostly an execution question.

We need mixed-quality data that has subtasks that are both successes and failures. I think ReWiND style mislabeling is a hack to get started, but truly disastrous trajectories are still out of distribution.

These annotations can be sourced from human annotation, or potentially from a sufficiently strong embodied video understanding model (e.g. a better version of Perceptron may be able to do this). If an existing video understanding model cannot do this, we should do a SAM-style human-plus-pseudolabel bootstrapping approach to produce a model that can.

Step 2: Large-scale training of language-conditioned short-horizon reward model

What do we want?

Using these diverse, large-scale (start time, end time, language description) annotations, we can train a general language-conditioned value function: V(x_t, z)

If we have action labels, we can also train a Q function Q(x_t, a, z)

How do we get it?

The details of the training objective are an open research question.

For value estimation I want to try:

For state representation x_t I want to try:

But at scale this ought to produce a general, short-horizon reward model.

Step 3: Offline and Online RL with short-horizon tasks

Given that we now have

We can explore offline algorithms first that we can incorporate into large-scale policy or world model training. In terms of likelihood to work and add value:

If the value function is causal, we can also explore online, on-policy approaches (e.g. RLT), and use that to produce policy rollouts we can add to the large-scale training corpus.