Executive Summary

I want to build a prediction engine that takes as input raw percepts and predicts future “state” of the world. I believe being able to make high quality statements across time is the critical component to understanding action in the world, and a high quality prediction engine is a useful backbone for the planning stack of generally capable embodied agents, from autonomous vehicles to service robots.

I believe that in order for this prediction engine to be effective, it needs to be highly data-driven, an approach that’s been massively successful in the language domain. In service of this, I am searching for the learning problem formulation that produces a prediction engine where its prediction quality scales with compute and data used to train it, without requiring human annotations. This means filling in important low level details: what are these raw percepts? What is this future “state”? Where are we going to get all this data?

My work is trying to answer these important questions. To my mind, a few answers are clear:

However, important questions remain:

There is useful prior art in this general direction: