I am a Senior Staff Research Scientist at Google DeepMind where I work on multi-modal generative models and world models.
I received my PhD from the Computer Science & Engineering Department at the University of Michigan, Ann Arbor under the supervision of Professor Honglak Lee. During my PhD, I mainly focused on building conditional video generation models and contributed to one of the first world models from pixels successfully used to improve sample-efficiency in model-based reinforcement learning.
Fun Facts: I played for my national basketball team (I am originally from Ecuador). I was also second best scorer in the nation in a national championship I played back in the day. I was part of a team that beat the media's projected champion during a championship in Quito (the guy that was best scorer in the national championship played for the other team :P). Let's have a Curry-range 3-point shootout. Ok, I'll stop now ...
A selection of my most impactful research. Full list on Google Scholar.
Pioneered causal spatio-temporal video tokens and masked transformers to synthesize variable-length video stories from sequential text prompts.
One of the first world models from pixels — agents learn complex behaviors with high sample efficiency by planning entirely inside learned video prediction models.
Demonstrated that scaling up stochastic video prediction networks dramatically improves quality — state-of-the-art across multiple driving and human motion benchmarks.
Neural kinematic chains for skeleton-aware motion transfer between characters without paired supervision.
Proposed predicting high-level structure (pose) first, then generating pixels — enabling much longer-horizon video generation.
Factorized video into motion and content streams for more structured and accurate natural video prediction.