← Search

Yusuf Aytar

23 accepted papers

2025

Motion Prompting: Controlling Video Generation with Motion Trajectories

CVPR 2025poster

Motion control is crucial for generating expressive and compelling video content; however, most existing video generation models rely mainly on text prompts for control, which struggle to capture the nuances of dynamic actions and temporal compositions. To this end, we train a video generation model…

Cited by 22SourcePDFScholar
2024

FlexCap: Describe Anything in Images in Controllable Detail

NeurIPS 2024poster

We introduce FlexCap, a vision-language model that generates region-specific descriptions of varying lengths. FlexCap is trained to produce length-conditioned captions for input boxes, enabling control over information density, with descriptions ranging from concise object labels to detailed caption…

2024

Genie: Generative Interactive Environments

ICML 2024oral

We introduce Genie, the first *generative interactive environment* trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketche…

Cited by 172SourcePDFScholar
2024

Learning from One Continuous Video Stream

CVPR 2024poster

We introduce a framework for online learning from a single continuous video stream - the way people and animals learn without mini-batches data augmentation or shuffling. This poses great challenges given the high correlation between consecutive video frames and there is very little prior work on it…

Cited by 3SourcePDFScholar
2024

Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models

NeurIPS 2024spotlight

We address the problem of multi-object 3D pose control in image diffusion models. Instead of conditioning on a sequence of text tokens, we propose to use a set of per-object representations, *Neural Assets*, to control the 3D pose of individual objects in a scene. Neural Assets are obtained by pooli…

Cited by 13SourcePDFScholar
2024

RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation

ICRA 2024poster

For robots to be useful outside labs and specialized factories we need a way to teach them new useful behaviors quickly. Current approaches lack either the generality to onboard new tasks without task-specific engineering, or else lack the data-efficiency to do so in an amount of time that enables p…

Cited by 45SourceScholar
2023

Lossless Adaptation of Pretrained Vision Models For Robotic Manipulation

ICLR 2023poster

Recent works have shown that large models pretrained on common visual learning tasks can provide useful representations for a wide range of specialized perception problems, as well as a variety of robotic manipulation tasks. While prior work on robotic manipulation has predominantly used frozen pre…

Cited by 33SourcePDFScholar
2023

Perception Test: A Diagnostic Benchmark for Multimodal Video Models

NeurIPS 2023poster

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, BEiT-3, or GPT-4). Compared to existing benchmarks that focus on computational tasks (e.g. classification, detection or tracking), th…

2023

TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal Refinement

ICCV 2023poster

We present a novel model for Tracking Any Point (TAP) that effectively tracks any queried point on any physical surface throughout a video sequence. Our approach employs two stages: (1) a matching stage, which independently locates a suitable candidate point match for the query point on every other…

Cited by 337PDFcodeScholar
2022

Learning transferable motor skills with hierarchical latent mixture policies

ICLR 2022spotlight

For robots operating in the real world, it is desirable to learn reusable abstract behaviours that can effectively be transferred across numerous tasks and scenarios. We propose an approach to learn skills from data using a hierarchical mixture latent variable model. Our method exploits a multi-leve…

Cited by 38SourcePDFScholar
2022

TAP-Vid: A Benchmark for Tracking Any Point in a Video

NeurIPS 2022accept

Generic motion understanding from video involves not only tracking objects, but also perceiving how their surfaces deform and move. This information is useful to make inferences about 3D shape, physical properties and object interactions. While the problem of tracking arbitrary physical points on su…

2022

Wish you were here: Hindsight Goal Selection for long-horizon dexterous manipulation

ICLR 2022poster

Complex sequential tasks in continuous-control settings often require agents to successfully traverse a set of ``narrow passages'' in their state space. Solving such tasks with a sparse reward in a sample-efficient manner poses a challenge to modern reinforcement learning (RL) due to the associated…

Cited by 19SourcePDFScholar
2021

With a Little Help From My Friends: Nearest-Neighbor Contrastive Learning of Visual Representations

ICCV 2021poster

Self-supervised learning algorithms based on instance discrimination train encoders to be invariant to pre-defined transformations of the same instance. While most methods treat different views of the same image as positives for a contrastive loss, we are interested in using positives from other ins…

Cited by 561PDFScholar
2020

Counting Out Time: Class Agnostic Video Repetition Counting in the Wild

CVPR 2020poster

We present an approach for estimating the period with which an action is repeated in a video. The crux of the approach lies in constraining the period prediction module to use temporal self-similarity as an intermediate representation bottleneck that allows generalization to unseen repetitions in vi…

Cited by 159PDFcodeScholar
2020

Learning rich touch representations through cross-modal self-supervision

CoRL 2020

The sense of touch is fundamental in several manipulation tasks, but rarely used in robot manipulation. In this work we tackle the problem of learning rich touch features from cross-modal self-supervision. We evaluate them identifying objects and their properties in a few-shot classification setting

2020

Scaling data-driven robotics with reward sketching and batch reinforcement learning

RSS 2020poster

By harnessing a growing dataset of robot experience, we learn control policies for a diverse and increasing set of related manipulation tasks. To make this possible, we introduce reward sketching: an effective way of eliciting human preferences to learn the reward function for a new task. This rewar…

2020

Self-Supervised Sim-to-Real Adaptation for Visual Robotic Manipulation

ICRA 2020poster

Collecting and automatically obtaining reward signals from real robotic visual data for the purposes of training reinforcement learning algorithms can be quite challenging and time-consuming. Methods for utilizing unlabeled data can have a huge potential to further accelerate robotic learning. We co…

Cited by 78SourceScholar
2018

Playing hard exploration games by watching YouTube

NeurIPS 2018spotlight

Deep reinforcement learning methods traditionally struggle with tasks where environment rewards are particularly sparse. One successful method of guiding exploration in these domains is to imitate trajectories provided by a human demonstrator. However, these demonstrations are typically collected un…

Cited by 329SourcePDFScholar
2017

Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

CVPR 2017poster

In this paper, we introduce Recipe1M, a new large-scale, structured corpus of over 1m cooking recipes and 800k food images. As the largest publicly available collection of recipe data, Recipe1M affords the ability to train high-capacity models on aligned, multi-modal data. Accordingly, we train a ne…

Cited by 758PDFScholar
2016

Learning Aligned Cross-Modal Representations From Weakly Aligned Data

CVPR 2016poster

People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer across modalities. To study this problem, we introduce a new cross-modal scene dataset. While convolutional neural networks c…

Cited by 204PDFScholar