← Search

Michael S. Ryoo

58 accepted papers

2026

IVRA: Improving Visual-Token Relations for Robot Action Policy with Training-Free Hint-Based Guidance

ICRA 2026poster

Many Vision-Language-Action (VLA) models flatten image patches into a 1D token sequence, weakening the 2D spatial cues needed for precise manipulation. We introduce IVRA, a lightweight, training-free method that improves spatial understanding by exploiting affinity hints already available in the mod…

2026

MotionV2V: Editing Motion in a Video

CVPR 2026

While generative video models have achieved remarkable fidelity and consistency, applying these capabilities to video editing remains a complex challenge. Recent research has extensively explored motion controllability as a means to enhance text-to-video generation or image animation; however, we id

Cited by 0SourcecodeScholar
2026

Pixel Motion Diffusion is What We Need for Robot Control

CVPR 2026

We present DAWN (Diffusion is All We Need for robot control), a unified diffusion-based framework for language-conditioned robotic manipulation that bridges high-level motion intent and low-level robot action via structured pixel motion representation. In DAWN, both the high-level and low-level cont

Cited by 0SourcecodeScholar
2025

Adaptive Caching for Faster Video Generation with Diffusion Transformers

ICCV 2025poster

Generating temporally-consistent high-fidelity videos can be computationally expensive, especially over longer temporal spans. More-recent Diffusion Transformers (DiTs)--- despite making significant headway in this context--- have only heightened such challenges as they rely on larger models and hea…

Cited by 0SourcePDFScholar
2025

LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback

ACL 2025finding

Large Action Models (LAMs) for AI Agents offer incredible potential but face challenges due to the need for high-quality training data, especially for multi-steps tasks that involve planning, executing tool calls, and responding to feedback. To address these issues, we present LAM SIMULATOR, a compr…

Cited by 0SourcePDFScholar
2025

LLaRA: Supercharging Robot Learning Data for Vision-Language Policy

ICLR 2025poster

Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting a pretrained VLM for robotic control remains challenging, particularly when constrained by a limited number of robot demonstrations. In this…

2025

Language Repository for Long Video Understanding

ACL 2025finding

Language has become a prominent modality in computer vision with the rise of LLMs. Despite supporting long context-lengths, their effectiveness in handling long-term information gradually declines with input length. This becomes critical, especially in applications such as long-form video understand…

2025

Understanding Long Videos with Multimodal Language Models

ICLR 2025poster

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that…

2024

CoPT: Unsupervised Domain Adaptive Segmentation using Domain-Agnostic Text Embeddings

ECCV 2024poster

"Unsupervised domain adaptation (UDA) involves learning class semantics from labeled data within a source domain that generalize to an unseen target domain. UDA methods are particularly impactful for semantic segmentation, where annotations are more difficult to collect than in image classification.…

2024

Crossway Diffusion: Improving Diffusion-based Visuomotor Policy via Self-supervised Learning

ICRA 2024poster

Diffusion models have been adopted for behavioral cloning in a sequence modeling fashion, benefiting from their exceptional capabilities in modeling complex data distributions. The standard diffusion-based policy iteratively denoises action sequences from random noise conditioned on the input states…

Cited by 29SourcecodeScholar
2024

Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs

CVPR 2024poster

Integration of Large Language Models (LLMs) into visual domain tasks resulting in visual-LLMs (V-LLMs) has enabled exceptional performance in vision-language tasks particularly for visual question answering (VQA). However existing V-LLMs (e.g. BLIP-2 LLaVA) demonstrate weak spatial reasoning and loc…

Cited by 22SourcePDFScholar
2024

MAGICK: A Large-scale Captioned Dataset from Matting Generated Images using Chroma Keying

CVPR 2024poster

We introduce MAGICK a large-scale dataset of generated objects with high-quality alpha mattes. While image generation methods have produced segmentations they cannot generate alpha mattes with accurate details in hair fur and transparencies. This is likely due to the small size of current alpha matt…

2024

Mirasol3B: A Multimodal Autoregressive Model for Time-Aligned and Contextual Modalities

CVPR 2024poster

One of the main challenges of multimodal learning is the need to combine heterogeneous modalities (e.g. video audio text). For example video and audio are obtained at much higher rates than text and are roughly aligned in time. They are often not synchronized with text which comes as a global contex…

Cited by 23SourcePDFScholar
2024

VicTR: Video-conditioned Text Representations for Activity Recognition

CVPR 2024poster

Vision-Language models (VLMs) have excelled in the image-domain--- especially in zero-shot settings--- thanks to the availability of vast pretraining data (i.e. paired image-text samples). However for videos such paired data is not as abundant. Therefore video-VLMs are usually designed by adapting p…

Cited by 27SourcePDFScholar
2023

Language-based Action Concept Spaces Improve Video Self-Supervised Learning

NeurIPS 2023poster

Recent contrastive language image pre-training has led to learning highly transferable and robust image representations. However, adapting these models to video domain with minimal supervision remains an open problem. We explore a simple step in that direction, using language tied self-supervised le…

Cited by 17SourcePDFScholar
2023

Open-vocabulary Queryable Scene Representations for Real World Planning

ICRA 2023poster

Large language models (LLMs) have unlocked new capabilities of task planning from human instructions. However, prior attempts to apply LLMs to real-world robotic tasks are limited by the lack of grounding in the surrounding scene. In this paper, we develop NLMap, an open-vocabulary and queryable sce…

Cited by 209SourcecodeScholar
2023

RT-1: Robotics Transformer for Real-World Control at Scale

RSS 2023poster

By transferring knowledge from large, diverse, task-agnostic datasets, modern machine learning models can solve specific downstream tasks either zero-shot or with small task-specific datasets to a high level of performance. While this capability has been demonstrated in other fields such as computer…

2023

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

CoRL 2023poster

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions a…

Cited by 1068SourceScholar
2023

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

ICLR 2023top-25%

We investigate how multimodal prompt engineering can use language as the intermediate representation to combine complementary knowledge from different pretrained (potentially multimodal) language models for a variety of tasks. This approach is both distinct from and complementary to the dominant par…

2023

Token Turing Machines

CVPR 2023poster

We propose Token Turing Machines (TTM), a sequential, autoregressive Transformer model with memory for real-world sequential visual understanding. Our model is inspired by the seminal Neural Turing Machine, and has an external memory consisting of a set of tokens which summarise the previous history…

2023

Weakly-Guided Self-Supervised Pretraining for Temporal Activity Detection

AAAI 2023technical

Temporal Activity Detection aims to predict activity classes per frame, in contrast to video-level predictions in Activity Classification (i.e., Activity Recognition). Due to the expensive frame-level annotations required for detection, the scale of detection datasets is limited. Thus, commonly, pre…

2022

Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?

NeurIPS 2022accept

We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervis…

Cited by 35SourcePDFScholar
2022

Hybrid Random Features

ICLR 2022poster

We propose a new class of random feature methods for linearizing softmax and Gaussian kernels called hybrid random features (HRFs) that automatically adapt the quality of kernel estimation to provide most accurate approximation in the defined regions of interest. Special instantiations of HRFs lead…

2022

Learning Viewpoint-Agnostic Visual Representations by Recovering Tokens in 3D Space

NeurIPS 2022accept

Humans are remarkably flexible in understanding viewpoint changes due to visual cortex supporting the perception of 3D structure. In contrast, most of the computer vision models that learn visual representation from a pool of 2D images often fail to generalize over novel camera viewpoints. Recently,…

2022

MS-TCT: Multi-Scale Temporal ConvTransformer for Action Detection

CVPR 2022poster

Action detection is an essential and challenging task, especially for densely labelled datasets of untrimmed videos. The temporal relation is complex in those datasets, including challenges like composite action, and co-occurring action. For detecting actions in those complex videos, efficiently cap…

Cited by 99PDFcodeScholar
2022

Self-Supervised Video Transformer

CVPR 2022oral

In this paper, we propose self-supervised training for video transformers using unlabeled video data. From a given video, we create local and global spatiotemporal views with varying spatial sizes and frame rates. Our self-supervised objective seeks to match the features of these different views rep…

Cited by 131PDFcodeScholar
2022

StARformer: Transformer with State-Action-Reward Representations for Visual Reinforcement Learning

ECCV 2022poster

"Reinforcement Learning (RL) can be considered as a sequence modeling task: given a sequence of past state-action-reward experiences, an agent predicts a sequence of next actions. In this work, we propose State-Action-Reward Transformer (StARformer) for visual RL, which explicitly models short-term…

2022

Video Question Answering with Iterative Video-Text Co-Tokenization

ECCV 2022poster

"Video question answering is a challenging task that requires understanding jointly the language input, the visual information in individual video frames, as well as the temporal information about the events occurring in the video. In this paper, we propose a novel multi-stream video encoder for vid…

Cited by 24SourcePDFScholar
2021

Self-Supervised Disentangled Representation Learning for Third-Person Imitation Learning

IROS 2021poster

Humans learn to imitate by observing others. However, robot imitation learning generally requires expert demonstrations in the first-person view (FPV). Collecting such FPV videos for every robot could be very expensive.Third-person imitation learning (TPIL) is the concept of learning action policies…

Cited by 25SourceScholar
2021

TokenLearner: Adaptive Space-Time Tokenization for Videos

NeurIPS 2021poster

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting strategies to obtain visual tokens and processing a large numb…

Cited by 179SourcePDFScholar
2021

Visionary: Vision architecture discovery for robot learning

ICRA 2021poster

We propose a vision-based architecture search algorithm for robot manipulation learning, which discovers interactions between low dimension action inputs and high dimensional visual inputs. Our approach automatically designs architectures while training on the task – discovering novel ways of combin…

Cited by 12SourceScholar
2020

Adversarial Generative Grammars for Human Activity Prediction

ECCV 2020poster

In this paper we propose an adversarial generative grammar model for future prediction. The objective is to learn a model that explicitly captures temporal dependencies, providing a capability to forecast multiple, distinct future activities. Our adversarial grammar is designed so that it can learn…

Cited by 36SourcePDFScholar
2020

AssembleNet++: Assembling Modality Representations via Attention Connections - Supplementary Material -

ECCV 2020poster

We create a family of powerful video models which are able to: (i) learn interactions between semantic object information and raw appearance and motion features, and (ii) deploy attention in order to better learn the importance of features at each convolutional block of the network. A new network co…

Cited by 1SourcePDFScholar
2020

AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures

ICLR 2020poster

Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to include the time dimension, using modules such as 3D convolutions, or by using…

Cited by 121SourcecodeScholar
2020

AttentionNAS: Spatiotemporal Attention Cell Search for Video Classification

ECCV 2020poster

Convolutional operations have two limitations: (1) do not explicitly model where to focus as the same filter is applied to all the positions, and (2) are unsuitable for modeling long-range dependencies as they only operate on a small neighborhood. While both limitations can be alleviated by attentio…

Cited by 56SourcePDFScholar
2020

Password-conditioned Anonymization and Deanonymization with Face Identity Transformers

ECCV 2020poster

Cameras are prevalent in our daily lives, and enable many useful systems built upon computer vision technologies such as smart cameras and home robots for service applications. However, there is also an increasing societal concern as the captured images/videos may contain privacy-sensitive informati…

Cited by 66SourcePDFScholar
2019

Privacy-Preserving Robot Vision with Anonymized Faces by Extreme Low Resolution

IROS 2019poster

As smart cameras are becoming ubiquitous in mobile robot systems, there is an increasing concern in camera devices invading people's privacy by recording unwanted images. We want to fundamentally protect privacy by blurring unwanted blocks in images, such as faces, yet ensure that the robots can und…

Cited by 14SourceScholar
2018

Joint Person Segmentation and Identification in Synchronized First- and Third-person Videos

ECCV 2018poster

In a world of pervasive cameras, public spaces are often captured from multiple perspectives by cameras of different types, both fixed and mobile. An important problem is to organize these heterogeneous collections of videos by finding connections between them, such as identifying correspondences be…

Cited by 46SourcePDFScholar
2018

Learning to Anonymize Faces for Privacy Preserving Action Detection

ECCV 2018poster

There is an increasing concern in computer vision devices invading the privacy of their users. We want the camera systems/robots to recognize important events and assist human daily life by understanding its videos, but we also want to ensure that they do not intrude people's privacy. In this paper,…

Cited by 271SourcePDFScholar
2017

Identifying First-Person Camera Wearers in Third-Person Videos

CVPR 2017poster

We consider scenarios in which we wish to perform joint scene understanding, object tracking, activity recognition, and other tasks in scenarios in which multiple people are wearing body-worn cameras while a third-person static camera also captures the scene. To do this, we need to establ…

Cited by 77PDFScholar
2017

Learning robot activities from first-person human videos using convolutional future regression

IROS 2017poster

We design a new approach that allows robot learning of new activities from unlabeled human example videos. Given videos of humans executing the same activity from a human's viewpoint (i.e., first-person videos), our objective is to make the robot learn the temporal structure of the activity as its f…

Cited by 70SourceScholar
2017

Learning social affordance grammar from videos: Transferring human interactions to human-robot interactions

ICRA 2017poster

In this paper, we present a general framework for learning social affordance grammar as a spatiotemporal AND-OR graph (ST-AOG) from RGB-D videos of human interactions, and transfer the grammar to humanoids to enable a real-time motion inference for human-robot interaction (HRI). Based on Gibbs sampl…

Cited by 54SourceScholar