← Search

Antonio Torralba

165 accepted papers

2026

Ambient Dataloops: Generative Models for Dataset Refinement

ICML 2026poster

We propose Ambient Dataloops, an iterative framework for refining datasets that makes it easier for diffusion models to learn the underlying data distribution. Modern datasets contain samples of highly varying quality, and training directly on such heterogeneous data often yields suboptimal models. …

Cited by 0SourceScholar
2026

MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

ICLR 2026poster

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are limited in size, language coverage, and task diversity. We introduce *MathNet*, a large-scale, high-quality, multilingual, and multimodal dataset of Olympiad-lev…

Cited by 0SourcecodeScholar
2026

Motion Attribution for Video Generation

ICML 2026oral

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework that scales to modern, large, high-quality video datasets and m…

Cited by 0SourceScholar
2026

PRISM: Controllable Diffusion for Compound Image Restoration with Scientific Fidelity

ICLR 2026poster

Scientific and environmental imagery are often degraded by multiple compounding factors related to sensor noise and environmental effects. Existing restoration methods typically treat these mixed effects by iteratively removing fixed categories, lacking the compositionality needed to handle real-wor…

Cited by 0SourceScholar
2026

Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis

ICML 2026poster

Strong semantic representations improve the convergence and generation quality of diffusion and flow models. Existing approaches largely rely on external models, which require separate training, operate on misaligned objectives, and exhibit unexpected scaling behavior. We argue that this dependence …

Cited by 0SourceScholar
2026

VirtualEnv: A Platform for Embodied AI Research

AAAI 2026technical

As large language models (LLMs) continue to improve in reasoning and decision-making, there is a growing need for realistic and interactive environments where their abilities can be rigorously evaluated. We present VirtualEnv, a next-generation simulation platform built on Unreal Engine 5 that enabl

Cited by 0SourcePDFScholar
2025

Adaptive Length Image Tokenization via Recurrent Allocation

ICLR 2025poster

Current vision systems typically assign fixed-length representations to images, regardless of the information content. This contrasts with human intelligence —and even large language models—which allocate varying representational capacities based on entropy, context and familiarity. Inspired by this…

2025

Ambient Diffusion Omni: Training Good Models with Bad Data

NeurIPS 2025spotlight

We show how to use low-quality, synthetic, and out-of-distribution images to improve the quality of a diffusion model. Typically, diffusion models are trained on curated datasets that emerge from highly filtered data pools from the Web and other sources. We show that there is immense value in the lo…

Cited by 0SourcecodeScholar
2025

Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent

NeurIPS 2025poster

When a vision model performs image recognition, which visual attributes drive its predictions? Detecting unintended reliance on specific visual features is critical for ensuring model robustness, preventing overfitting, and avoiding spurious correlations. We introduce an automated framework for dete…

Cited by 0SourceScholar
2025

Dataset Distillation for Pre-Trained Self-Supervised Vision Models

NeurIPS 2025poster

The task of dataset distillation aims to find a small set of synthetic images such that training a model on them reproduces the performance of the same model trained on a much larger dataset of real samples. Existing distillation methods focus on synthesizing datasets that enable training randomly i…

Cited by 0SourcecodeScholar
2025

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

CVPR 2025poster

Despite the unprecedented progress in the field of 3D generation, current systems still often fail to produce high-quality 3D assets that are visually appealing and geometrically and semantically consistent across multiple viewpoints. To effectively assess the quality of the generated 3D data, there…

Cited by 1SourcePDFScholar
2025

Feed-Forward Bullet-Time Reconstruction of Dynamic Scenes from Monocular Videos

NeurIPS 2025poster

Recent advancements in static feed-forward scene reconstruction have demonstrated significant progress in high-quality novel view synthesis. However, these models often struggle with generalizability across diverse environments and fail to effectively handle dynamic content. We present BTimer (short…

Cited by 0SourceScholar
2025

LoRA vs Full Fine-tuning: An Illusion of Equivalence

NeurIPS 2025poster

Fine-tuning is a crucial paradigm for adapting pre-trained large language models to downstream tasks. Recently, methods like Low-Rank Adaptation (LoRA) have been shown to effectively fine-tune LLMs with an extreme reduction in trainable parameters. But, \emph{are their learned solutions really equiv…

Cited by 0SourceScholar
2025

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

ICLR 2025poster

Large language models (LLMs) have achieved remarkable performance in recent years but are fundamentally limited by the underlying training data. To improve models beyond the training data, recent works have explored how LLMs can be used to generate synthetic data for autonomous self-improvement. How…

Cited by 12SourcePDFScholar
2025

Separating Knowledge and Perception with Procedural Data

ICML 2025poster

We train representation models with procedural data only, and apply them on visual similarity, classification, and semantic segmentation tasks without further training by using visual memory---an explicit database of reference image embeddings. Unlike prior work on visual memory, our approach achiev…

Cited by 0SourcePDFScholar
2025

Single-pass Adaptive Image Tokenization for Minimum Program Search

NeurIPS 2025poster

According to Algorithmic Information Theory (AIT), intelligent representations compress data into the shortest possible program while remaining predictive of its content—exhibiting low Kolmogorov Complexity (KC). In contrast, most visual representation learning systems assign fixed-length representa…

Cited by 0SourceScholar
2025

SketchAgent: Language-Driven Sequential Sketch Generation

CVPR 2025poster

Sketching serves as a versatile tool for externalizing ideas, enabling rapid exploration and visual communication that spans various disciplines. While artificial systems have driven substantial advances in content creation and human-computer interaction, capturing the dynamic and abstract nature of…

Cited by 5SourcePDFScholar
2024

A Multimodal Automated Interpretability Agent

ICML 2024poster

This paper describes MAIA, a Multimodal Automated Interpretability Agent. MAIA is a system that uses neural models to automate neural model understanding tasks like feature interpretation and failure mode discovery. It equips a pre-trained vision-language model with a set of tools that support itera…

Cited by 67SourcePDFScholar
2024

A Vision Check-up for Language Models

CVPR 2024highlight

What does learning to model relationships between strings teach Large Language Models (LLMs) about the visual world? We systematically evaluate LLMs' abilities to generate and recognize an assortment of visual concepts of increasing complexity and then demonstrate how a preliminary visual representa…

Cited by 29SourcePDFScholar
2024

Align Your Gaussians: Text-to-4D with Dynamic 3D Gaussians and Composed Diffusion Models

CVPR 2024highlight

Text-guided diffusion models have revolutionized image and video generation and have also been successfully used for optimization-based 3D object synthesis. Here we instead focus on the underexplored text-to-4D setting and synthesize dynamic animated 3D objects using score distillation methods with…

Cited by 110SourcePDFScholar
2024

Characterizing Model Robustness via Natural Input Gradients

ECCV 2024poster

"Adversarially robust models are locally smooth around each data sample so that small perturbations cannot drastically change model outputs. In modern systems, such smoothness is usually obtained via Adversarial Training, which explicitly enforces models to perform well on perturbed examples. In thi…

2024

Concept Sliders: LoRA Adaptors for Precise Control in Diffusion Models

ECCV 2024poster

"We present a method to create interpretable concept sliders that enable precise control over attributes in image generations from diffusion models. Our approach identifies a low-rank parameter direction corresponding to one concept while minimizing interference with other attributes. A slider is cr…

2024

ConceptGraphs: Open-Vocabulary 3D Scene Graphs for Perception and Planning

ICRA 2024poster

For robots to perform a wide variety of tasks, they require a 3D representation of the world that is semantically rich, yet compact and efficient for task-driven perception and planning. Recent approaches have attempted to leverage features from large vision-language models to encode semantics in 3D…

Cited by 202SourceScholar
2024

Efficient 3D Instance Mapping and Localization with Neural Fields

ICRA 2024poster

We tackle the problem of learning an implicit scene representation for 3D instance segmentation from a sequence of posed RGB images. Towards this, we introduce 3DIML, a novel framework that efficiently learns a label field that may be rendered from novel viewpoints to produce view-consistent instanc…

Cited by 6SourceScholar
2024

Follow Anything: Open-Set Detection, Tracking, and Following in Real-Time

RA-L 2024

Tracking and following objects of interest is critical to several robotics use cases, ranging from industrial automation to logistics and warehousing, to healthcare and security. In this paper, we present a robotic system to detect, track, and follow any object in real-time. Our approach, dubbed <it

Cited by 41SourcecodeScholar
2024

GOMA: Proactive Embodied Cooperative Communication via Goal-Oriented Mental Alignment

IROS 2024poster

Verbal communication plays a crucial role in human cooperation, particularly when the partners only have incomplete information about the task, environment, and each other’s mental state. In this paper, we propose a novel cooperative communication framework, Goal-Oriented Mental Alignment (GOMA). GO…

Cited by 10SourceScholar
2024

Improving Factuality and Reasoning in Language Models through Multiagent Debate

ICML 2024poster

Large language models (LLMs) have demonstrated remarkable capabilities in language generation, understanding, and few-shot learning in recent years. An extensive body of work has explored how their performance may be further improved through the tools of prompting, ranging from verification, self-co…

2024

L4GM: Large 4D Gaussian Reconstruction Model

NeurIPS 2024poster

We present L4GM, the first 4D Large Reconstruction Model that produces animated objects from a single-view video input -- in a single feed-forward pass that takes only a second. Key to our success is a novel dataset of multiview videos containing curated, rendered animated objects from Objaverse. Th…

Cited by 38SourcePDFScholar
2024

LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis

ECCV 2024poster

"Recent text-to-3D generation approaches produce impressive 3D results but require time-consuming optimization that can take up to an hour per prompt. Amortized methods like ATT3D optimize multiple prompts simultaneously to improve efficiency, enabling fast text-to-3D synthesis. However, they cannot…

2024

Learning to Jointly Understand Visual and Tactile Signals

ICLR 2024poster

Modeling and analyzing object and shape has been well studied in the past. However, manipulation of these complex tools and articulated objects remains difficult for autonomous agents. Our human hands, however, are dexterous and adaptive. We can easily adapt a manipulation skill on one object to all…

Cited by 6SourcePDFScholar
2024

MMToM-QA: Multimodal Theory of Mind Question Answering

ACL 2024long

Theory of Mind (ToM), the ability to understand people’s mental states, is an essential ingredient for developing machines with human-level social intelligence. Recent machine learning models, particularly large language models, seem to show some aspects of ToM understanding. However, existing ToM b…

2023

3D-IntPhys: Towards More Generalized 3D-grounded Visual Intuitive Physics under Challenging Scenes

NeurIPS 2023poster

Given a visual scene, humans have strong intuitions about how a scene can evolve over time under given actions. The intuition, often termed visual intuitive physics, is a critical ability that allows us to make effective plans to manipulate the scene to achieve desired outcomes without relying on ex…

Cited by 8SourcePDFScholar
2023

BT^2: Backward-compatible Training with Basis Transformation

ICCV 2023poster

Modern retrieval system often requires recomputing the representation of every piece of data in the gallery when updating to a better representation model. This process is known as backfilling and can be especially costly in the real world where the gallery often contains billions of samples. Recent…

Cited by 6PDFcodeScholar
2023

Composing Ensembles of Pre-trained Models via Iterative Consensus

ICLR 2023poster

Large pre-trained models exhibit distinct and complementary capabilities dependent on the data they are trained on. Language models such as GPT-3 are capable of textual reasoning but cannot understand visual information, while vision models such as DALL-E can generate photorealistic photos but fail…

Cited by 33SourcePDFScholar
2023

ConceptFusion: Open-set multimodal 3D mapping

RSS 2023poster

Building 3D maps of the environment is central to robot navigation, planning, and interaction with objects in a scene. Most existing approaches that integrate semantic concepts with 3D maps largely remain confined to the closed-set setting: they can only reason about a finite set of concepts, pre-de…

2023

Detecting Everything in the Open World: Towards Universal Object Detection

CVPR 2023poster

In this paper, we formally address universal object detection, which aims to detect every scene and predict every category. The dependence on human annotations, the limited visual information, and the novel categories in the open world severely restrict the universality of traditional detectors. We…

2023

DreamTeacher: Pretraining Image Backbones with Deep Generative Models

ICCV 2023poster

In this work, we introduce a self-supervised feature representation learning framework DreamTeacher that utilizes generative networks for pre-training downstream image backbones. We propose to distill knowledge from a trained generative model into standard image backbones that have been well enginee…

Cited by 23PDFScholar
2023

FIND: A Function Description Benchmark for Evaluating Interpretability Methods

NeurIPS 2023poster

Labeling neural network submodules with human-legible descriptions is useful for many downstream tasks: such descriptions can surface failures, guide interventions, and perhaps even explain important model behaviors. To date, most mechanistic descriptions of trained networks have involved small mode…

2023

FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

ICLR 2023top-25%

Humans manipulate various kinds of fluids in their everyday life: creating latte art, scooping floating objects from water, rolling an ice cream cone, etc. Using robots to augment or replace human labors in these daily settings remain as a challenging task due to the multifaceted complexities of flu…

2023

Generalizing Dataset Distillation via Deep Generative Prior

CVPR 2023poster

Dataset Distillation aims to distill an entire dataset's knowledge into a few synthetic images. The idea is to synthesize a small number of synthetic data points that, when given to a learning algorithm as training data, result in a model approximating one trained on the original data. Despite a rec…

2023

NOPA: Neurally-guided Online Probabilistic Assistance for Building Socially Intelligent Home Assistants

ICRA 2023poster

In this work, we study how to build socially intelligent robots to assist people in their homes. In particular, we focus on assistance with online goal inference, where robots must simultaneously infer humans' goals and how to help them achieve those goals. Prior assistance methods either lack the a…

Cited by 24SourceScholar
2023

NeuralField-LDM: Scene Generation With Hierarchical Latent Diffusion Models

CVPR 2023poster

Automatically generating high-quality real world 3D scenes is of enormous interest for applications such as virtual reality and robotics simulation. Towards this goal, we introduce NeuralField-LDM, a generative model capable of synthesizing complex 3D environments. We leverage Latent Diffusion Model…

2023

Open-vocabulary Panoptic Segmentation with Embedding Modulation

ICCV 2023poster

Open-vocabulary segmentation is attracting increasing attention due to its critical applications in the real world. Traditional closed-vocabulary segmentation methods are not able to characterize novel objects, whereas several recent open-vocabulary attempts obtain unsatisfactory results, i.e., nota…

Cited by 34PDFScholar
2023

Optimal Goal-Reaching Reinforcement Learning via Quasimetric Learning

ICML 2023poster

In goal-reaching reinforcement learning (RL), the optimal value function has a particular geometry, called quasimetrics structure. This paper introduces Quasimetric Reinforcement Learning (QRL), a new RL method that utilizes quasimetric models to learn optimal value functions. Distinct from prior ap…

2023

Physics-Driven Diffusion Models for Impact Sound Synthesis From Videos

CVPR 2023poster

Modeling sounds emitted from physical object interactions is critical for immersive perceptual experiences in real and virtual worlds. Traditional methods of impact sound synthesis use physics simulation to obtain a set of physics parameters that could represent and synthesize the sound. However, th…

Cited by 30SourcePDFScholar
2023

Structure from Duplicates: Neural Inverse Graphics from a Pile of Objects

NeurIPS 2023poster

Abstract Our world is full of identical objects (\emph{e.g.}, cans of coke, cars of same model). These duplicates, when seen together, provide additional and strong cues for us to effectively reason about 3D. Inspired by this observation, we introduce Structure from Duplicates (SfD), a novel inverse…

2023

Unsupervised Compositional Concepts Discovery with Text-to-Image Generative Models

ICCV 2023poster

Text-to-image generative models have enabled high-resolution image synthesis across different domains, but require users to specify the content they wish to generate. In this paper, we consider the inverse problem - given a collection of different images, can we discover the generative concepts that…

Cited by 14PDFScholar
2022

ActionSense: A Multimodal Dataset and Recording Framework for Human Activities Using Wearable Sensors in a Kitchen Environment

NeurIPS 2022accept

This paper introduces ActionSense, a multimodal dataset and recording framework with an emphasis on wearable sensing in a kitchen environment. It provides rich, synchronized data streams along with ground truth data to facilitate learning pipelines that could extract insights about how humans inter…

Cited by 59SourcePDFScholar
2022

BigDatasetGAN: Synthesizing ImageNet With Pixel-Wise Annotations

CVPR 2022poster

Annotating images with pixel-wise labels is a time-consuming and costly process. Recently, DatasetGAN showcased a promising alternative - to synthesize a large labeled dataset via a generative adversarial network (GAN) by exploiting a small set of manually labeled, GAN-generated images. Here, we sca…

Cited by 121PDFScholar
2022

ComPhy: Compositional Physical Reasoning of Objects and Events from Videos

ICLR 2022poster

Objects' motions in nature are governed by complex interactions and their properties. While some properties, such as shape and material, can be identified via the object's visual appearances, others like mass and electric charge are not directly visible. The compositionality between the visible and…

Cited by 58SourcePDFScholar
2022

Compositional Visual Generation with Composable Diffusion Models

ECCV 2022poster

"Large text-guided diffusion models, such as DALLE-2, are able to generate stunning photorealistic images given natural language descriptions. While such models are highly flexible, they struggle to understand the composition of certain concepts, such as confusing the attributes of different objects…

2022

Correcting Robot Plans with Natural Language Feedback

RSS 2022poster

When humans design cost or goal specifications for robots, they often produce specifications that are ambiguous, under-specified, or beyond planners’ ability to solve. In these cases, corrections provide a valuable tool for human-in-the-loop robot control. Corrections might take the form of new goal…

Cited by 110SourcePDFScholar
2022

Dataset Distillation by Matching Training Trajectories

CVPR 2022oral

Dataset distillation is the task of synthesizing a small dataset such that a model trained on the synthetic set will match the test accuracy of the model trained on the full dataset. The task is extremely challenging as it often involves backpropagating through the full training process or assuming…

Cited by 451PDFcodeScholar
2022

Denoised MDPs: Learning World Models Better Than the World Itself

ICML 2022spotlight

The ability to separate signal from noise, and reason with clean abstractions, is critical to intelligence. With this ability, humans can efficiently perform real world tasks without considering all possible nuisance factors. How can artificial agents do the same? What kind of information can agents…

2022

Ego4D: Around the World in 3,000 Hours of Egocentric Video

CVPR 2022oral

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countri…

Cited by 1162PDFcodeScholar
2022

Finding Fallen Objects via Asynchronous Audio-Visual Integration

CVPR 2022poster

The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in…

Cited by 20PDFScholar
2022

Fixing Malfunctional Objects With Learned Physical Simulation and Functional Prediction

CVPR 2022poster

This paper studies the problem of fixing malfunctional 3D objects. While previous works focus on building passive perception models to learn the functionality from static 3D objects, we argue that functionality is reckoned with respect to the physical interactions between the object and the user. Gi…

Cited by 6PDFScholar
2022

GAN-Supervised Dense Visual Alignment

CVPR 2022oral

We propose GAN-Supervised Learning, a framework for learning discriminative models and their GAN-generated training data jointly end-to-end. We apply our framework to the dense visual alignment problem. Inspired by the classic Congealing method, our GANgealing algorithm trains a Spatial Transformer…

Cited by 78PDFcodeScholar
2022

Learning Neural Acoustic Fields

NeurIPS 2022accept

Our environment is filled with rich and dynamic acoustic information. When we walk into a cathedral, the reverberations as much as appearance inform us of the sanctuary's wide open space. Similarly, as an object moves around us, we expect the sound emitted to also exhibit this movement. While recent…

Cited by 79SourcePDFScholar
2022

Learning Program Representations for Food Images and Cooking Recipes

CVPR 2022oral

In this paper, we are interested in modeling a how-to instructional procedure, such as a cooking recipe, with a meaningful and rich high-level representation. Specifically, we propose to represent cooking recipes and food images as cooking programs. Programs provide a structured representation of th…

Cited by 41PDFScholar
2022

MTFormer: Multi-task Learning via Transformer and Cross-Task Reasoning

ECCV 2022poster

"In this paper, we explore the advantages of utilizing transformer structures for addressing multi-task learning (MTL). Specifically, we demonstrate that models with transformer structures are more appropriate for MTL than convolutional neural networks (CNNs), and we propose a novel transformer-base…

Cited by 67SourcePDFScholar
2022

Natural Language Descriptions of Deep Visual Features

ICLR 2022oral

Some neurons in deep networks specialize in recognizing highly specific perceptual, structural, or semantic features of inputs. In computer vision, techniques exist for identifying neurons that respond to individual concept categories like colors, textures, and object classes. But these techniques a…

Cited by 134SourcePDFScholar
2022

Noisy Agents: Self-supervised Exploration by Predicting Auditory Events

IROS 2022poster

Humans integrate multiple sensory modalities (e.g., visual and audio) to build a causal understanding of the physical world. In this work, we propose a novel type of intrinsic motivation for Reinforcement Learning (RL) that encourages the agent to understand the causal effect of its actions through…

Cited by 8SourceScholar
2022

Polymorphic-GAN: Generating Aligned Samples Across Multiple Domains With Learned Morph Maps

CVPR 2022oral

Modern image generative models show remarkable sample quality when trained on a single domain or class of objects. In this work, we introduce a generative adversarial network that can simultaneously generate aligned image samples from multiple related domains. We leverage the fact that a variety of…

Cited by 9PDFScholar
2022

Pre-Trained Language Models for Interactive Decision-Making

NeurIPS 2022accept

Language model (LM) pre-training is useful in many language processing tasks. But can pre-trained LMs be further leveraged for more general machine learning problems? We propose an approach for using LMs to scaffold learning and generalization in general sequential decision-making problems. In this…

Cited by 229SourcePDFScholar
2022

Procedural Image Programs for Representation Learning

NeurIPS 2022accept

Learning image representations using synthetic data allows training neural networks without some of the concerns associated with real images, such as privacy and bias. Existing work focuses on a handful of curated generative processes which require expert knowledge to design, making it hard to scale…

2022

Robust Contrastive Learning Against Noisy Views

CVPR 2022poster

Contrastive learning relies on an assumption that positive pairs contain related views that share certain underlying information about an instance, e.g., patches of an image or co-occurring multimodal signals of a video. What if this assumption is violated? The literature suggests that contrastive l…

Cited by 100PDFcodeScholar
2022

The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark Towards Physically Realistic Embodied AI

ICRA 2022poster

We introduce a visually-guided task-and-motion planning benchmark, which we call the ThreeDWorld Trans-port Challenge. In this challenge, an embodied agent is spawned randomly in a simulated physical home environment and required to transport a small set of objects scattered around the house with co…

Cited by 46SourceScholar
2022

Totems: Physical Objects for Verifying Visual Integrity

ECCV 2022poster

"We introduce a new approach to image forensics: placing physical refractive objects, which we call totems, into a scene so as to protect any photograph taken of that scene. Totems bend and redirect light rays, thus providing multiple, albeit distorted, views of the scene within a single image. A de…

Cited by 3SourcePDFScholar
2022

Virtual Correspondence: Humans as a Cue for Extreme-View Geometry

CVPR 2022poster

Recovering the spatial layout of the cameras and the geometry of the scene from extreme-view images is a longstanding challenge in computer vision. Prevailing 3D reconstruction algorithms often adopt the image matching paradigm and presume that a portion of the scene is co-visible across images, yie…

Cited by 27PDFScholar
2021

3D Neural Scene Representations for Visuomotor Control

CoRL 2021oral

Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us to perform a wide range of manipulation tasks that are far beyond the reach of current robots. In this work, we desire t…

Cited by 155SourceScholar
2021

DatasetGAN: Efficient Labeled Data Factory With Minimal Human Effort

CVPR 2021poster

We introduce DatasetGAN: an automatic procedure to generate massive datasets of high-quality semantically segmented images requiring minimal human effort. Current deep networks are extremely data-hungry, benefiting from training on large-scale datasets, which are time-consuming to annotate. Our meth…

Cited by 393PDFcodeScholar
2021

DriveGAN: Towards a Controllable High-Quality Neural Simulation

CVPR 2021poster

Realistic simulators are critical for training and verifying robotics systems. While most of the contemporary simulators are hand-crafted, a scaleable way to build simulators is to use machine learning to learn how the environment behaves in response to an action, directly from data. In this work, w…

Cited by 119PDFScholar
2021

Dynamic Modeling of Hand-Object Interactions via Tactile Sensing

IROS 2021poster

Tactile sensing is critical for humans to perform everyday tasks. While significant progress has been made in analyzing object grasping from vision, it remains unclear how we can utilize tactile sensing to reason about and model the dynamics of hand-object interactions. In this work, we employ a hig…

Cited by 19SourceScholar
2021

EditGAN: High-Precision Semantic Image Editing

NeurIPS 2021poster

Generative adversarial networks (GANs) have recently found applications in image editing. However, most GAN-based image editing methods often require large-scale datasets with semantic segmentation annotations for training, only provide high-level control, or merely interpolate between different ima…

Cited by 280SourcePDFScholar
2021

Editing a classifier by rewriting its prediction rules

NeurIPS 2021poster

We propose a methodology for modifying the behavior of a classifier by directly rewriting its prediction rules. Our method requires virtually no additional data collection and can be applied to a variety of settings, including adapting a model to new environments, and modifying it to ignore spurious…

2021

Image GANs meet Differentiable Rendering for Inverse Graphics and Interpretable 3D Neural Rendering

ICLR 2021oral

Differentiable rendering has paved the way to training neural networks to perform “inverse graphics” tasks such as predicting 3D geometry from monocular photographs. To train high performing models, most of the current approaches rely on multi-view imagery which are not readily available in practice…

Cited by 145SourcePDFScholar
2021

Intelligent Carpet: Inferring 3D Human Pose From Tactile Signals

CVPR 2021poster

Daily human activities, e.g., locomotion, exercises, and resting, are heavily guided by the tactile interactions between the human and the ground. In this work, leveraging such tactile interactions, we propose a 3D human pose estimation approach using the pressure maps recorded by a tactile carpet a…

Cited by 67PDFScholar
2021

Learning to See by Looking at Noise

NeurIPS 2021spotlight

Current vision systems are trained on huge datasets, and these datasets come with costs: curation is expensive, they inherit human biases, and there are concerns over privacy and usage rights. To counter these costs, interest has surged in learning from cheaper data sources, such as unlabeled images…

2021

Measuring Generalization with Optimal Transport

NeurIPS 2021spotlight

Understanding the generalization of deep neural networks is one of the most important tasks in deep learning. Although much progress has been made, theoretical error bounds still often behave disparately from empirical observations. In this work, we develop margin-based generalization bounds, where…

2021

OPEn: An Open-ended Physics Environment for Learning Without a Task

IROS 2021poster

Humans have mental models that allow them to plan, experiment, and reason in the physical world. How should an intelligent agent go about learning such models? In this paper, we will study if models of the world learned in an open-ended physics environment, without any specific tasks, can be reused…

Cited by 3SourceScholar
2021

PTR: A Benchmark for Part-based Conceptual, Relational, and Physical Reasoning

NeurIPS 2021poster

A critical aspect of human visual perception is the ability to parse visual scenes into individual objects and further into object parts, forming part-whole hierarchies. Such composite structures could induce a rich set of semantic concepts and relations, thus playing an important role in the interp…

Cited by 49SourcePDFScholar
2021

Semantic Segmentation With Generative Models: Semi-Supervised Learning and Strong Out-of-Domain Generalization

CVPR 2021poster

Training deep networks with limited labeled data while achieving a strong generalization ability is key in the quest to reduce human annotation efforts. This is the goal of semi-supervised learning, which exploits more widely available unlabeled data to complement small labeled data sets. In this pa…

Cited by 241PDFcodeScholar
2021

ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation

NeurIPS 2021poster

We introduce ThreeDWorld (TDW), a platform for interactive multi-modal physical simulation. TDW enables the simulation of high-fidelity sensory data and physical interactions between mobile agents and objects in rich 3D environments. Unique properties include real-time near-photo-realistic image ren…

Cited by 342SourcecodeScholar
2021

Toward a Visual Concept Vocabulary for GAN Latent Space

ICCV 2021poster

A large body of recent work has identified transformations in the latent spaces of generative adversarial networks (GANs) that consistently and interpretably transform generated images. But existing techniques for identifying these transformations rely on either a fixed vocabulary of pre-specified v…

Cited by 19PDFcodeScholar
2021

Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration

ICLR 2021spotlight

In this paper, we introduce Watch-And-Help (WAH), a challenge for testing social intelligence in agents. In WAH, an AI agent needs to help a human-like agent perform a complex household task efficiently. To succeed, the AI agent needs to i) understand the underlying goal of the task by watching a si…

2021

Weakly Supervised Human-Object Interaction Detection in Video via Contrastive Spatiotemporal Regions

ICCV 2021poster

We introduce the task of weakly supervised learning for detecting human and object interactions in videos. Our task poses unique challenges as a system does not know what types of human-object interactions are present in a video or the actual spatiotemporal location of the human and object. To addre…

Cited by 13PDFcodeScholar
2021

What You Can Learn by Staring at a Blank Wall

ICCV 2021poster

We present a passive non-line-of-sight method that infers the number of people or activity of a person from the observation of a blank wall in an unknown room. Our technique analyzes complex imperceptible changes in indirect illumination in a video of the wall to reveal a signal that is correlated w…

Cited by 19PDFScholar
2020

CLEVRER: Collision Events for Video Representation and Reasoning

ICLR 2020spotlight

The ability to reason about temporal and causal events from videos lies at the core of human intelligence. Most video reasoning benchmarks, however, focus on pattern recognition from complex visual and language input, instead of on causal structure. We study the complementary problem, exploring the…

Cited by 559SourceScholar
2020

Causal Discovery in Physical Systems from Videos

NeurIPS 2020poster

Causal discovery is at the core of human cognition. It enables us to reason about the environment and make counterfactual predictions about unseen scenarios that can vastly differ from our previous experiences. We consider the task of causal discovery from videos in an end-to-end fashion without sup…

Cited by 126SourcePDFScholar
2020

Debiased Contrastive Learning

NeurIPS 2020spotlight

A prominent technique for self-supervised representation learning has been to contrast semantically similar and dissimilar pairs of samples. Without access to labels, dissimilar (negative) points are typically taken to be randomly sampled datapoints, implicitly accepting that these points may, in re…

2020

Deep Audio Priors Emerge From Harmonic Convolutional Networks

ICLR 2020poster

Convolutional neural networks (CNNs) excel in image recognition and generation. Among many efforts to explain their effectiveness, experiments show that CNNs carry strong inductive biases that capture natural image priors. Do deep networks also have inductive biases for audio signals? In this paper,…

Cited by 40SourceScholar
2020

Deep Feedback Inverse Problem Solver

ECCV 2020poster

We present an efficient, effective, and generic approach towards solving inverse problems. The key idea is to leverage the feedback signal provided by the forward process and learn an iterative update model. Specifically, in each iteration, the neural network takes the feedback as input and outputs…

2020

Detecting Natural Disasters, Damage, and Incidents in the Wild

ECCV 2020poster

damage, and incidents in the wild","Responding to natural disasters, such as earthquakes, floods, and wildfires, is a laborious task performed by on-the-ground emergency responders and analysts. Social media has emerged as a low-latency data source to quickly understand disaster situations. While mo…

Cited by 77SourcePDFScholar
2020

Diverse Image Generation via Self-Conditioned GANs

CVPR 2020poster

We introduce a simple but effective unsupervised method for generating diverse images. We train a class-conditional GAN model without using manually annotated class labels. Instead, our model is conditional on labels automatically derived from clustering in the discriminator's feature space. Our clu…

Cited by 132PDFcodeScholar
2020

Estimating Generalization under Distribution Shifts via Domain-Invariant Representations

ICML 2020poster

When machine learning models are deployed on a test distribution different from the training distribution, they can perform poorly, but overestimate their performance. In this work, we aim to better estimate a model’s performance under distribution shift, without supervision. To do so, we use a set…

2020

Foley Music: Learning to Generate Music from Videos

ECCV 2020poster

In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key intermediate representations for a successful video to music generator: body keypoints from videos and MIDI events from au…

Cited by 168SourcePDFScholar
2020

Learning Compositional Koopman Operators for Model-Based Control

ICLR 2020spotlight

Finding an embedding space for a linear approximation of a nonlinear dynamical system enables efficient system identification and control synthesis. The Koopman operator theory lays the foundation for identifying the nonlinear-to-linear coordinate transformations with data-driven methods. Recently,…

Cited by 152SourceScholar
2020

Learning to Simulate Dynamic Environments With GameGAN

CVPR 2020poster

Simulation is a crucial component of any robotic system. In order to simulate correctly, we need to write complex rules of the environment: how dynamic agents behave, and how the actions of each of the agents affect the behavior of others. In this paper, we aim to learn a simulator by simply watchin…

Cited by 138PDFScholar
2020

The Hessian Penalty: A Weak Prior for Unsupervised Disentanglement

ECCV 2020poster

Existing popular methods for disentanglement rely on hand-picked priors and complex encoder-based architectures. In this paper, we propose the Hessian Penalty, a simple regularization function that encourages the input Hessian of a function to be diagonal. Our method is completely model-agnostic and…

2020

Visual Grounding of Learned Physical Models

ICML 2020poster

Humans intuitively recognize objects’ physical properties and predict their motion, even when the objects are engaged in complicated interactions. The abilities to perform physical reasoning and to adapt to new environments, while intrinsic to humans, remain challenging to state-of-the-art computati…

Cited by 82SourcePDFScholar
2019

GAN Dissection: Visualizing and Understanding Generative Adversarial Networks

ICLR 2019poster

Generative Adversarial Networks (GANs) have recently achieved impressive results for many real-world applications, and many GAN variants have emerged with improvements in sample quality and training stability. However, visualization and understanding of GANs is largely missing. How does a GAN repres…

2019

Gaze360: Physically Unconstrained Gaze Estimation in the Wild

ICCV 2019poster

Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale remote gaze-tracking dataset and method for robust 3D gaze estimation in unconstrained images. Our dataset consists of 238 subjects in indoor and outdoor environments with labelled 3D…

Cited by 475PDFScholar
2019

HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization

ICCV 2019poster

This paper presents a new large-scale dataset for recognition and temporal localization of human actions collected from Web videos. We refer to it as HACS (Human Action Clips and Segments). We leverage consensus and disagreement among visual classifiers to automatically mine candidate short clips fr…

Cited by 347PDFScholar
2019

How to Make a Pizza: Learning a Compositional Layer-Based GAN Model

CVPR 2019poster

A food recipe is an ordered set of instructions for preparing a particular dish. From a visual perspective, every instruction step can be seen as a way to change the visual appearance of the dish by adding extra objects (e.g., adding an ingredient) or changing the appearance of the existing ones (e.…

Cited by 48PDFScholar
2019

Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids

ICLR 2019poster

Real-life control tasks involve matters of various substances---rigid or soft bodies, liquid, gas---each with distinct physical behaviors. This poses challenges to traditional rigid-body physics engines. Particle-based simulators have been developed to model the dynamics of these complex scenes; how…

Cited by 435SourcePDFScholar
2019

Meta-Sim: Learning to Generate Synthetic Datasets

ICCV 2019oral

Training models to high-end performance requires availability of large labeled datasets, which are expensive to get. The goal of our work is to automatically synthesize labeled datasets that are relevant for a downstream task. We propose Meta-Sim, which learns a generative model of synthetic scenes,…

Cited by 316PDFScholar
2019

Neural Turtle Graphics for Modeling City Road Layouts

ICCV 2019oral

We propose Neural Turtle Graphics (NTG), a novel generative model for spatial graphs, and demonstrate its applications in modeling city road layouts. Specifically, we represent the road layout using a graph where nodes in the graph represent control points and edges in the graph represents road segm…

Cited by 104PDFScholar
2019

Propagation Networks for Model-Based Control Under Partial Observation

ICRA 2019poster

There has been an increasing interest in learning dynamics simulators for model-based control. Compared with off-the-shelf physics engines, a learnable simulator can quickly adapt to unseen objects, scenes, and tasks. However, existing models like interaction networks only work for fully observable…

Cited by 170SourcecodeScholar
2019

Seeing What a GAN Cannot Generate

ICCV 2019oral

Despite the success of Generative Adversarial Networks (GANs), mode collapse remains a serious issue during GAN training. To date, little work has focused on understanding and quantifying which modes have been dropped by a model. In this work, we visualize mode collapse at both the distribution leve…

Cited by 457PDFcodeScholar
2019

Self-Supervised Moving Vehicle Tracking With Stereo Sound

ICCV 2019poster

Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We introduce a system that can leverage unlabeled audiovisual data to learn to localize objects (moving vehicles) in a visual re…

Cited by 174PDFScholar
2019

Self-supervised Audio-visual Co-segmentation

ICASSP 2019accepted

Segmenting objects in images and separating sound sources in audio are challenging tasks, in part because traditional approaches require large amounts of labeled data. In this paper we develop a neural network model for visual object segmentation and sound source separation that learns from natural…

Cited by 0SourceScholar
2019

Synthesizing Environment-Aware Activities via Activity Sketches

CVPR 2019poster

In order to learn to perform activities from demonstrations or descriptions, agents need to distill what the essence of the given activity is, and how it can be adapted to new environments. In this work, we address the problem: environment-aware program generation. Given a visual demonstration or a…

Cited by 45PDFScholar
2019

Through-Wall Human Mesh Recovery Using Radio Signals

ICCV 2019poster

This paper presents RF-Avatar, a neural network model that can estimate 3D meshes of the human body in the presence of occlusions, baggy clothes, and bad lighting conditions. We leverage that radio frequency (RF) signals in the WiFi range traverse clothes and occlusions and bounce off the human body…

Cited by 123PDFScholar
2018

3D-Aware Scene Manipulation via Inverse Graphics

NeurIPS 2018poster

We aim to obtain an interpretable, expressive, and disentangled scene representation that contains comprehensive structural and textural information for each object. Previous scene representations learned by neural networks are often uninterpretable, limited to a single object, or lacking 3D knowled…

2018

Inferring Light Fields From Shadows

CVPR 2018poster

We present a method for inferring a 4D light field of a hidden scene from 2D shadows cast by a known occluder on a diffuse wall. We do this by determining how light naturally reflected off surfaces in the hidden scene interacts with the occluder. By modeling the light transport as a linear system, a…

2018

Interpretable Basis Decomposition for Visual Explanation

ECCV 2018poster

Explanations of the decisions made by a deep neural network are important for human end-users to be able to understand and diagnose the trustworthiness of the system. Current neural networks used for visual recognition are generally used as black boxes that do not provide any human interpretable jus…

2018

Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input

ECCV 2018poster

In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer to. We demonstrate that these audio-visual associative localizations emerge from network-internal representations learne…

Cited by 254SourcePDFScholar
2018

Learning to Act Properly: Predicting and Explaining Affordances From Images

CVPR 2018poster

We address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent’s actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the scene, and aim to reason about action-object affordances that…

Cited by 125SourcePDFScholar
2018

Learning to Zoom: a Saliency-Based Sampling Layer for Neural Networks

ECCV 2018poster

We introduce a saliency-based distortion layer for convolutional neural networks that helps to improve the spatial sampling of input data for a given task. Our differentiable layer can be added as a preprocessing block to existing task networks and trained altogether in an end-to-end fashion. The ef…

2018

Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding

NeurIPS 2018spotlight

We marry two powerful ideas: deep representation learning for visual recognition and language understanding, and symbolic program execution for reasoning. Our neural-symbolic visual question answering (NS-VQA) system first recovers a structural scene representation from the image and a program trace…

2018

Real-Time Object Pose Estimation with Pose Interpreter Networks

IROS 2018poster

In this work, we introduce pose interpreter networks for 6-DoF object pose estimation. In contrast to other CNN-based approaches to pose estimation that require expensively annotated object pose data, our pose interpreter network is trained entirely on synthetic pose data. We use object masks as an…

Cited by 59SourcecodeScholar
2018

Single Image Intrinsic Decomposition without a Single Intrinsic Image

ECCV 2018poster

Intrinsic image decomposition---decomposing a natural image into a set of images corresponding to different physical causes---is one of the key and fundamental problems of computer vision. Previous intrinsic decomposition approaches either address the problem in a fully supervised manner, or require…

Cited by 81SourcePDFScholar
2018

The Sound of Pixels

ECCV 2018poster

We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a set of components that represents the sound from each pixel. Our approach capitalizes on the natural synchronization of t…

Cited by 638SourcePDFScholar
2018

Through-Wall Human Pose Estimation Using Radio Signals

CVPR 2018poster

This paper demonstrates accurate human pose estimation through walls and occlusions. We leverage the fact that wireless signals in the WiFi frequencies traverse walls and reflect off the human body. We introduce a deep neural network approach that parses such radio signals to estimate 2D poses. Sinc…

Cited by 731SourcePDFScholar
2018

VirtualHome: Simulating Household Activities via Programs

CVPR 2018poster

In this paper, we are interested in modeling complex activities that occur in a typical household. We propose to use programs, i.e., sequences of atomic actions and interactions, as a high level representation of complex tasks. Programs are interesting because they provide a non-ambiguous representa…

Cited by 620SourcePDFScholar
2018

Visual Object Networks: Image Generation with Disentangled 3D Representations

NeurIPS 2018poster

Recent progress in deep generative models has led to tremendous breakthroughs in image generation. While being able to synthesize photorealistic images, existing models lack an understanding of our underlying 3D world. Different from previous works built on 2D datasets and models, we present a new g…

2017

A Compositional Object-Based Approach to Learning Physical Dynamics

ICLR 2017poster

We present the Neural Physics Engine (NPE), a framework for learning simulators of intuitive physics that naturally generalize across variable object count and different scene configurations. We propose a factorization of a physical scene into composable object-based representations and a neural net…

Cited by 530SourceScholar
2017

Learning Cross-Modal Embeddings for Cooking Recipes and Food Images

CVPR 2017poster

In this paper, we introduce Recipe1M, a new large-scale, structured corpus of over 1m cooking recipes and 800k food images. As the largest publicly available collection of recipe data, Recipe1M affords the ability to train high-capacity models on aligned, multi-modal data. Accordingly, we train a ne…

Cited by 758PDFScholar
2017

Network Dissection: Quantifying Interpretability of Deep Visual Representations

CVPR 2017oral

We propose a general framework called Network Dissection for quantifying the interpretability of latent representations of CNNs by evaluating the alignment between individual hidden units and a set of semantic concepts. Given any CNN model, the proposed method draws on a data set of concepts to scor…

Cited by 1943PDFcodeScholar
2017

SegICP: Integrated deep semantic segmentation and pose estimation

IROS 2017poster

Recent robotic manipulation competitions have highlighted that sophisticated robots still struggle to achieve fast and reliable perception of task-relevant objects in complex, realistic scenarios. To improve these systems' perceptive speed and robustness, we present SegICP, a novel integrated soluti…

Cited by 189SourceScholar
2017

Turning Corners Into Cameras: Principles and Methods

ICCV 2017spotlight

We show that walls and other obstructions with edges can be exploited as naturally-occurring "cameras" that reveal the hidden scenes beyond them. In particular, we demonstrate methods for using the subtle spatio-temporal radiance variations that arise on the ground at the base of edges to construct…

Cited by 156PDFScholar
2016

Eye Tracking for Everyone

CVPR 2016poster

From scientific research to commercial applications, eye tracking is an important tool across many domains. Despite its range of applications, eye tracking has yet to become a pervasive technology. We believe that we can put the power of eye tracking in everyone's palm by building eye tracking softw…

Cited by 1274PDFcodeScholar
2016

Learning Aligned Cross-Modal Representations From Weakly Aligned Data

CVPR 2016poster

People can recognize scenes across many different modalities beyond natural images. In this paper, we investigate how to learn cross-modal scene representations that transfer across modalities. To study this problem, we introduce a new cross-modal scene dataset. While convolutional neural networks c…

Cited by 204PDFScholar
2016

Learning Deep Features for Discriminative Localization

CVPR 2016poster

In this work, we revisit the global average pooling layer proposed in [13], and shed light on how it explicitly enables the convolutional neural network (CNN) to have remarkable localization ability despite being trained on image-level labels. While this technique was previously proposed as a means…

Cited by 13283PDFcodeScholar
2016

MovieQA: Understanding Stories in Movies Through Question-Answering

CVPR 2016spotlight

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who" did "What" to "Whom", to "Why" and "How" certain events occur…

Cited by 875PDFScholar
2016

Unsupervised Learning of Spoken Language with Visual Context

NeurIPS 2016poster

Humans learn to speak before they can read or write, so why can’t computers do the same? In this paper, we present a deep neural network model capable of rudimentary spoken language acquisition using untranscribed audio training data, whose only supervision comes in the form of contextually relevant…

2016

Visually Indicated Sounds

CVPR 2016oral

Objects make distinctive sounds when they are hit or scratched. These sounds reveal aspects of an object's material properties, as well as the actions that produced them. In this paper, we propose the task of predicting what sound an object makes when struck as a way of studying physical interaction…

Cited by 489PDFScholar
2015

Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books

ICCV 2015oral

Books are a rich source of both fine-grained information, how a character, an object or a scene looks like, as well as high-level semantics, what someone is thinking, feeling and how these states evolve through a story. This paper aims to align books to their movie releases in order to provide rich…

Cited by 3512PDFScholar
2015

Skip-Thought Vectors

NeurIPS 2015poster

We describe an approach for unsupervised learning of a generic, distributed sentence encoder. Using the continuity of text from books, we train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded passage. Sentences that share semantic and syntactic properties a…

2015

Understanding and Predicting Image Memorability at a Large Scale

ICCV 2015poster

Progress in estimating visual memorability has been limited by the small scale and lack of variety of benchmark data. Here, we introduce a novel experimental procedure to objectively measure human memory, building the largest annotated image memorability dataset to date (with 60,000 labeled images f…

Cited by 435PDFScholar