← Search

Svetlana Lazebnik

29 accepted papers

2026

Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping

ICLR 2026poster

Multimodal large language models (MLLMs) often miss small details and spatial relations in cluttered scenes, leading to errors in fine-grained perceptual grounding. We introduce AttWarp, a lightweight method that allocates more resolution to query-relevant content while compressing less informative…

Cited by 0SourceScholar
2026

Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations

ICLR 2026poster

This work introduces Robots Imitating Generated Videos (RIGVid), a system that enables robots to perform complex manipulation tasks—such as pouring, wiping, and mixing—purely by imitating AI-generated videos, without requiring any physical demonstrations or robot-specific training. Given a language…

Cited by 0SourcecodeScholar
2025

A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

ICRA 2025

Task specification for robotic manipulation in open-world environments is challenging, requiring flexible and adaptive objectives that align with human intentions and can evolve through iterative feedback. We introduce Iterative Keypoint Reward (IKER), a visually grounded, Python-based reward functi

Cited by 1SourcecodeScholar
2024

Shadows Don't Lie and Lines Can't Bend! Generative Models don't know Projective Geometry...for now

CVPR 2024poster

Generative models can produce impressively realistic images. This paper demonstrates that generated images have geometric features different from those of real images. We build a set of collections of generated images prequalified to fool simple signal-based classifiers into believing they are real.…

2024

ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs

ECCV 2024poster

"Methods for finetuning generative models for concept-driven personalization generally achieve strong results for subject-driven or style-driven generation. Recently, low-rank adaptations () have been proposed as a parameter-efficient way of achieving concept-driven personalization. While recent wor…

2021

Bridging the Imitation Gap by Adaptive Insubordination

NeurIPS 2021poster

In practice, imitation learning is preferred over pure reinforcement learning whenever it is possible to design a teaching agent to provide expert supervision. However, we show that when the teaching agent makes decisions with access to privileged information that is unavailable to the student, this…

Cited by 41SourcePDFScholar
2021

Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-On and Outfit Editing

ICCV 2021poster

We proposes a flexible person generation framework called Dressing in Order (DiOr), which supports 2D pose transfer, virtual try-on, and several fashion editing tasks. The key to DiOr is a novel recurrent generation pipeline to sequentially put garments on a person, so that trying on the same garmen…

Cited by 116PDFcodeScholar
2021

GridToPix: Training Embodied Agents With Minimal Supervision

ICCV 2021poster

While deep reinforcement learning (RL) promises freedom from hand-labeled data, great successes, especially for Embodied AI, require significant work to create supervision via carefully shaped rewards. Indeed, without shaped rewards, i.e., with only terminal rewards, present-day Embodied AI results…

Cited by 24PDFcodeScholar
2021

Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied Agents

ICCV 2021poster

Communication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of emergent forms of communication, we first introduce the collaborat…

Cited by 39PDFScholar
2020

A Cordial Sync: Going Beyond Marginal Policies for Multi-Agent Embodied Tasks

ECCV 2020poster

Autonomous agents must learn to collaborate. It is not scalable to develop a new centralized agent every time a task’s difficulty outpaces a single agent’s abilities. While multi-agent collaboration research has flourished in gridworld-like environments, relatively little work has considered visuall…

2020

Memory-Efficient Incremental Learning Through Feature Adaptation

ECCV 2020poster

We introduce an approach for incremental learning that preserves feature descriptors of training images from previously learned classes, instead of the images themselves, unlike most existing work. Keeping the much lower-dimensional feature embeddings of images reduces the memory footprint significa…

Cited by 226SourcePDFScholar
2019

Two Body Problem: Collaborative Visual Task Completion

CVPR 2019oral

Collaboration is a necessary skill to perform tasks that are beyond one agent's capabilities. Addressed extensively in both conventional and modern AI, multi-agent collaboration has often been studied in the context of simple grid worlds. We argue that there are inherently visual aspects to collabor…

Cited by 98PDFScholar
2018

Conditional Image-Text Embedding Networks

ECCV 2018poster

This paper presents an approach for grounding phrases in images which jointly learns multiple text-conditioned embeddings in a single end-to-end model. In order to differentiate text phrases into semantically distinct subspaces, we propose a concept weight branch that automatically assigns phrases t…

2018

PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning

CVPR 2018poster

This paper presents a method for adding multiple tasks to a single deep neural network while avoiding catastrophic forgetting. Inspired by network pruning techniques, we exploit redundancies in large deep networks to free up parameters that can then be employed to learn new tasks. By performing iter…

2018

Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights

ECCV 2018poster

This work presents a method for adapting a single, fixed deep neural network to multiple tasks without affecting performance on already learned tasks. By building upon ideas from network quantization and pruning, we learn binary masks that ``piggyback'' on an existing network, or are applied to unmo…

2018

Two Can Play This Game: Visual Dialog With Discriminative Question Generation and Answering

CVPR 2018poster

Human conversation is a complex mechanism with subtle nuances. It is hence an ambitious goal to develop artificial intelligence agents that can participate fluently in a conversation. While we are still far from achieving this goal, recent progress in visual question answering, image captioning, and…

Cited by 101SourcePDFScholar
2017

Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding Space

NeurIPS 2017poster

This paper explores image caption generation using conditional variational auto-encoders (CVAEs). Standard CVAEs with a fixed Gaussian prior yield descriptions with too little variability. Instead, we propose two models that explicitly structure the latent space around K components corresponding to…

2017

Phrase Localization and Visual Relationship Detection With Comprehensive Image-Language Cues

ICCV 2017poster

This paper presents a framework for localization or grounding of phrases in images using a large collection of linguistic and visual cues. We model the appearance, size, and position of entity bounding boxes, adjectives that contain attribute information, and spatial relationships between pairs of e…

Cited by 230PDFcodeScholar
2015

Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models

ICCV 2015poster

The Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated boundi…

Cited by 2475PDFScholar
2015

Where to Buy It: Matching Street Clothing Photos in Online Shops

ICCV 2015oral

In this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differences between street photos (pictures of people wearing clothing in everyday unco…

Cited by 578PDFScholar