← Search

Yezhou Yang

58 accepted papers

2026

Domain Expansion: A Latent Space Construction Framework for Multi-Task Learning

ICLR 2026poster

Training a single network with multiple objectives often leads to conflicting gradients that degrade shared representations, forcing them into a compromised state that is suboptimal for any single task—a problem we term latent representation collapse. We introduce Domain Expansion, a framework that…

Cited by 0SourceScholar
2026

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

CVPR 2026

We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffusion models at scale. At its core is VibeToken, a novel resolution-agnostic 1D Transformer-based image tokenizer that enc

Cited by 0SourcecodeScholar
2025

AcT2I: Evaluating and Improving Action Depiction in Text-to-Image Models

EMNLP 2025

Text-to-Image (T2I) models have recently achieved remarkable success in generating images from textual descriptions. However, challenges still persist in accurately rendering complex scenes where actions and interactions form the primary semantic focus. Our key observation in this work is that T2I m

2025

DeepShade: Enable Shade Simulation by Text-conditioned Image Generation

IJCAI 2025

Heatwaves pose a significant threat to public health, especially as global warming intensifies. However, current routing systems (e.g., online maps) fail to incorporate shade information due to the difficulty of estimating shades directly from noisy satellite imagery and the limited availability of

2025

EraseFlow: Learning Concept Erasure Policies via GFlowNet-Driven Alignment

NeurIPS 2025spotlight

Erasing harmful or proprietary concepts from powerful text‑to‑image generators is an emerging safety requirement, yet current ``concept erasure'' techniques either collapse image quality, rely on brittle adversarial losses, or demand prohibitive retraining cycles. We trace these limitations to a myo…

Cited by 0SourceScholar
2025

FlowChef: Steering of Rectified Flow Models for Controlled Generations

ICCV 2025poster

Despite recent advances in Rectified Flow Models (RFMs), unlocking their full potential for controlled generation tasks--such as inverse problems and image editing--remains a significant hurdle. Although RFMs and Diffusion Models (DMs) represent state-of-the-art approaches in generative modeling, th…

2025

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions

ICCV 2025poster

Despite recent advances in inversion and instruction-based image editing, existing approaches primarily excel at editing single, prominent objects but significantly struggle when applied to complex scenes containing multiple entities. To quantify this gap, we first introduce **`RefEdit-Bench`**, a r…

Cited by 0SourcePDFScholar
2025

Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation

NeurIPS 2025poster

Recent advances in video generation have enabled high-fidelity video synthesis from user provided prompts. However, existing models and benchmarks fail to capture the complexity and requirements of professional video generation. Towards that goal, we introduce Stable Cinemetrics, a structured evalua…

Cited by 0SourceScholar
2025

VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning

ICLR 2025poster

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly across multiple images remains a significant challenge. To addre…

Cited by 0SourcePDFScholar
2024

ConceptBed: Evaluating Concept Learning Abilities of Text-to-Image Diffusion Models

AAAI 2024technical

The ability to understand visual concepts and replicate and compose these concepts from images is a central goal for computer vision. Recent advances in text-to-image (T2I) models have lead to high definition and realistic image quality generation by learning from large databases of images and their…

2024

ECLIPSE: A Resource-Efficient Text-to-Image Prior for Image Generations

CVPR 2024poster

Text-to-image (T2I) diffusion models notably the unCLIP models (e.g. DALL-E-2) achieve state-of-the-art (SOTA) performance on various compositional T2I benchmarks at the cost of significant computational resources. The unCLIP stack comprises T2I prior and diffusion image decoder. The T2I prior model…

Cited by 21SourcePDFScholar
2024

Getting it Right: Improving Spatial Consistency in Text-to-Image Models

ECCV 2024poster

"One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive investigation of this limitation, while also developing datase…

2024

Lost in Translation? Translation Errors and Challenges for Fair Assessment of Text-to-Image Models on Multilingual Concepts

NAACL 2024short

Benchmarks of the multilingual capabilities of text-to-image (T2I) models compare generated images prompted in a test language to an expected image distribution over a concept set. One such benchmark, “Conceptual Coverage Across Languages” (CoCo-CroLa), assesses the tangible noun inventory of T2I mo…

Cited by 3SourcePDFScholar
2024

On the Robustness of Language Guidance for Low-Level Vision Tasks: Findings from Depth Estimation

CVPR 2024poster

Recent advances in monocular depth estimation have been made by incorporating natural language as additional guidance. Although yielding impressive results the impact of the language prior particularly in terms of generalization and robustness remains unexplored. In this paper we address this gap by…

2024

Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model

EMNLP 2024finding

Despite advancements in text-to-image models, generating images that precisely align with textual descriptions remains challenging due to misalignment in training data. In this paper, we analyze the critical role of caption precision and recall in text-to-image model training. Our analysis of human-…

2024

R.A.C.E.: Robust Adversarial Concept Erasure for Secure Text-to-Image Diffusion Model

ECCV 2024oral

"In the evolving landscape of text-to-image (T2I) diffusion models, the remarkable capability to generate high-quality images from textual descriptions faces challenges with the potential misuse of reproducing sensitive content. To address this critical issue, we introduce Robust Adversarial Concept…

2024

REVISION: Rendering Tools Enable Spatial Fidelity in Vision-Language Models

ECCV 2024poster

"Text-to-Image (T2I) and multimodal large language models (MLLMs) have been adopted in solutions for several computer vision and multimodal learning tasks. However, it has been found that such vision-language models lack the ability to correctly reason over spatial relationships. To tackle this shor…

2024

TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning

EMNLP 2024finding

Zero-shot inference, where pre-trained models perform tasks without specific training data, is an exciting emergent ability of large models like CLIP. Although there has been considerable exploration into enhancing zero-shot abilities in image captioning (IC) for popular datasets such as MSCOCO and…

2024

TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

NeurIPS 2024poster

Contrastive Language-Image Pretraining (CLIP) models maximize the mutual information between text and visual modalities to learn representations. This makes the nature of the training data a significant factor in the efficacy of CLIP for downstream tasks. However, the lack of compositional diversity…

2024

WOUAF: Weight Modulation for User Attribution and Fingerprinting in Text-to-Image Diffusion Models

CVPR 2024poster

The rapid advancement of generative models facilitating the creation of hyper-realistic images from textual descriptions has concurrently escalated critical societal concerns such as misinformation. Although providing some mitigation traditional fingerprinting mechanisms fall short in attributing re…

2024

eTraM: Event-based Traffic Monitoring Dataset

CVPR 2024highlight

Event cameras with their high temporal and dynamic range and minimal memory usage have found applications in various fields. However their potential in static traffic monitoring remains largely unexplored. To facilitate this exploration we present eTraM - a first-of-its-kind fully event-based traffi…

2023

Attributing Image Generative Models using Latent Fingerprints

ICML 2023poster

Generative models have enabled the creation of contents that are indistinguishable from those taken from nature. Open-source development of such models raised concerns about the risks of their misuse for malicious purposes. One potential risk mitigation strategy is to attribute generative models via…

2023

CAROM Air - Vehicle Localization and Traffic Scene Reconstruction from Aerial Videos

ICRA 2023poster

Road traffic scene reconstruction from videos has been desirable by road safety regulators, city planners, researchers, and autonomous driving technology developers. However, it is expensive and unnecessary to cover every mile of the road with cameras mounted on the road infrastructure. This paper p…

Cited by 12SourcecodeScholar
2023

End-to-end Knowledge Retrieval with Multi-modal Queries

ACL 2023long

We investigate knowledge retrieval with multi-modal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval. We curate a new dataset called ReMuQ for benchmarking progress on this task. ReMuQ require…

2022

CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering

EMNLP 2022main

Videos often capture objects, their visible properties, their motion, and the interactions between different objects. Objects also have physical properties such as mass, which the imaging pipeline is unable to directly capture. However, these properties can be estimated by utilizing cues from relati…

2022

Injecting Semantic Concepts Into End-to-End Image Captioning

CVPR 2022poster

Tremendous progress has been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more fl…

Cited by 137PDFcodeScholar
2022

Learning Action-Effect Dynamics for Hypothetical Vision-Language Reasoning Task

EMNLP 2022finding

‘Actions’ play a vital role in how humans interact with the world. Thus, autonomous agents that would assist us in everyday tasks also require the capability to perform ‘Reasoning about Actions & Change’ (RAC). This has been an important research direction in Artificial Intelligence (AI) in general,…

2022

Semantically Distributed Robust Optimization for Vision-and-Language Inference

ACL 2022findings

Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms. While data augmentation techniques have been designed to mitigate against these failure modes, method…

2022

Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual Patterns

ICRA 2022poster

Recent studies demonstrated the vulnerability of control policies learned through deep reinforcement learning against adversarial attacks, raising concerns about the application of such models to risk-sensitive tasks such as autonomous driving. Threat models for these demonstrations are limited to (…

Cited by 15SourcecodeScholar
2022

To Find Waldo You Need Contextual Cues: Debiasing Who’s Waldo

ACL 2022short

We present a debiased dataset for the Person-centric Visual Grounding (PCVG) task first proposed by Cui et al. (2021) in the Who’s Waldo dataset. Given an image and a caption, PCVG requires pairing up a person’s name mentioned in a caption with a bounding box that points to the person in the image.…

2021

Attribute-Guided Adversarial Training for Robustness to Natural Perturbations

AAAI 2021technical

While existing work in robust deep learning has focused on small pixel-level norm-based perturbations, this may not account for perturbations encountered in several real world settings. In many such cases although test data might not be available, broad specifications about the types of perturbation…

2021

CAROM - Vehicle Localization and Traffic Scene Reconstruction from Monocular Cameras on Road Infrastructures

ICRA 2021poster

Traffic monitoring cameras are powerful tools for traffic management and essential components of intelligent road infrastructure systems. In this paper, we present a vehicle localization and traffic scene reconstruction framework using these cameras, dubbed as CAROM, i.e., "CARs On the Map". CAROM p…

Cited by 27SourcecodeScholar
2021

CLEVR_HYP: A Challenge Dataset and Baselines for Visual Question Answering with Hypothetical Actions over Images

NAACL 2021long

Most existing research on visual question answering (VQA) is limited to information explicitly present in an image or a video. In this paper, we take visual understanding to a higher level where systems are challenged to answer questions that involve mentally simulating the hypothetical consequences…

2021

Compressing Visual-Linguistic Model via Knowledge Distillation

ICCV 2021poster

Despite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation(KD) to effectively compress a transformer-based large VL model into a small VL model. The major challenge arises from the inconsis…

Cited by 100PDFcodeScholar
2021

SEED: Self-supervised Distillation For Visual Representation

ICLR 2021poster

This paper is concerned with self-supervised learning for small models. The problem is motivated by our empirical studies that while the widely used contrastive self-supervised learning method has shown great progress on large model training, it does not work well for small models. To address this p…

2021

SMURF: SeMantic and linguistic UndeRstanding Fusion for Caption Evaluation via Typicality Analysis

ACL 2021long

The open-ended nature of visual captioning makes it a challenging area for evaluation. The majority of proposed models rely on specialized training to improve human-correlation, resulting in limited adoption, generalizability, and explainabilty. We introduce “typicality”, a new formulation of evalua…

2021

Weakly Supervised Relative Spatial Reasoning for Visual Question Answering

ICCV 2021poster

Vision-and-language (V&L) reasoning necessitates perception of visual concepts such as objects and actions, understanding semantics and language grounding, and reasoning about the interplay between the two modalities. One crucial aspect of visual reasoning is spatial understanding, which involves un…

Cited by 26PDFScholar
2020

Learning hierarchical behavior and motion planning for autonomous driving

IROS 2020poster

Learning-based driving solution, a new branch for autonomous driving, is expected to simplify the modeling of driving by learning the underlying mechanisms from data. To improve the tactical decision-making for learning-based driving solution, we introduce hierarchical behavior and motion planning (…

Cited by 48SourceScholar
2020

VQA-LOL: Visual Question Answering under the Lens of Logic

ECCV 2020poster

Logical connectives and their implications on the meaning of a natural language sentence are a fundamental aspect of understanding. In this paper, we investigate whether visual question answering (VQA) systems trained to answer a question about an image, are able to answer the logical composition of…

Cited by 103SourcePDFScholar
2020

ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language

ECCV 2020poster

Person search by natural language aims at retrieving a specific person in a large-scale image pool that matches given textual descriptions. While most of the current methods treat the task as a holistic visual and textual feature matching one, we approach it from an attribute-aligning perspective th…

2019

GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment

RA-L 2019

We study the problem of learning a generalizable action policy for an intelligent agent to actively approach an object of interest, in an indoor environment, solely from its visual inputs. While scene-driven or recognition-driven visual navigation has been widely studied, prior efforts suffer severe

Cited by 20SourceScholar
2019

How Shall I Drive? Interaction Modeling and Motion Planning towards Empathetic and Socially-Graceful Driving

ICRA 2019poster

While intelligence of autonomous vehicles (AVs) has significantly advanced in recent years, accidents involving AVs suggest that these autonomous systems lack gracefulness in driving when interacting with human drivers. In the setting of a two-player game, we propose model predictive control based o…

Cited by 20SourceScholar
2018

Active Object Perceiver: Recognition-Guided Policy Learning for Object Searching on Mobile Robots

IROS 2018poster

We study the problem of learning a navigation policy for a robot to actively search for an object of interest in an indoor environment solely from its visual inputs. While scene-driven visual navigation has been widely studied, prior efforts on learning navigation policies for robots to find objects…

Cited by 58SourceScholar
2018

Extrinsic Dexterity Through Active Slip Control Using Deep Predictive Models

ICRA 2018poster

We present a machine learning methodology for actively controlling slip, in order to increase robot dexterity. Leveraging recent insights in deep learning, we propose a Deep Predictive Model that uses tactile sensor information to reason about slip and its future influence on the manipulated object.…

Cited by 13SourceScholar
2018

Stroke Controllable Fast Style Transfer with Adaptive Receptive Fields

ECCV 2018poster

The Fast Style Transfer methods have been recently proposed to transfer a photograph to an artistic style in real-time. This task involves controlling the stroke size in the stylized results, which remains an open challenge. In this paper, we present a stroke controllable style transfer network that…

Cited by 148SourcePDFScholar
2018

Transductive Unbiased Embedding for Zero-Shot Learning

CVPR 2018poster

Most existing Zero-Shot Learning (ZSL) methods have the strong bias problem, in which instances of unseen (target) classes tend to be categorized as one of the seen (source) classes. So they yield poor performance after being deployed in the generalized ZSL settings. In this paper, we propose a stra…

Cited by 254SourcePDFScholar
2017

Fast task-specific target detection via graph based constraints representation and checking

ICRA 2017poster

We present a framework for fast target detection in real-world robotics applications. Considering that an intelligent agent attends to a task-specific object target during execution, our goal is to detect the object efficiently. We propose the concept of early recognition, which influences the candi…

Cited by 1SourceScholar
2017

Unsupervised Linking of Visual Features to Textual Descriptions in Long Manipulation Activities

RA-L 2017

We present a novel unsupervised framework, which links continuous visual features and symbolic textual descriptions of manipulation activity videos. First, we extract the semantic representation of visually observed manipulations by applying a bottom-up approach to the continuous image streams. We t

Cited by 11SourceScholar
2017

What can i do around here? Deep functional scene understanding for cognitive robots

ICRA 2017poster

For robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding…

Cited by 59SourceScholar
2015

Grasp Type Revisited: A Modern Perspective on a Classical Feature for Vision

CVPR 2015poster

The grasp type provides crucial information about human action. However, recognizing the grasp type in unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distortions. In this paper, first we present a convolutional neural network to classify…

Cited by 103SourcePDFScholar
2015

Learning the spatial semantics of manipulation actions through preposition grounding

ICRA 2015poster

In this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcom…

Cited by 64SourceScholar