← Search

Xueting Wang

4 accepted papers

2026

Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data

AAAI 2026technical

Generative models have become a powerful tool for synthesizing training data in computer vision tasks. Current approaches solely focus on aligning generated images with the target dataset distribution. As a result, they capture only the common features in the real dataset and mostly generate "easy s

Cited by 0SourcePDFScholar
2024

ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions

ACL 2024findings

Image retrieval from contextual descriptions (IRCD) aims to identify an image within a set of minimally contrastive candidates based on linguistically complex text. Despite the success of VLMs, they still significantly lag behind human performance in IRCD. The main challenges lie in aligning key con…

2024

Self-supervised 6-DoF Robot Grasping by Demonstration via Augmented Reality Teleoperation System

ICRA 2024poster

Most existing 6-DoF robot grasping solutions depend on strong supervision on grasp pose to ensure satisfactory performance, which could be laborious and impractical when the robot works in some restricted area. To this end, we propose a self-supervised 6-DoF grasp pose detection framework via an Aug…

Cited by 3SourceScholar
2024

The Lottery Ticket Hypothesis in Denoising: Towards Semantic-Driven Initialization

ECCV 2024poster

"Text-to-image diffusion models allow users control over the content of generated images. Still, text-to-image generation occasionally leads to generation failure requiring users to generate dozens of images under the same text prompt before they obtain a satisfying result. We formulate the lottery…