← Search

Yuheng Li

27 accepted papers

2026

Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing

AAAI 2026technical

Recent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we i

Cited by 0SourcePDFScholar
2026

From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning

ICML 2026poster

Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive “shortcut” issue: tasks can reach up to 99% …

Cited by 0SourceScholar
2026

Group Diffusion: Enhancing Image Generation by Unlocking Cross-Sample Collaboration

CVPR 2026

In this work, we explore an untapped signal in diffusion model inference. While all previous methods generate images independently at inference, we instead ask if samples can be generated collaboratively. We propose Group Diffusion, unlocking the attention mechanism to be shared across images, rathe

Cited by 0SourcecodeScholar
2026

Learning an Image Editing Model without Image Editing Pairs

ICLR 2026poster

Recent image editing models have achieved impressive results while following natural language editing instructions, but they rely on supervised fine-tuning with large datasets of input-target pairs. This is a critical bottleneck, as such naturally occurring pairs are hard to curate at scale. Curren…

Cited by 0SourcecodeScholar
2026

Meta-FC: Meta-Learning with Feature Consistency for Robust and Generalizable Watermarking

CVPR 2026

Deep learning-based watermarking has made remarkable progress in recent years. To achieve robustness against various distortions, current methods commonly adopt a training strategy where a \underline s ingle \underline r andom \underline d istortion (SRD) is chosen as the noise layer in each trainin

Cited by 0SourcecodeScholar
2026

PAMotion: Physics-Aware Motion Generation for Full-Body Interaction with Multiple Objects

CVPR 2026

We present PAMotion, a physics-aware diffusion framework for generating realistic full-body human interactions with multiple objects.Existing diffusion-based methods that jointly synthesize human and object motions often struggle to capture the intricate physical interactions--especially those invol

Cited by 0SourcecodeScholar
2025

Can Reinforcement Learning Solve Asymmetric Combinatorial-Continuous Zero-Sum Games?

ICLR 2025poster

There have been extensive studies on learning in zero-sum games, focusing on the analysis of the existence and algorithmic convergence of Nash equilibrium (NE). Existing studies mainly focus on symmetric games where the strategy spaces of the players are of the same type and size. For the few studie…

2025

Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks

ICCV 2025poster

Diffusion Transformers (DiTs) can generate short photorealistic videos, yet directly training and sampling longer videos with full attention across the video remains computationally challenging. Alternative methods break long videos down into sequential generation of short video segments, requiring…

2025

X-Fusion: Introducing New Modality to Frozen Large Language Models

ICCV 2025poster

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights, keeping the LLM's parameters frozen while integrating vision-specific informat…

Cited by 0SourcePDFScholar
2025

Yo'Chameleon: Personalized Vision and Language Generation

CVPR 2025poster

Large Multimodal Models (e.g., GPT-4, Gemini, Chameleon) have evolved into powerful tools with millions of users. However, they remain generic models and lack personalized knowledge of specific user concepts. Previous work has explored personalization for text generation, yet it remains unclear how…

Cited by 1SourcePDFScholar
2024

AnatoMask: Enhancing Medical Image Segmentation with Reconstruction-guided Self-masking

ECCV 2024poster

"Due to the scarcity of labeled data, self-supervised learning (SSL) has gained much attention in 3D medical image segmentation, by extracting semantic representations from unlabeled data. Among SSL strategies, Masked image modeling (MIM) has shown effectiveness by reconstructing randomly masked ima…

2024

Edit One for All: Interactive Batch Image Editing

CVPR 2024poster

In recent years image editing has advanced remarkably. With increased human control it is now possible to edit an image in a plethora of ways; from specifying in text what we want to change to straight up dragging the contents of the image in an interactive point-based manner. However most of the fo…

Cited by 4SourcePDFScholar
2024

Towards Automatic Boundary Detection for Human-AI Collaborative Hybrid Essay in Education

AAAI 2024technical

The recent large language models (LLMs), e.g., ChatGPT, have been able to generate human-like and fluent responses when provided with specific instructions. While admitting the convenience brought by technological advancement, educators also have concerns that students might leverage LLMs to complet…

2024

Yo'LLaVA: Your Personalized Language and Vision Assistant

NeurIPS 2024poster

Large Multimodal Models (LMMs) have shown remarkable capabilities across a variety of tasks (e.g., image captioning, visual question answering). While broad, their knowledge remains generic (e.g., recognizing a dog), and they are unable to handle personalized subjects (e.g., recognizing a user's pet…

2023

GLIGEN: Open-Set Grounded Text-to-Image Generation

CVPR 2023poster

Large-scale text-to-image diffusion models have made amazing advances. However, the status quo is to use text input alone, which can impede controllability. In this work, we propose GLIGEN: Open-Set Grounded Text-to-Image Generation, a novel approach that builds upon and extends the functionality of…

2023

Towards Universal Fake Image Detectors That Generalize Across Generative Models

CVPR 2023poster

With generative models proliferating at a rapid rate, there is a growing need for general purpose fake image detectors. In this work, we first show that the existing paradigm, which consists of training a deep network for real-vs-fake classification, fails to detect fake images from newer breeds of…

2023

Visual Instruction Inversion: Image Editing via Image Prompting

NeurIPS 2023poster

Text-conditioned image editing has emerged as a powerful tool for editing images. However, in many situations, language can be ambiguous and ineffective in describing specific image edits. When faced with such challenges, visual prompts can be a more informative and intuitive way to convey ideas. We…

Cited by 48SourcePDFScholar
2023

What Knowledge Gets Distilled in Knowledge Distillation?

NeurIPS 2023poster

Knowledge distillation aims to transfer useful information from a teacher network to a student network, with the primary goal of improving the student's performance for the task at hand. Over the years, there has a been a deluge of novel techniques and use cases of knowledge distillation. Yet, despi…

Cited by 30SourcePDFScholar
2022

Bigger Data or Fairer Data? Augmenting BERT via Active Sampling for Educational Text Classification

COLING 2022main

Pretrained Language Models (PLMs), though popular, have been diagnosed to encode bias against protected groups in the representations they learn, which may harm the prediction fairness of downstream models. Given that such bias is believed to be related to the amount of demographic information carri…

2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

Contrastive Learning for Diverse Disentangled Foreground Generation

ECCV 2022poster

"We introduce a new method for diverse foreground generation with explicit control over various factors. Existing image inpainting based foreground generation methods often struggle to generate diverse results and rarely allow users to explicitly control specific factors of variation (e.g., varying…

2021

Collaging Class-Specific GANs for Semantic Image Synthesis

ICCV 2021poster

We propose a new approach for high resolution semantic image synthesis. It consists of one base image generator and multiple class-specific generators. The base generator generates high quality images based on a segmentation map. To further improve the quality of different objects, we create a bank…

Cited by 43PDFScholar
2020

MixNMatch: Multifactor Disentanglement and Encoding for Conditional Image Generation

CVPR 2020poster

We present MixNMatch, a conditional generative model that learns to disentangle and encode background, object pose, shape, and texture from real images with minimal supervision, for mix-and-match image generation. We build upon FineGAN, an unconditional generative model, to learn the desired disenta…

Cited by 101PDFcodeScholar