← Search

Xianghao Kong

9 accepted papers

2026

Composing Concepts from Images and Videos via Concept-prompt Binding

CVPR 2026

Visual concept composition, which aims to integrate different elements from images and videos into a single, coherent visual output, still falls short in accurately extracting complex concepts from visual inputs and flexibly combining concepts from both images and videos. We introduce Bind & Compose

Cited by 0SourcecodeScholar
2026

SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment

ICLR 2026poster

Recent studies on 3D hand reconstruction have demonstrated the effectiveness of synthetic training data to improve estimation performance. However, most methods rely on game engines to synthesize hand images, which often lack diversity in textures and environments, and fail to include crucial compon…

Cited by 0SourceScholar
2025

Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget

CVPR 2025poster

As scaling laws in generative AI push performance, they simultaneously concentrate the development of these models among actors with large computational resources. With a focus on text-to-image (T2I) generative models, we aim to unlock this bottleneck by demonstrating very low-cost training of large…

2024

Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks

ACL 2024findings

The widespread use of Text-to-Image (T2I) models in content generation requires careful examination of their safety, including their robustness to adversarial attacks. Despite extensive research on adversarial attacks, the reasons for their effectiveness remain underexplored. This paper presents an…

2024

Controllable Navigation Instruction Generation with Chain of Thought Prompting

ECCV 2024poster

"Instruction generation is a vital and multidisciplinary research area with broad applications. Existing instruction generation models are limited to generating instructions in a single style from a particular dataset, and the style and content of generated instructions cannot be controlled. Moreove…

2024

Interpretable Diffusion via Information Decomposition

ICLR 2024poster

Denoising diffusion models enable conditional generation and density modeling of complex relationships like images and text. However, the nature of the learned relationships is opaque making it difficult to understand precisely what relationships between words and parts of an image are captured, or…

2024

Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training

NeurIPS 2024poster

Diffusion models learn to denoise data and the trained denoiser is then used to generate new samples from the data distribution. In this paper, we revisit the diffusion sampling process and identify a fundamental cause of sample quality degradation: the denoiser is poorly estimated in regions that…

2022

3D-SPS: Single-Stage 3D Visual Grounding via Referred Point Progressive Selection

CVPR 2022oral

3D visual grounding aims to locate the referred target object in 3D point cloud scenes according to a free-form language description. Previous methods mostly follow a two-stage paradigm, i.e., language-irrelevant detection and cross-modal matching, which is limited by the isolated architecture. In s…

Cited by 69PDFcodeScholar