← Search

Luping Liu

12 accepted papers

2026

SpatialHand: Generative Object Manipulation from 3D Prespective

ICLR 2026poster

We introduce SpatialHand, a novel framework for generative object insertion with precise 3D control. Current generative object manipulation methods primarily operate within the 2D image plane, but often fail to grasp 3D scene complexities, leading to ambiguities in an object's 3D position, orientati…

Cited by 0SourceScholar
2025

Improving Long-Text Alignment for Text-to-Image Diffusion Models

ICLR 2025poster

The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitations, and aligning the generated images with long texts becomes challenging. To ta…

2025

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025poster

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Meanwhile, multimodal representation models have emerged as the foundation for these versatile multimodal understanding and generation pipeline. Models like CLIP, CLAP and ImageBind…

Cited by 11SourcePDFScholar
2025

T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback

ACL 2025long

Text-to-audio (T2A) generation has achieved remarkable progress in generating a variety of audio outputs from language prompts. However, current state-of-the-art T2A models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio…

Cited by 0SourcePDFScholar
2024

Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

NeurIPS 2024poster

Recent advancements in 3D Large Language Models (LLMs) have demonstrated promising capabilities for 3D scene understanding. However, previous methods exhibit deficiencies in general referencing and grounding capabilities for intricate scene comprehension. In this paper, we introduce the use of objec…

2024

Extending Multi-modal Contrastive Representations

NeurIPS 2024poster

Multi-modal contrastive representation (MCR) of more than three modalities is critical in multi-modal learning. Although recent methods showcase impressive achievements, the high dependence on large-scale, high-quality paired data and the expensive training costs limit their further development. Ins…

2024

FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion

ICML 2024poster

Unified multi-model representation spaces are the foundation of multimodal understanding and generation. However, the billions of model parameters and catastrophic forgetting problems make it challenging to further enhance pre-trained unified spaces. In this work, we propose FreeBind, an idea that t…

2024

InstructSpeech: Following Speech Editing Instructions via Large Language Models

ICML 2024poster

Instruction-guided speech editing aims to follow the user's natural language instruction to manipulate the semantic and acoustic attributes of a speech. In this work, we construct triplet paired data (instruction, input speech, output speech) to alleviate data scarcity and train a multi-task large l…

2024

Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners

ACL 2024long

Large language models (LLMs) have successfully served as a general-purpose interface across multiple tasks and languages, while the adaptation of voice LLMs is mostly designed for specific purposes (either single-task or monolingual), where the advantages of LLMs especially for low-resource language…

2023

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

ICML 2023poster

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high-quality text-audio pairs, and the complexity of modeling long continuous audio…

2023

PTQD: Accurate Post-Training Quantization for Diffusion Models

NeurIPS 2023poster

Diffusion models have recently dominated image synthesis and other related generative tasks. However, the iterative denoising process is expensive in computations at inference time, making diffusion models less practical for low-latency and scalable real-world applications. Post-training quantizati…