← Search

Qing Liu

50 accepted papers

2026

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

ICML 2026poster

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation enc…

Cited by 0SourceScholar
2026

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning

ICLR 2026oral

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due…

Cited by 0SourcecodeScholar
2026

EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question Answering

ICML 2026poster

Knowledge-based Visual Question Answering (KB-VQA) remains a challenging task, particularly when queries require precise identification and grounding of fine-grained entities within large-scale knowledge base. Existing methods often treat visual and textual signals in isolation and rely heavily on i…

Cited by 0SourceScholar
2026

HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation

CVPR 2026

Recent unified models integrate understanding experts (e.g., LLMs) with generative experts (e.g., diffusion models), achieving strong multimodal performance. However, recent advanced methods such as BAGEL and LMFusion follow the Mixture-of-Transformers (MoT) paradigm, adopting a symmetric design tha

Cited by 0SourceScholar
2026

UniSER: A Foundation Model for Unified Soft Effects Removal

CVPR 2026

Digital images are often degraded by soft effects such as lens flare, haze, shadows, and reflections, which reduce aesthetics even though the underlying pixels remain partially visible. The prevailing works address these degradations in isolation, developing highly specialized, specialist models tha

Cited by 0SourceScholar
2025

A Triangular Stable Node Network based on Self-supervised Learning for personalized prediction

ICASSP 2025accepted

In recent years, research has illuminated the potency of implicit data processing in enhancing user preferences. Nevertheless, barriers remain in breaking through the constraints of implicit information. This study aims to bridge this gap by firstly constructing a triangular stable node network mode…

Cited by 0SourceScholar
2025

Baking Gaussian Splatting into Diffusion Denoiser for Fast and Scalable Single-stage Image-to-3D Generation and Reconstruction

ICCV 2025poster

Existing feedforward image-to-3D methods mainly rely on 2D multi-view diffusion models that cannot guarantee 3D consistency. These methods easily collapse when changing the prompt view direction and mainly handle object-centric cases. In this paper, we propose a novel single-stage 3D diffusion model…

2025

Causal fMRI-Mamba: Causal State Space Model for Neural Decoding and Brain Task States Recognition

ICASSP 2025accepted

Deep learning advances neural decoding in functional magnetic resonance imaging (fMRI) tasks with convolution and attention-based methods. However, these methods struggle with capturing global spatiotemporal information due to high dimensionality, noise and inter-individual difference of fMRI, which…

Cited by 0SourceScholar
2025

CompleteMe: Reference-based Human Image Completion

ICCV 2025poster

Recent methods for human image completion can reconstruct plausible body shapes but often fail to preserve unique details, such as specific clothing patterns or distinctive accessories, without explicit reference images. Even state-of-the-art reference-based inpainting approaches struggle to accurat…

Cited by 0SourcePDFScholar
2025

Enhancing Facial Privacy Protection via Weakening Diffusion Purification

CVPR 2025poster

The rapid growth of social media has led to the widespread sharing of individual portrait images, which pose serious privacy risks due to the capabilities of automatic face recognition (AFR) systems for mass surveillance. Hence, protecting facial privacy against unauthorized AFR systems is essential…

2025

FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity

CVPR 2025poster

The advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate integration of visual and textual information across various applications, including image and video captioning, visual question answering, and cross-modal retrieva…

Cited by 6SourcePDFScholar
2025

Generative Image Layer Decomposition with Visual Effects

CVPR 2025poster

Recent advancements in large generative models, particularly diffusion-based methods, have significantly enhanced the capabilities of image editing. However, achieving precise control over image composition tasks remains a challenge. Layered representations, which allow for independent editing of im…

Cited by 1SourcePDFScholar
2025

Generative Video Propagation

CVPR 2025poster

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be addressed in a unified way by leveraging the generative power of su…

Cited by 1SourcePDFScholar
2025

HydraRAG: Structured Cross-Source Enhanced Large Language Model Reasoning

EMNLP 2025

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by incorporating external knowledge. Current hybrid RAG system retrieves evidence from both knowledge graphs (KGs) and text documents to support LLM reasoning. However, it faces challenges like handling multi-hop reasoning, m

2025

MMAPG: A Training-Free Framework for Multimodal Multi-hop Question Answering via Adaptive Planning Graphs

EMNLP 2025

Multimodal Multi-hop question answering requires integrating information from diverse sources, such as images and texts, to derive answers. Existing methods typically rely on sequential retrieval and reasoning, where each step builds on the previous output. However, this single-path paradigm makes t

Cited by 0SourcePDFScholar
2025

ObjectMover: Generative Object Movement with Video Prior

CVPR 2025poster

Simple as it seems, moving an object to another location within an image is, in fact, a challenging image-editing task that requires re-harmonizing the lighting, adjusting the pose based on perspective, accurately filling occluded regions, and ensuring coherent synchronization of shadows and reflect…

Cited by 1SourcePDFScholar
2025

UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

CVPR 2025highlight

We introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation…

2025

What If We Recaption Billions of Web Images with LLaMA-3?

ICML 2025poster

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investi…

Cited by 38SourcePDFScholar
2024

Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask Completion

AAAI 2024technical

Amodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a…

2024

Binary Amplitude-Only Hologram Generation for Acoustic End-Effector Design by Physics-based deep learning

IROS 2024poster

Acoustic holography has emerged as a cutting-edge technique for constructing a micro-robot acoustic end-effector for non-contact manipulation. As one of typical implementations of acoustic holography, Binary Amplitude-Only Hologram (BAOH) featured with a simple structure provides an efficient altern…

Cited by 0SourceScholar
2024

GroupDiff: Diffusion-based Group Portrait Editing

ECCV 2024poster

"Group portrait editing is highly desirable since users constantly want to add a person, delete a person, or manipulate existing persons. It is also challenging due to the intricate dynamics of human interactions and the diverse gestures. In this work, we present GroupDiff, a pioneering effort to ta…

2024

PaCaS-WAA: Patch-Based Contrastive Semi-Supervised Learning with Wavelet Guidance and Adaptive Augmentation for Tumour Segmentation

ICASSP 2024accepted

In many image-guided clinical approaches, tumor segmentation is a fundamental and critical step for locating tumor involvement. However, the scarcity of annotated data and the low contrast of medical imaging techniques make it challenging to accurately segment tumors from surrounding tissues using s…

Cited by 0SourceScholar
2024

SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis

ECCV 2024poster

"We present , a new data generation approach that pushes the performance boundaries of state-of-the-art image segmentation models. One major bottleneck of previous data synthesis methods for segmentation is the design of “segmentation labeler module”, which is used to synthesize segmentation masks f…

Cited by 11SourcePDFScholar
2024

SmartMask: Context Aware High-Fidelity Mask Generation for Fine-grained Object Insertion and Layout Control

CVPR 2024poster

The field of generative image inpainting and object insertion has made significant progress with the recent advent of latent diffusion models. Utilizing a precise object mask can greatly enhance these applications. However due to the challenges users encounter in creating high-fidelity masks there i…

Cited by 9SourcePDFScholar
2024

SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing

ECCV 2024poster

"Effective editing of personal content holds a pivotal role in enabling individuals to express their creativity, weaving captivating narratives within their visual stories, and elevate the overall quality and impact of their visual content. Therefore, in this work, we introduce , a novel framework t…

Cited by 16SourcePDFScholar
2024

UniHuman: A Unified Model For Editing Human Images in the Wild

CVPR 2024poster

Human image editing includes tasks like changing a person's pose their clothing or editing the image according to a text prompt. However prior work often tackles these tasks separately overlooking the benefit of mutual reinforcement from learning them jointly. In this paper we propose UniHuman a uni…

2023

A Novel Transformer-Based Pipeline for Lung Cytopathological Whole Slide Image Classification

ICASSP 2023accepted

We propose a novel three-stage Transformer-based methodology for entire cytopathological whole slide image (WSI) classification. The key idea is to leverage Transformer to extract the fine-grained lesion-level features and then progressively aggregate them into intermediate-grained patch-level featu…

Cited by 0SourceScholar
2023

PHOTOSWAP: Personalized Subject Swapping in Images

NeurIPS 2023poster

In an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity. Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the…

Cited by 36SourcePDFScholar
2023

Perceptual Artifacts Localization for Image Synthesis Tasks

ICCV 2023poster

Recent advancements in deep generative models have facilitated the creation of photo-realistic images across various tasks. However, these generated images often exhibit perceptual artifacts in specific regions, necessitating manual correction. In this study, we present a comprehensive empirical exa…

Cited by 23PDFcodeScholar
2023

SceneComposer: Any-Level Semantic Image Synthesis

CVPR 2023highlight

We propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More specifically, the input layout consists of one or more semantic regions with free-form text descriptions and adjustable p…

2023

TOT:Topology-Aware Optimal Transport for Multimodal Hate Detection

AAAI 2023technical

Multimodal hate detection, which aims to identify the harmful content online such as memes, is crucial for building a wholesome internet environment. Previous work has made enlightening exploration in detecting explicit hate remarks. However, most of their approaches neglect the analysis of implicit…

Cited by 14SourcePDFScholar
2023

Ultrafast Acoustic Holography with Physics-Reinforced Self-Supervised Learning for Precise Robotic Manipulation

IROS 2023poster

Ultrafast acoustic holography (AH) enabling dynamic contactless micro-nano robotic manipulation has recently attracted wide attention. As an advanced technique, AH encodes specific three-dimensional (3D) acoustic field on a two-dimensional (2D) hologram whereby realizing holographic reconstruction w…

Cited by 1SourceScholar
2022

Few-shot Named Entity Recognition with Self-describing Networks

ACL 2022long

Few-shot NER needs to effectively capture information from limited instances and transfer useful knowledge from external resources. In this paper, we propose a self-describing mechanism for few-shot NER, which can effectively leverage illustrative instances and precisely transfer knowledge from exte…

2022

Learning Deep Pathological Features for WSI-Level Cervical Cancer Grading

ICASSP 2022accepted

Fully automated cervical cancer grading on the level of Whole Slide Images (WSI) is a challenge task. As WSIs are in gigapixel resolution, it is impossible to train a deep classification neural network with the entire WSIs as inputs. To bypass this problem, we propose a two-stage learning framework.…

Cited by 0SourceScholar
2022

Learning Part Segmentation Through Unsupervised Domain Adaptation From Synthetic Vehicles

CVPR 2022oral

Part segmentations provide a rich and detailed part-level description of objects. However, their annotation requires an enormous amount of work, which makes it difficult to apply standard deep learning methods. In this paper, we propose the idea of learning part segmentation through unsupervised dom…

Cited by 28PDFcodeScholar
2022

MEJIGCLU: More Effective Jigsaw Clustering For Unsupervised Visual Representation Learning

ICASSP 2022accepted

Unsupervised visual representation learning aims to learn general features from unlabelled data. Early methods design intra-image pretext tasks as learning targets and can be achieved with low computational overhead but unsatisfactory performance. Recent methods introduce contrastive learning and ac…

Cited by 0SourceScholar
2022

PolygonE: Modeling N-ary Relational Data as Gyro-Polygons in Hyperbolic Space

AAAI 2022technical

N-ary relational knowledge base (KBs) embedding aims to map binary and beyond-binary facts into low-dimensional vector space simultaneously. Existing approaches typically decompose n-ary relational facts into subtuples (entity pairs, triples or quintuples, etc.), and they generally model n-ary relat…

Cited by 6SourcePDFScholar
2022

Rethinking Computer-Aided Pelvis Segmentation

ICASSP 2022accepted

As an important structure connecting the spine and lower limbs, the abnormal pelvis is one of the threats to human health worldwide, leading to millions of deaths every year. Although early diagnosis and treatment can greatly improve the chances of survival, it remains a major challenge, especially…

Cited by 0SourceScholar
2022

Unified Structure Generation for Universal Information Extraction

ACL 2022long

Information extraction suffers from its varying targets, heterogeneous structures, and demand-specific schemas. In this paper, we propose a unified text-to-structure generation framework, namely UIE, which can universally model different IE tasks, adaptively generate targeted structures, and collabo…

2021

Alexa Conversations: An Extensible Data-driven Approach for Building Task-oriented Dialogue Systems

NAACL 2021system demonstrations

Traditional goal-oriented dialogue systems rely on various components such as natural language understanding, dialogue state tracking, policy learning and response generation. Training each component requires annotations which are hard to obtain for every new domain, limiting scalability of such sys…

Cited by 22SourcePDFScholar
2021

Fine-grained Entity Typing via Label Reasoning

EMNLP 2021main

Conventional entity typing approaches are based on independent classification paradigms, which make them difficult to recognize inter-dependent, long-tailed and fine-grained entity types. In this paper, we argue that the implicitly entailed extrinsic and intrinsic dependencies between labels can pro…

2021

Weakly Supervised Instance Segmentation for Videos With Temporal Mask Consistency

CVPR 2021poster

Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of objects and (b) missing object predictions. We show that thes…

Cited by 31PDFScholar
2020

A Bidirectional Context Propagation Network for Urine Sediment Particle Detection in Microscopic Images

ICASSP 2020accepted

The microscopic urine sediment examination is a crucial part in the evaluation of renal and urinary tract diseases. Recently, there are emerging CNNs-based detectors to detect the urine sediment particles in an end-to-end manner. However, it is not very compatible to transfer CNNs-based detector dir…

Cited by 0SourceScholar
2020

Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial Occlusion

CVPR 2020poster

Recent work has shown that deep convolutional neural networks (DCNNs) do not generalize well under partial occlusion. Inspired by the success of compositional models at classifying partially occluded objects, we propose to integrate compositional models and DCNNs into a unified deep model with innat…

Cited by 120PDFcodeScholar
2020

Incremental Few-Shot Meta-Learning via Indirect Discriminant Alignment

ECCV 2020poster

We propose a method to train a model so it can learn new classification tasks while improving with each task solved. This amounts to combining meta-learning with incremental learning. Different tasks can have disjoint classes, so one cannot directly align different classifiers as done in model disti…

Cited by 30SourcePDFScholar
2019

Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints From Limited Training Data

ICCV 2019poster

Detecting semantic parts of an object is a challenging task, particularly because it is hard to annotate semantic parts and construct large datasets. In this paper, we present an approach which can learn from a small annotated dataset containing a limited range of viewpoints and generalize to detect…

Cited by 11PDFcodeScholar
2019

Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval

ICCV 2019poster

Sketch-based image retrieval (SBIR) is widely recognized as an important vision problem which implies a wide range of real-world applications. Recently, research interests arise in solving this problem under the more realistic and challenging setting of zero-shot learning. In this paper, we investig…

Cited by 137PDFcodeScholar