← Search

Xueming Qian

21 accepted papers

2026

Dynamic Multi-Path Retrieval for Knowledge-based Visual Question Answering

IJCAI 2026

Knowledge-based Visual Question Answering (KB-VQA) requires models to answer visual questions by reasoning over external knowledge beyond the given image. Existing approaches suffer from two main limitations. First, candidate knowledge is often retrieved in a single modality, either textual or visua

Cited by 0Scholar
2026

Whole-Field Action Sensing via Wearable Single-Channel EMG Sensors and Resource-Efficient Motion Network

AAAI 2026technical

The proliferation of collaborative training and multi-person sports has underscored the necessity for concurrent whole-field action sensing. However, Electromyography (EMG) recognition, which plays a pivotal role in Wearable Human Activity Recognition (WHAR) for analyzing muscle activity and decodin

Cited by 0SourcePDFScholar
2025

A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs

ACL 2025finding

Temporal knowledge graph reasoning aims to predict future events with knowledge of existing facts and plays a key role in various downstream tasks. Previous methods focused on either graph structure learning or semantic reasoning, failing to integrate dual reasoning perspectives to handle different…

2025

Frequency Domain-Based Diffusion Model for Unpaired Image Dehazing

ICCV 2025poster

Unpaired image dehazing has attracted increasing attention due to its flexible data requirements during model training. Dominant methods based on contrastive learning not only introduce haze-unrelated content information, but also ignore haze-specific properties in the frequency domain (i.e., haze-r…

Cited by 0SourcePDFScholar
2025

Learning Deblurring Texture Prior from Unpaired Data with Diffusion Model

ICCV 2025poster

Since acquiring large amounts of realistic blurry-sharp image pairs is difficult and expensive, learning blind image deblurring from unpaired data is a more practical and promising solution. Unfortunately, most existing approaches only use adversarial learning to bridge the gap from blurry domains t…

Cited by 0SourcePDFScholar
2025

M3Net: Efficient Time-Frequency Integration Network with Mirror Attention for Audio Classification on Edge

AAAI 2025technical

Audio classification plays a crucial role within fields such as human-machine interaction and intelligent robotics. However, high-performance audio classification systems typically demand significant computational and storage resources, posing substantial challenges when deploying to the resource-co…

2024

Decoupling Degradations with Recurrent Network for Video Restoration in Under-Display Camera

AAAI 2024technical

Under-display camera (UDC) systems are the foundation of full-screen display devices in which the lens mounts under the display. The pixel array of light-emitting diodes used for display diffracts and attenuates incident light, causing various degradations as the light intensity changes. Unlike gene…

2024

Learning to Paraphrase for Alignment with LLM Preference

EMNLP 2024finding

Large Language Models (LLMs) exhibit the issue of paraphrase divergence. This means that when a question is phrased in a slightly different but semantically similar way, LLM may output a wrong response despite being able to answer the original question correctly. Previous research has regarded this…

2024

Motion-adaptive Separable Collaborative Filters for Blind Motion Deblurring

CVPR 2024poster

Eliminating image blur produced by various kinds of motion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in feature space. These practices not only prevent the capture of spatially variable mot…

2024

Pseudo-Label Enhanced Prototypical Contrastive Learning for Uniformed Intent Discovery

EMNLP 2024finding

New intent discovery is a crucial capability for task-oriented dialogue systems. Existing methods focus on transferring in-domain (IND) prior knowledge to out-of-domain (OOD) data through pre-training and clustering stages. They either handle the two processes in a pipeline manner, which exhibits a…

2024

QueryCDR: Query-based Controllable Distortion Rectification Network for Fisheye Images

ECCV 2024poster

"Fisheye image rectification aims to correct distortions in images taken with fisheye cameras. Although current models show promising results on images with a similar degree of distortion as the training data, they will produce sub-optimal results when the degree of distortion changes and without re…

2023

A Diffusion Model with Contrastive Learning for ICU False Arrhythmia Alarm Reduction

IJCAI 2023poster

The high rate of false arrhythmia alarms in intensive care units (ICUs) can negatively impact patient care and lead to slow staff response time due to alarm fatigue. To reduce false alarms in ICUs, previous works proposed conventional supervised learning methods which have inherent limitations in de…

2023

CSDA: Learning Category-Scale Joint Feature for Domain Adaptive Object Detection

ICCV 2023poster

Domain Adaptive Object Detection (DAOD) aims to improve the detection performance of target domains by minimizing the feature distribution between the source and target domain. Recent approaches usually align such distributions in terms of categories through adversarial learning and some progress ha…

Cited by 14PDFScholar
2023

FSI: Frequency and Spatial Interactive Learning for Image Restoration in Under-Display Cameras

ICCV 2023poster

Under-display camera (UDC) systems remove the screen notch for bezel-free displays and provide a better interactive experience. The main challenge is that the pixel array of light-emitting diodes used for display diffracts and attenuates the incident light, leading to complex degradation. Existing m…

Cited by 22PDFScholar
2023

Learning Data-Driven Vector-Quantized Degradation Model for Animation Video Super-Resolution

ICCV 2023poster

Existing real-world video super-resolution (VSR) methods focus on designing a general degradation pipeline for open-domain videos while ignoring data intrinsic characteristics which strongly limit their performance when applying to some specific domains (e.g., animation videos). In this paper, we th…

Cited by 5PDFcodeScholar
2023

Space-time Prompting for Video Class-incremental Learning

ICCV 2023oral

Recently, prompt-based learning has made impressive progress on image class-incremental learning, but it still lacks sufficient exploration in the video domain. In this paper, we will fill this gap by learning multiple prompts based on a powerful image-language pre-trained model, i.e., CLIP, making…

Cited by 11PDFScholar
2022

Learning a Condensed Frame for Memory-Efficient Video Class-Incremental Learning

NeurIPS 2022accept

Recent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning…

Cited by 20SourcePDFScholar
2022

Semi-supervised New Slot Discovery with Incremental Clustering

EMNLP 2022finding

Discovering new slots is critical to the success of dialogue systems. Most existing methods rely on automatic slot induction in unsupervised fashion or perform domain adaptation across zero or few-shot scenarios. They have difficulties in providing high-quality supervised signals to learn clustering…

Cited by 10SourcePDFScholar
2021

AINet: Association Implantation for Superpixel Segmentation

ICCV 2021poster

Recently, some approaches are proposed to harness deep convolutional networks to facilitate superpixel segmentation. The common practice is to first evenly divide the image into a pre-defined number of grids and then learn to associate each pixel with its surrounding grids. However, simply applying…

Cited by 53PDFcodeScholar