← Search

Shuang Xu

30 accepted papers

2026

MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios

AAAI 2026technical

Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different s

Cited by 0SourcePDFScholar
2025

Beyond Low-rankness: Guaranteed Matrix Recovery via Modified Nuclear Norm

IJCAI 2025

The nuclear norm (NN) has been widely explored in matrix recovery problems, such as Robust PCA and matrix completion, leveraging the inherent global low-rank structure of the data. In this study, we introduce a new modified nuclear norm (MNN) framework, where the MNN family norms are defined by adop

2025

DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning

NeurIPS 2025poster

Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitat…

Cited by 0SourcecodeScholar
2025

Enhancing Image Editing with Chain-of-Thought Reasoning and Multimodal Large Language Models

ICASSP 2025accepted

Image editing in our daily lives often requires models to first understand user’s intention and then proceed with the editing. Despite significant advancements in image editing technology, understanding and executing complex instructions remains a substantial challenge. Existing image editing models…

Cited by 0SourceScholar
2025

Fast Guaranteed Tensor Recovery with Adaptive Tensor Nuclear Norm

IJCAI 2025

Real-world datasets like multi-spectral images and videos are naturally represented as tensors. However, limitations in data acquisition often lead to corrupted or incomplete tensor data, making tensor recovery a critical challenge. Solving this problem requires exploiting inherent structural patter

2025

Hipandas: Hyperspectral Image Joint Denoising and Super-Resolution by Image Fusion with the Panchromatic Image

ICCV 2025poster

Hyperspectral images (HSIs) are frequently noisy and of low resolution due to the constraints of imaging devices. Recently launched satellites can concurrently acquire HSIs and panchromatic (PAN) images, enabling the restoration of HSIs to generate clean and high-resolution imagery through fusing PA…

2025

Retinex-MEF: Retinex-based Glare Effects Aware Unsupervised Multi-Exposure Image Fusion

ICCV 2025poster

Multi-exposure image fusion (MEF) synthesizes multiple, differently exposed images of the same scene into a single, well-exposed composite. Retinex theory, which separates image illumination from scene reflectance, provides a natural framework to ensure consistent scene representation and effective…

2025

Task-driven Image Fusion with Learnable Fusion Loss

CVPR 2025highlight

Multi-modal image fusion aggregates information from multiple sensor sources, achieving superior visual quality and perceptual features compared to single-source images, often improving downstream tasks. However, current fusion methods for downstream tasks still use predefined fusion objectives that…

2025

VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution

NeurIPS 2025poster

Conventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution…

Cited by 0SourceScholar
2024

Equivariant Multi-Modality Image Fusion

CVPR 2024poster

Multi-modality image fusion is a technique that combines information from different sensors or modalities enabling the fused image to retain complementary features from each modality such as functional highlights and texture details. However effective training of such fusion models is challenging du…

2024

MaDE: Multi-Scale Decision Enhancement for Multi-Agent Reinforcement Learning

ICASSP 2024accepted

In the domain of multi-agent reinforcement learning (MARL), the limited information availability, complex agent interactions, and individual capabilities among agents often pose a bottleneck for effective decision-making. Previous studies frequently fall short due to insufficient consideration of th…

Cited by 0SourceScholar
2023

CDDFuse: Correlation-Driven Dual-Branch Feature Decomposition for Multi-Modality Image Fusion

CVPR 2023poster

Multi-modality (MM) image fusion aims to render fused images that maintain the merits of different modalities, e.g., functional highlight and detailed textures. To tackle the challenge in modeling cross-modality features and decomposing desirable modality-specific and modality-shared features, we pr…

2023

DDFM: Denoising Diffusion Model for Multi-Modality Image Fusion

ICCV 2023oral

Multi-modality image fusion aims to combine different modalities to produce fused images that retain the complementary features of each modality, such as functional highlights and texture details. To leverage strong generative priors and address challenges such as unstable training and lack of inter…

Cited by 210PDFcodeScholar
2023

Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding

ICASSP 2023accepted

Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of…

Cited by 0SourceScholar
2023

Spherical Space Feature Decomposition for Guided Depth Map Super-Resolution

ICCV 2023poster

Guided depth map super-resolution (GDSR), as a hot topic in multi-modal image processing, aims to upsample low-resolution (LR) depth maps with additional information involved in high-resolution (HR) RGB images from the same scene. The critical step of this task is to effectively extract domain-share…

Cited by 35PDFcodeScholar
2022

A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue Systems

ICASSP 2022accepted

Building a human-machine conversational agent is a core problem in Artificial Intelligence, where knowledge has to be integrated into the model effectively. In this paper, we propose a Multi Domain Knowledge Enhanced Matching Network (MDKEMN) to build retrievalbased dialogue systems that could lever…

Cited by 0SourceScholar
2022

Discrete Cosine Transform Network for Guided Depth Map Super-Resolution

CVPR 2022oral

Guided depth super-resolution (GDSR) is an essential topic in multi-modal image processing, which reconstructs high-resolution (HR) depth maps from low-resolution ones collected with suboptimal conditions with the help of HR RGB images of the same scene. To solve the challenges in interpreting the w…

Cited by 130PDFcodeScholar
2022

Improving Cross-Modal Understanding in Visual Dialog Via Contrastive Learning

ICASSP 2022accepted

Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog history. Though existing methods try to deal with the cross-modal understanding in visual dialog, they are still not enough i…

Cited by 0SourceScholar
2021

Consecutive Decoding for Speech-to-text Translation

AAAI 2021technical

Speech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal cross…

2021

Counterfactual Supporting Facts Extraction for Explainable Medical Record Based Diagnosis with Graph Network

NAACL 2021long

Providing a reliable explanation for clinical diagnosis based on the Electronic Medical Record (EMR) is fundamental to the application of Artificial Intelligence in the medical field. Current methods mostly treat the EMR as a text sequence and provide explanations based on a precise medical knowledg…

2021

Deep Gradient Projection Networks for Pan-sharpening

CVPR 2021poster

Pan-sharpening is an important technique for remote sensing imaging systems to obtain high resolution multispectral images. Recently, deep learning has become the most popular tool for pan-sharpening. This paper develops a model-based deep pan-sharpening approach. Specifically, two optimization prob…

Cited by 194PDFcodeScholar
2021

Learning Flexibly Distributional Representation for Low-quality 3D Face Recognition

AAAI 2021technical

Due to the superiority of using geometric information, 3D Face Recognition (FR) has achieved great successes. Existing methods focus on high-quality 3D FR which is unpractical in real scenarios. Low-quality 3D FR is a more realistic scenario but the low-quality data are born with heavy noises. There…

Cited by 16SourcePDFScholar
2021

Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text Translation

AAAI 2021technical

An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding syst…

2020

Asymmetric Two-Stream Architecture for Accurate RGB-D Saliency Detection

ECCV 2020poster

Most existing RGB-D saliency detection methods adopt symmetric two-stream architectures for learning discriminative RGB and depth representations. In fact, there is another level of ambiguity that is often overlooked: if RGB and depth data are necessary to fit into the same network. In this paper, w…

2020

DIDFuse: Deep Image Decomposition for Infrared and Visible Image Fusion

IJCAI 2020poster

Infrared and visible image fusion, a hot topic in the field of image processing, aims at obtaining fused images keeping the advantages of source images. This paper proposes a novel auto-encoder (AE) based fusion network. The core idea is that the encoder decomposes an image into background and detai…

2020

Knowledge Aware Emotion Recognition in Textual Conversations via Multi-Task Incremental Transformer

COLING 2020main

Emotion recognition in textual conversations (ERTC) plays an important role in a wide range of applications, such as opinion mining, recommender systems, and so on. ERTC, however, is a challenging task. For one thing, speakers often rely on the context and commonsense knowledge to express emotions;…

Cited by 54SourcePDFScholar
2018

CBLDNN-Based Speaker-Independent Speech Separation Via Generative Adversarial Training

ICASSP 2018accepted

In this paper, we propose a speaker-independent multi-speaker monaural speech separation system (CBLDNN-GAT) based on convolutional, bidirectional long short-term memory, deep feedforward neural network (CBLDNN) with generative adversarial training (GAT). Our system aims at obtaining better speech q…

Cited by 0SourceScholar
2016

Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis

ICASSP 2016accepted

Though recurrent neural networks (RNNs) using long short-term memory (LSTM) units can address the issue of long-span dependencies across the linguistic inputs and have achieved the state-of-the-art performance for statistical parametric speech synthesis (SPSS), another limitation of the intrinsic un…

Cited by 0SourceScholar