← Search

Rui Xu

52 accepted papers

2026

A TEXT-IMAGE FUSION METHOD WITH DATA AUGMENTATION CAPABILITIES FOR REFERRING MEDICAL IMAGE SEGMENTATION

ICASSP 2026poster

Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, known as referring or text-guided image segmentation. However, common augmentations like rotation and flipping disrupt sp…

Cited by 0SourcePDFScholar
2026

CLIP-Guided Unsupervised Semantic-Aware Exposure Correction

ICASSP 2026poster

Improper exposure often leads to severe loss of details, color distortion, and reduced contrast. Exposure correction still faces two critical challenges: (1) the ignorance of object-wise regional semantic information causes the color shift artifacts; (2) real-world exposure images generally have no…

Cited by 0SourcePDFScholar
2026

MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment

AAAI 2026technical

Multimodal Action Quality Assessment (AQA) has recently emerged as a promising paradigm. By leveraging complementary information across shared contextual cues, it enhances the discriminative evaluation of subtle intra-class variations in highly similar action sequences. However, partial modalities a

Cited by 0SourcePDFScholar
2026

MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly

CVPR 2026

Scaling artist-designed meshes to high triangle numbers remains challenging for autoregressive generative models. Existing transformer-based methods suffer from long-sequence bottlenecks and limited quantization resolution, primarily due to the large number of tokens required and constrained quantiz

Cited by 0SourcecodeScholar
2026

PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data

ICLR 2026poster

Segmenting 3D objects into parts is a long-standing challenge in computer vision. To overcome taxonomy constraints and generalize to unseen 3D objects, recent works turn to open-world part segmentation. These approaches typically transfer supervision from 2D foundation models, such as SAM, by liftin…

Cited by 0SourcecodeScholar
2025

2.5D Top-K Ranked Multiple Instance Learning to Classify NSCLC PD-L1 Status on CT Images

ICASSP 2025accepted

Classifying the status of NSCLC PD-L1 on chest CT is a cost-effective and non-invasive method. The existing multiple instance learning (MIL) methods are not effective for this task, due to the lack of an efficient feature encoder for 3D instances and ignoring the importance of representative instanc…

Cited by 0SourceScholar
2025

A2ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization

ACL 2025finding

Long context large language models (LLMs) pose significant challenges for efficient serving due to the large memory footprint and high access overhead of KV cache.Retrieval-based KV cache reduction methods can mitigate these challenges, typically by offloading the complete KV cache to CPU and retrie…

2025

Character is Destiny: Can Persona-assigned Language Models Make Personal Choices?

EMNLP 2025

Can Large Language Models (LLMs) simulate humans in making important decisions? Recent research has unveiled the potential of using LLMs to develop role-playing language agents (RPLAs), mimicking mainly the knowledge and tones of various characters. However, imitative decision-making necessitates a

Cited by 0SourcePDFScholar
2025

CoSER: Coordinating LLM-Based Persona Simulation of Established Roles

ICML 2025poster

Role-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper…

2025

Curse of Knowledge: Your Guidance and Provided Knowledge are biasing LLM Judges in Complex Evaluation

EMNLP 2025

As large language models (LLMs) grow more capable, they face increasingly diverse and complex tasks, making reliable evaluation challenging. The paradigm of LLMs as judges has emerged as a scalable solution, yet prior work primarily focuses on simple settings. Their reliability in complex tasks—wher

Cited by 0SourcePDFScholar
2025

DEFOM-Stereo: Depth Foundation Model Based Stereo Matching

CVPR 2025poster

Stereo matching is a key technique for metric depth estimation in computer vision and robotics. Real-world challenges like occlusion and non-texture hinder accurate disparity estimation from binocular matching cues. Recently, monocular relative depth estimation has shown remarkable generalization us…

2025

DanceFix: An Exploration in Group Dance Neatness Assessment Through Fixing Abnormal Challenges of Human Pose

AAAI 2025technical

The fair and objective assessment of performances and competitions is a common pursuit and challenge in human society. The application of computer vision technology offers hope for this purpose, but it still faces obstacles such as occlusion and motion blur. To address these hindrances, our DanceFix…

Cited by 0SourcePDFScholar
2025

Delving into Transformer-based Network Architecture for Guided Depth Super-Resolution

ICASSP 2025accepted

Guided Depth Super-Resolution (GDSR) enhances low-resolution (LR) depth maps by leveraging high-resolution (HR) color images. The primary challenges involve achieving effective cross-modal data alignment and fusion, as well as incorporating multi-scale information within the Transformer architecture…

Cited by 0SourceScholar
2025

FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data

NAACL 2025long

Prior research on training grounded factuality classification models to detect hallucinations in large language models (LLMs) has relied on public natural language inference (NLI) data and synthetic data. However, conventional NLI datasets are not well-suited for document-level reasoning, which is c…

2025

Few-shot Image Classification based on Attribute Prediction and Selection

ICASSP 2025accepted

Few-shot learning addresses the challenges of image classification with limited samples, but current methods often fail to fully utilize sample correlations and external semantic information, leading to low accuracy. To overcome these limitations, we propose a few-shot image classification method ba…

Cited by 0SourceScholar
2025

Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents

EMNLP 2025

Recent advances in Large Language Model (LLM)-based Role-Playing Language Agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in many tasks for LLMs, the internal thinking processes of RPLAs remain unexplored. Understanding cha

Cited by 0SourcePDFScholar
2025

Language-Guided Audio-Visual Learning for Long-Term Sports Assessment

CVPR 2025poster

Long-term sports assessment is a challenging task in video understanding since it requires judging complex movement variations and action-music coordination. However, there is no direct correlation between the diverse background music and movements in sporting events. Previous works require a large…

2025

Lifelong Test-Time Adaptation via Online Learning in Tracked Low-Dimensional Subspace

NeurIPS 2025poster

Test-time adaptation (TTA) aims to adapt a source model to a target domain using only test data. Existing methods predominantly rely on unsupervised entropy minimization or its variants, which suffer from degeneration, leading to trivial solutions with low-entropy but inaccurate predictions. In this…

Cited by 0SourceScholar
2025

Mining Scene Structural Guidance for Thermal Images in Self-Supervised Monocular Depth Estimation

ICASSP 2025accepted

Self-supervised monocular depth estimation from RGB images has seen significant advancements recently, primarily because it eliminates the need for ground truth data during training. However, applying this technique to thermal images remains challenging due to their inherent characteristics, such as…

Cited by 0SourceScholar
2025

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

NeurIPS 2025spotlight

Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces chall…

Cited by 0SourceScholar
2025

Self-Supervised Monocular Depth Estimation from Videos via Pose-Adaptive Reconstruction

ICASSP 2025accepted

Self-supervised depth estimation from videos involves predicting the depth map of a target frame and the pose changes between source and target frames. The reconstructed source frame is aligned with the target view using the predicted pose and depth information. Precise pose estimation significantly…

Cited by 0SourceScholar
2025

Wasserstein-Regularized Conformal Prediction under General Distribution Shift

ICLR 2025poster

Conformal prediction yields a prediction set with guaranteed $1-\alpha$ coverage of the true target under the i.i.d. assumption, which can fail and lead to a gap between $1-\alpha$ and the actual coverage. Prior studies bound the gap using total variation distance, which cannot identify the gap cha…

Cited by 0SourcePDFScholar
2024

Bridging the Gap: Sketch to Color Diffusion Model with Semantic Prompt Learning

ICASSP 2024accepted

Automatic anime sketch colorization aims to generate a color image from a sketch image, which is challenging due to limited structure and semantic understanding, leading to constrained style, and semantic color inconsistency. In this paper, we introduce a sketch to color diffusion model with semanti…

Cited by 0SourceScholar
2024

Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data

EMNLP 2024finding

Role-playing agents (RPA) have been a popular application area for large language models (LLMs), attracting significant interest from both industry and academia. While existing RPAs well portray the characters’ knowledge and tones, they face challenges in capturing their minds, especially for small…

2024

Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works

EMNLP 2024main

Large language models (LLMs) have demonstrated impressive performance and spurred numerous AI applications, in which role-playing agents (RPAs) are particularly popular, especially for fictional characters. The prerequisite for these RPAs lies in the capability of LLMs to understand characters from…

2024

InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews

ACL 2024long

Role-playing agents (RPAs), powered by large language models, have emerged as a flourishing field of applications. However, a key challenge lies in assessing whether RPAs accurately reproduce the personas of target characters, namely their character fidelity. Existing methods mainly focus on the kno…

2024

Masked Motion Prediction with Semantic Contrast for Point Cloud Sequence Learning

ECCV 2024poster

"Self-supervised representation learning on point cloud sequences is a challenging task due to the complex spatio-temporal structure. Most recent attempts aim to train the point cloud sequences representation model by reconstructing the point coordinates or designing frame-level contrastive learning…

2024

Vision-Language Action Knowledge Learning for Semantic-Aware Action Quality Assessment

ECCV 2024poster

"Action quality assessment (AQA) is a challenging vision task that requires discerning and quantifying subtle differences in actions from the same class. While recent research has made strides in creating fine-grained annotations for more precise analysis, existing methods primarily focus on coarse…

Cited by 6SourcePDFScholar
2024

Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation

AAAI 2024technical

New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs). We present Xiezhi, the most comprehensive evaluation suite designed to assess holistic domain knowledge.Xiezhi comprises multiple-choice questions across 516 diverse…

2023

Adaptive Submanifold-Preserving Sparse Regression for Feature Selection And Multiclass Classification

ICASSP 2023accepted

In this paper, we propose a novel embedded feature selection method, which is able to select the informative and discriminative features with the underlying submanifolds of data in intra-class being well preserved so as to improve the classification performance. Specifically, we first impose the l <…

Cited by 0SourceScholar
2023

Converge to the Truth: Factual Error Correction via Iterative Constrained Editing

AAAI 2023technical

Given a possibly false claim sentence, how can we automatically correct it with minimal editing? Existing methods either require a large number of pairs of false and corrected claims for supervised training or do not handle well errors spanning over multiple tokens within an utterance. In this paper…

2023

Cross-Modality depth Estimation via Unsupervised Stereo RGB-to-infrared Translation

ICASSP 2023accepted

Existing depth estimation methods infer scene depth only from stereo visible light (RGB) images. Since RGB imaging is sensitive to changes in light, it’s difficult to estimate depth information accurately in some degraded visibility conditions. In contrast, infrared (IR) imaging captures thermal rad…

Cited by 0SourceScholar
2023

Spteae: A Soft Prompt Transfer Model for Zero-Shot Cross-Lingual Event Argument Extraction

ICASSP 2023accepted

In zero-shot cross-lingual event argument extraction(EAE) task, a model is typically trained on source language datasets and then applied on task language datasets. There is a trend to regard the zero-shot cross-lingual EAE task as a sequence generation task with manual prompts or discrete prompts.…

Cited by 0SourceScholar
2022

Domain Disentangled Generative Adversarial Network for Zero-Shot Sketch-Based 3D Shape Retrieval

AAAI 2022technical

Sketch-based 3D shape retrieval is a challenging task due to the large domain discrepancy between sketches and 3D shapes. Since existing methods are trained and evaluated on the same categories, they cannot effectively recognize the categories that have not been used during training. In this paper,…

Cited by 27SourcePDFScholar
2022

E-KAR: A Benchmark for Rationalizing Natural Language Analogical Reasoning

ACL 2022findings

The ability to recognize analogies is fundamental to human cognition. Existing benchmarks to test word analogy do not reveal the underneath process of analogical reasoning of neural models. Holding the belief that models capable of reasoning should be right for the right reasons, we propose a first-…

Cited by 35SourcePDFScholar
2022

Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical Constraints

NAACL 2022long

Lexically constrained neural machine translation (NMT) draws much industrial attention for its practical usage in specific domains. However, current autoregressive approaches suffer from high latency. In this paper, we focus on non-autoregressive translation (NAT) for this problem for its efficiency…

2022

Pixel-Level and Affinity-Level Knowledge Distillation for Unsupervised Segmentation of Covid-19 Lesions

ICASSP 2022accepted

Automatic segmentation of COVID-19 lesions is essential for computer-aided diagnosis. However, this task remains challenging because widely-used supervised based methods require large-scale annotated data that is difficult to obtain. Although an unsupervised method based on anomaly detection has sho…

Cited by 0SourceScholar
2022

Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation Networks

ICASSP 2022accepted

Stereo matching has been widely used to estimate depth maps in terrestrial environments. However, it is difficult to achieve appealing performance in underwater environments, since adequate underwater stereo data with groundtruth depth information is not easily available for training an underwater d…

Cited by 0SourceScholar
2021

Learning Scene Structure Guidance via Cross-Task Knowledge Transfer for Single Depth Super-Resolution

CVPR 2021poster

Existing color-guided depth super-resolution (DSR) approaches require paired RGB-D data as training examples where the RGB image is used as structural guidance to recover the degraded depth map due to their geometrical similarity. However, the paired data may be limited or expensive to be collected…

Cited by 55PDFScholar
2020

Discovering Symbolic Models from Deep Learning with Inductive Biases

NeurIPS 2020poster

We develop a general approach to distill symbolic representations of a learned deep model by introducing strong inductive biases. We focus on Graph Neural Networks (GNNs). The technique works as follows: we first encourage sparse latent representations when we train a GNN in a supervised setting, th…

2020

Progressive Point Cloud Deconvolution Generation Network

ECCV 2020poster

In this paper, we propose an effective point cloud generation method, which can generate multi-resolution point clouds of the same shape from a latent vector. Specifically, we develop a novel progressive deconvolution network with the learning-based bilateral interpolation. The learning-based bilate…

2020

Retinal Vessel Segmentation via a Semantics and Multi-Scale Aggregation Network

ICASSP 2020accepted

Precise segmentation of retinal vessels is crucial for a computer-aided diagnosis system of retinal fundus images. However, this task remains challenging due to large variations in scales and poor segmentation of capillary vessels. In this paper, we propose a semantics and multi-scale aggregation ne…

Cited by 0SourceScholar
2020

Unsupervised Content-Preserved Adaptation Network for Classification of Pulmonary Textures from Different CT Scanners

ICASSP 2020accepted

Deep network based methods have been proposed for accurate classification of pulmonary textures on CT images. However, such methods well-trained on CT data from one scanner cannot perform well when they are directly applied to the data from other scanners. This domain shift problem is caused by diff…

Cited by 0SourceScholar
2020

When NAS Meets Robustness: In Search of Robust Architectures Against Adversarial Attacks

CVPR 2020poster

Recent advances in adversarial attacks uncover the intrinsic vulnerability of modern deep neural networks. Since then, extensive efforts have been devoted to enhancing the robustness of deep networks via specialized learning algorithms and loss functions. In this work, we take an architectural persp…

Cited by 210PDFcodeScholar
2018

Pulmonary Textures Classification Using A Deep Neural Network with Appearance and Geometry Cues

ICASSP 2018accepted

Classification of pulmonary textures on CT images is essential for the development of a computer-aided diagnosis system of diffuse lung diseases. In this paper, we propose a novel method to classify pulmonary textures by using a deep neural network, which can make full use of appearance and geometry…

Cited by 0SourceScholar