← Search

Kele Xu

42 accepted papers

2026

Decouple Your Discovery and Memory in Continual Generalized Category Discovery

CVPR 2026

Continual Generalized Category Discovery (C-GCD) seeks to incrementally discover new categories from unlabeled data and memorize old categories' knowledge, fostering model adaptability in real-world scenarios. Especially, the unlabeled data is from both old and new classes, requiring the model to re

Cited by 0SourceScholar
2026

Geometry-driven OOD Detectors Are Class-Incremental Learners

CVPR 2026

Class-Incremental Learning (CIL) seeks to acquire new classes over time without erasing prior knowledge. While recent methods leverage pre-trained models (PTMs) to curb forgetting, they largely optimize the feature extractor and overlook the crucial classification head. In this work, we advance a si

Cited by 0SourcecodeScholar
2026

Perturbing to Preserve: Defending Fragile Knowledge in Online Continual Learning

AAAI 2026technical

Online continual learning requires models to learn from non‑stationary data streams while retaining prior knowledge. We identify an overlooked phenomenon—knowledge fragility—where correctly learned instances are rapidly forgotten after minor parameter updates. Our analysis attributes this fragility

Cited by 0SourcePDFScholar
2026

Re-evaluating Continual VQA: Toward Fair and Robust Evaluation for Multimodal Continual Learning

CVPR 2026

Continual Visual Question Answering (Continual VQA) poses unique challenges for multimodal continual learning, requiring models to incrementally acquire new knowledge while preserving visual-semantic grounding across tasks. However, existing benchmarks hinder fair and robust evaluation of such capab

Cited by 0SourcecodeScholar
2026

Reliable Confidence Alignment for Generalized Category Discovery

ICML 2026poster

Generalized Category Discovery (GCD) requires models to categorize an unlabeled pool containing both known and novel classes under sparse supervision. We identify a systemic confidence bias inherent in existing parametric methods: while entropy regularization prevents class collapse, it indiscrimina…

Cited by 0SourceScholar
2025

A Counterfactual Ultrasound Anti-Interference Self-Supervised Network for B-mode Ultrasound Tongue Extraction

ICASSP 2025accepted

B-mode ultrasound tongue imaging is a non-invasive and real-time method for visualizing vocal tract deformation. However, accurately extracting the tongue’s surface contour remains a significant challenge due to the low signal-to-noise ratio (SNR) and prevalent speckle noise in ultrasound images. Tr…

Cited by 0SourceScholar
2025

Complementary Learning System Theory-based Active Learning for Audio Classification

ICASSP 2025accepted

Deep learning has significantly advanced the audio classification, achieving remarkable results. However, these successes often rely on extensive manual annotation of audio, a labor-intensive and costly process. Active Learning (AL) presents a promising solution by minimizing the required amount of…

Cited by 0SourceScholar
2025

Debiased Active Learning with Variational Gradient Rectifier

AAAI 2025technical

The strategy of selecting ``most informative'' hard samples in active learning has proven a boon for alleviating the challenges of few-shot learning and costly data annotation in deep learning. However, this very preference towards hard samples engenders bias issues, thereby impeding the full potent…

Cited by 0SourcePDFScholar
2025

Enhancing Decision-Making for LLM Agents via Step-Level Q-Value Models

AAAI 2025technical

Agents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific ta…

Cited by 8SourcePDFScholar
2025

HyperMST: Multi-scale Spatio-Temporal Hypercorrelation Network for POI Recommendation

ICASSP 2025accepted

Point-of-Interest (POI) recommendation has become increasingly important in the trajectory prediction domain. However, most existing approaches focus on a single scale and tend to overemphasize either spatial or temporal aspects. These methods often overlook the temporal dependencies in movement beh…

Cited by 0SourceScholar
2025

Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement Learning

ICML 2025poster

This paper addresses the challenge of discontinuity in goal-achievement capabilities observed in Goal-conditioned Reinforcement Learning (GCRL) algorithms. Through a theoretical analysis, we identify that the reuse of successful trajectories or policies during training can aid in achieving adjacent…

Cited by 0SourcePDFScholar
2025

Investigating the Role of Weight Decay in Enhancing Nonconvex SGD

CVPR 2025poster

Weight decay is a widely used technique in training machine learning models, known to empirically enhance the generalization of Stochastic Gradient Descent (SGD). While intuitively weight decay allows SGD to train a regularized model rather than the original one, there is limited theoretical underst…

Cited by 0SourcePDFScholar
2025

Knowledge Memorization and Rumination for Pre-trained Model-based Class-Incremental Learning

CVPR 2025poster

Class-Incremental Learning (CIL) enables models to continuously learn new classes while mitigating catastrophic forgetting. Recently, Pre-Trained Models (PTMs) have greatly enhanced CIL performance, even when fine-tuning is limited to the first task. This advantage is particularly beneficial for CIL…

2025

Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental Learning

AAAI 2025technical

Logit-based knowledge distillation (KD) is commonly used to mitigate catastrophic forgetting in class-incremental learning (CIL) caused by data distribution shifts. However, the strict match of logit values between student and teacher models conflicts with the cross-entropy (CE) loss objective of le…

2025

Mutual-View Contrastive Generative Framework for Attribute-Missing Graph Clustering

ICASSP 2025accepted

Attribute-Missing Graph Clustering addresses the challenging problem of incomplete node attribute information in graphs. Recent advancements in self-supervised learning techniques, particularly contrastive learning and generative approaches, have shown effectiveness in tackling tasks involving missi…

Cited by 0SourceScholar
2025

Relieving Universal Label Noise for Unsupervised Visible-Infrared Person Re-Identification by Inferring from Neighbors

AAAI 2025technical

Unsupervised visible-infrared person re-identification (USL-VI-ReID) is of great research and practical significance yet remains challenging due to the absence of annotations. Existing approaches aim to learn modality-invariant representations in an unsupervised setting. However, these methods often…

2025

SPEA: Large-Scale Entity Alignment via Self-Partitioning

ICASSP 2025accepted

The task of entity alignment (EA) seeks to identify corresponding entities across different knowledge graphs (KGs). However, in large-scale KG alignment tasks, the complexity of the problem renders traditional entity structure representation methods, designed for small-scale KGs, ineffective. Partit…

Cited by 0SourceScholar
2025

SSAST-Adapter: A Parameter-efficient Incremental Learning Algorithm for Underwater Acoustic Target Recognition

ICASSP 2025accepted

Underwater acoustic target recognition involves identifying and classifying targets in underwater environments using acoustic signals. In recent years, deep learning has made significant progress in this field. However, the models require the entire dataset to be available upfront, and classificatio…

Cited by 0SourceScholar
2025

Scaling Bioacoustic Signal Pre-training with Million Samples Via Mask-Modeling

ICASSP 2025accepted

Deep learning-based bioacoustic audio analysis holds immense potential across various applications. However, existing studies in bioacoustics often focus on a limited number of species, potentially hindering the transferability of models across different species. Furthermore, the manual annotation o…

Cited by 0SourceScholar
2025

Text-guided Multimodal Fusion for the Multimodal Emotion and Intent Joint Understanding

ICASSP 2025accepted

Emotion and Intent Joint Understanding in Multi-modal Conversation is a challenging task in the field of affective computing, aiming to decode the semantic information manifested in the multimodal conversational while simultaneously inferring the emotions and intents of the utterance. To address thi…

Cited by 0SourceScholar
2025

V-Pilot: A Velocity Vector Control Agent for Fixed-Wing UAVs from Imperfect Demonstrations

ICRA 2025

This paper addresses the challenge of Velocity Vector Control (VVC) for fixed-wing UAVs using Reinforcement Learning (RL) in the presence of imperfect demonstrations. The multi-objective and long-horizon nature of VVC introduces significant spatial and temporal complexities, complicating RL's explor

Cited by 0SourceScholar
2025

VVC-Gym: A Fixed-Wing UAV Reinforcement Learning Environment for Multi-Goal Long-Horizon Problems

ICLR 2025poster

Multi-goal long-horizon problems are prevalent in real-world applications. The additional goal space introduced by multi-goal problems intensifies the spatial complexity of exploration; meanwhile, the long interaction sequences in long-horizon problems exacerbate the temporal complexity of explorati…

Cited by 0SourcePDFScholar
2024

Adapter-Based Incremental Learning for Face Forgery Detection

ICASSP 2024accepted

Many existing face forgery detection methods primarily revolve around learning general representations on predefined datasets and subsequently crossing these static representations to other datasets. However, these approaches could lead to catastrophic forgetting in real-world scenarios, especially…

Cited by 0SourceScholar
2024

Iterative Regularized Policy Optimization with Imperfect Demonstrations

ICML 2024poster

Imitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employ…

2024

Optimistic Model Rollouts for Pessimistic Offline Policy Optimization

AAAI 2024technical

Model-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decis…

Cited by 1SourcePDFScholar
2024

Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language Tasks

NeurIPS 2024poster

Continual learning (CL) empowers pre-trained vision-language (VL) models to efficiently adapt to a sequence of downstream tasks. However, these models often encounter challenges in retaining previously acquired skills due to parameter shifts and limited access to historical data. In response, recent…

2024

Transformer-Inspired Lightweight Model for Efficient Time Series Forecasting

ICASSP 2024accepted

Accuracy and efficiency are pivotal considerations in the field of time series forecasting. Through the integration of meticulously designed temporal components, the Transformer-based models have significantly enhanced the accuracy of time series prediction. However, due to the utilization of attent…

Cited by 0SourceScholar
2023

Complementary Learning System Based Intrinsic Reward in Reinforcement Learning

ICASSP 2023accepted

Deep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learni…

Cited by 0SourceScholar
2023

Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm Regularization

ICASSP 2023accepted

The use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challeng…

Cited by 0SourceScholar
2023

Progressive Diversifying Policy for Multi-Agent Reinforcement Learning

ICASSP 2023accepted

Multi-Agent Reinforcement Learning (MARL) has recently achieved promising performance in many collaborative decision making tasks. However, one of the main bottleneck challenges for MARL is the sparsity of the team reward, which can lead to the homogenization of agents’ behaviors. To address these i…

Cited by 0SourceScholar
2023

Raw Ultrasound-Based Phonetic Segments Classification Via Mask Modeling

ICASSP 2023accepted

Ultrasound tongue imaging is widely used in clinical linguistics and phonetics. Recently, deep neural networks, especially convolutional neural networks, have been widely used in the interpretation and analysis of ultrasound tongue images (UTI). Despite achieving satisfactory performance, deep model…

Cited by 0SourceScholar
2022

FINT: Field-Aware Interaction Neural Network for Click-Through Rate Prediction

ICASSP 2022accepted

As a critical component for online advertising and marketing, click-through rate (CTR) prediction has drawn lots of attention from both industry and academia. Recently, deep learning has become the mainstream methodological choice for CTR. Despite sustainable efforts have been made, existing approac…

Cited by 0SourceScholar
2022

Improving the Classification of Phonetic Segments from Raw Ultrasound Using Self-Supervised Learning and Hard Example Mining

ICASSP 2022accepted

Ultrasound tongue imaging is an attractive way for speech production study as it provides an effective visualization for the vocal tract. Automatic classification of phonetic segments (tongue shapes) from raw ultrasound data is vital for further interpretation. Recently, deep learning-based approach…

Cited by 0SourceScholar
2022

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

IJCAI 2022poster

We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introd…

2021

Fden: Mining Effective Information of Features in Detecting Network Anomalies

ICASSP 2021accepted

Network anomaly detection is important for detecting and reacting to the presence of network attacks. In this paper, we propose a novel method to effectively leverage the features in detecting network anomalies, named FDEn, consisting of flow-based Feature Derivation (FD) and prior knowledge incorpo…

Cited by 0SourceScholar
2021

Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based Regularizer

ICASSP 2021accepted

B-mode ultrasound tongue imaging is widely used to visualize the tongue motion, due to its appearing properties. Extracting the tongue surface contour in the B-mode ultrasound image is still a challenge, while it is a prerequisite for further quantitative analysis. Recently, deep learning-based appr…

Cited by 0SourceScholar
2019

Denoising Convolutional Autoencoder Based B-mode Ultrasound Tongue Image Feature Extraction

ICASSP 2019accepted

B-mode ultrasound tongue imaging is widely used in the speech production field. However, efficient interpretation is in a great need for the tongue image sequences. Inspired by the recent success of unsupervised deep learning approach, we explore unsupervised convolutional network architecture for t…

Cited by 0SourceScholar
2019

Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional Lstm Neural Networks

ICASSP 2019accepted

A challenge in speech production research is to predict future tongue movements based on a short period of past tongue movements. This study tackles speaker-dependent tongue motion prediction problem in unlabeled ultrasound videos with convolutional long short-term memory (ConvLSTM) networks. The mo…

Cited by 0SourceScholar
2016

Contour-based 3D tongue motion visualization using ultrasound image sequences

ICASSP 2016accepted

This article describes a contour-based 3D tongue deformation visualization framework using B-mode ultrasound image sequences. A robust, automatic tracking algorithm characterizes tongue motion via a contour, which is then used to drive a generic 3D Finite Element Model (FEM). A novel contour-based 3…

Cited by 0SourceScholar