← Search

Wei Xia

36 accepted papers

2026

Class-Guided Network with Rare-Class Amplification for Sea State Estimation Based on Ship Motion Data

ICRA 2026poster

Accurate, real-time Sea State Estimation (SSE) is crucial for the safety and operational efficiency of Autonomous Surface Vessels (ASVs). However, existing deep learning methods for this task commonly face three major challenges: the inherent class imbalance of marine environments, the ambiguous bou…

Cited by 0Scholar
2026

Evolutionary Generation of Multi-Agent Systems

ICML 2026poster

Large language model (LLM)–based multi-agent systems (MAS) show strong promise for complex reasoning, planning, and tool-augmented tasks, but designing effective MAS architectures remains labor-intensive, brittle, and hard to generalize. Existing automatic MAS generation methods either rely on code …

Cited by 0SourceScholar
2026

Gated KalmaNet: A Fading Memory Layer through Test-time Ridge Regression

CVPR 2026

As efficient alternatives to softmax Attention, linear state space models (SSMs) achieve constant memory and linear compute, but maintain only a lossy, fading summary of the past, often leading to inferior performance in recall oriented settings. We propose Gated KalmaNet (GKA), a layer that reduces

Cited by 0SourcecodeScholar
2026

Learning When to Attend: Conditional Memory Access for Long-Context LLMs

ICML 2026poster

Language models struggle to generalize beyond the context lengths seen during pretraining, limiting performance on long-horizon reasoning and retrieval. Continued pretraining on long-context data can mitigate this limitation, but it is prohibitively expensive due to the quadratic scaling of Attentio…

Cited by 0SourceScholar
2026

Re-FORC: Adaptive Reward Prediction for Efficient Chain-of-Thought Reasoning

ICML 2026poster

We propose Re-FORC, an adaptive reward prediction method that, given a context, enables prediction of the expected future rewards as a function of the number of future thinking tokens. Re-FORC trains a lightweight adapter on reasoning models, demonstrating improved prediction with longer reasoning a…

Cited by 0SourceScholar
2026

Reinforcement-aware Knowledge Distillation for LLM Reasoning

ICML 2026poster

Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs), but the high inference cost of such models motivates distillation into smaller students. Most existing knowledge distillation (KD) methods are designed for super…

Cited by 0SourceScholar
2026

Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

CVPR 2026

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for multimodal generation systems. While existing text-based manip

Cited by 0SourcecodeScholar
2025

ANASETC: Automatic Neural Architecture Search for Encrypted Traffic Classification

ICASSP 2025accepted

The widespread adoption of encrypted network protocols has made traffic encryption ubiquitous, creating substantial challenges for network management and security. This paper introduces a novel encrypted traffic classification system, ANASETC, which combines traffic burst features with Neural Archit…

Cited by 0SourceScholar
2025

CoIR: A Comprehensive Benchmark for Code Information Retrieval Models

ACL 2025long

Despite the substantial success of Information Retrieval (IR) in various NLP tasks, most IR systems predominantly handle queries and corpora in natural language, neglecting the domain of code retrieval. Code retrieval is critically important yet remains under-explored, with existing methods and benc…

2025

RethinkMCTS: Refining Erroneous Thoughts in Monte Carlo Tree Search for Code Generation

EMNLP 2025

Tree search methods have demonstrated impressive performance in code generation. Previous methods combine tree search with reflection that summarizes past mistakes to achieve iterative improvement. However, these methods face significant challenges. First, they search directly within the code langua

2024

A Full-duplex Speech Dialogue Scheme Based On Large Language Model

NeurIPS 2024poster

We present a generative dialogue system capable of operating in a full-duplex manner, allowing for seamless interaction. It is based on a large language model (LLM) carefully aligned to be aware of a perception module, a motor function module, and the concept of a simple finite state machine (called…

Cited by 15SourcePDFScholar
2024

Large Scale Self-Supervised Pretraining for Active Speaker Detection

ICASSP 2024accepted

In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show…

Cited by 0SourceScholar
2023

Centerless Multi-View K-means Based on the Adjacency Matrix

AAAI 2023technical

Although K-Means clustering has been widely studied due to its simplicity, these methods still have the following fatal drawbacks. Firstly, they need to initialize the cluster centers, which causes unstable clustering performance. Secondly, they have poor performance on non-Gaussian datasets. Inspir…

2023

Orthogonal Non-negative Tensor Factorization based Multi-view Clustering

NeurIPS 2023poster

Multi-view clustering (MVC) based on non-negative matrix factorization (NMF) and its variants have attracted much attention due to their advantages in clustering interpretability. However, existing NMF-based multi-view clustering methods perform NMF on each view respectively and ignore the impact of…

Cited by 32SourcePDFScholar
2023

Set-to-Sequence Ranking-Based Concept-Aware Learning Path Recommendation

AAAI 2023technical

With the development of the online education system, personalized education recommendation has played an essential role. In this paper, we focus on developing path recommendation systems that aim to generating and recommending an entire learning path to the given user in each session. Noticing that…

Cited by 11SourcePDFScholar
2022

Multi-Dimensional, Nuanced and Subjective - Measuring the Perception of Facial Expressions

CVPR 2022poster

Humans can perceive multiple expressions, each one with varying intensity, in the picture of a face. We propose a methodology for collecting and modeling multidimensional modulated expression annotations from human annotators. Our data reveals that the perception of some expressions can be quite dif…

Cited by 9PDFScholar
2022

Stochastic Backpropagation: A Memory Efficient Strategy for Training Video Models

CVPR 2022oral

We propose a memory efficient method, named Stochastic Backpropagation (SBP), for training deep neural networks on videos. It is based on the finding that gradients from incomplete execution for backpropagation can still effectively train the models with minimal accuracy loss, which attributes to th…

Cited by 22PDFcodeScholar
2022

Towards Regression-Free Neural Networks for Diverse Compute Platforms

ECCV 2022poster

"With the shift towards on-device deep learning, ensuring a consistent behavior of an AI service across diverse compute platforms becomes tremendously important. Our work tackles the emergent problem of reducing predictive in-consistencies arising as negative flips: test samples that are correctly p…

Cited by 4SourcePDFScholar
2022

Turn-to-Diarize: Online Speaker Diarization Constrained by Transformer Transducer Speaker Turn Detection

ICASSP 2022accepted

In this paper, we present a novel speaker diarization system for streaming on-device applications. In this system, we use a transformer transducer to detect the speaker turns, represent each speaker turn by a speaker embedding, then cluster these embeddings with constraints from the detected speaker…

Cited by 61SourceScholar
2022

Unsupervised and Semi-Supervised Bias Benchmarking in Face Recognition

ECCV 2022poster

"We introduce Semi-supervised Performance Evaluation for Face Recognition (SPE-FR). SPE-FR is a statistical method for evaluating the performance and algorithmic bias of face verification systems when identity labels are unavailable or incomplete. The method is based on parametric Bayesian modeling…

Cited by 14SourcePDFScholar
2021

DEAAN: Disentangled Embedding and Adversarial Adaptation Network for Robust Speaker Representation Learning

ICASSP 2021accepted

Despite speaker verification has achieved significant performance improvement with the development of deep neural networks, do-main mismatch is still a challenging problem in this field. In this study, we propose a novel framework to disentangle speaker-related and domain-specific features and apply…

Cited by 0SourceScholar
2021

Learning Hierarchical Graph Neural Networks for Image Clustering

ICCV 2021poster

We propose a hierarchical graph neural network (GNN) model that learns how to cluster a set of images into an unknown number of identities using a training set of images annotated with labels belonging to a disjoint set of identities. Our hierarchical GNN uses a novel approach to merge connected com…

Cited by 53PDFcodeScholar
2021

Long Short-Term Transformer for Online Action Detection

NeurIPS 2021spotlight

We present Long Short-term TRansformer (LSTR), a temporal modeling algorithm for online action detection, which employs a long- and short-term memory mechanism to model prolonged sequence data. It consists of an LSTR encoder that dynamically leverages coarse-scale historical information from an exte…

2021

Positive-Congruent Training: Towards Regression-Free Model Updates

CVPR 2021poster

Reducing inconsistencies in the behavior of different versions of an AI system can be as important in practice as reducing its overall error. In image classification, sample-wise inconsistencies appear as "negative flips": A new model incorrectly predicts the output for a test sample that was correc…

Cited by 63PDFScholar
2021

Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning

ICASSP 2021accepted

In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a momentum contrastive (MoCo) learning framework, where the MoCo speaker embedding system utilizes a queue to maintain a large s…

Cited by 0SourceScholar
2020

Towards causal benchmarking of bias in face analysis algorithms

ECCV 2020poster

Measuring algorithmic bias is crucial both to assess algorithmic fairness, and to guide the improvement of algorithms. Current bias measurement methods in computer vision are based on observational datasets, and conflate algorithmic bias with dataset bias. To address this problem we develop an exper…

Cited by 103SourcePDFScholar
2019

Cross-lingual Text-independent Speaker Verification Using Unsupervised Adversarial Discriminative Domain Adaptation

ICASSP 2019accepted

Speaker verification systems often degrade significantly when there is a language mismatch between training and testing data. Being able to improve cross-lingual speaker verification system using unlabeled data can greatly increase the robustness of the system and reduce human labeling costs. In thi…

Cited by 69SourceScholar
2019

UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation

ICASSP 2019accepted

In this study, we present systems submitted by the Center for Robust Speech Systems (CRSS) from UTDallas to NIST SRE 2018 (SRE18). Three alternative front-end speaker embedding frameworks are investigated, that includes: (i) i-vector, (ii) x-vector, (iii) and a modified triplet speaker embedding sys…

Cited by 21SourceScholar
2017

Average SCR loss analysis for polarimetric STAP with Kronecker structured covariance matrix

ICASSP 2017accepted

The paper presents the average signal-to-clutter loss (SCRL) analysis for polarimetric space-time adaptive processing by exploiting the Kronecker structure of the clutter covariance matrix (CM). An expression for the average SCRL as a function of the mean square error of the corresponding CM estimat…

Cited by 0SourceScholar