← Search

Xuewei Li

19 accepted papers

2026

REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion

CVPR 2026

As an important and challenging problem in computer vision, Panoramic Semantic Segmentation (PASS) aims to provide complete scene perception based on an ultra-wide angle of view. Most PASS methods often focus on spherical geometry with RGB input or use the depth information in original or HHA format

Cited by 0SourceScholar
2025

A Novel Network for Short-Term Wind Speed Prediction: Mitigating Distribution Shift and Feature Loss

ICASSP 2025accepted

Accurate wind speed forecasting is essential for mitigating the challenges of wind power grid integration. However, existing wind speed prediction models overlook the distributional shift problem within wind speed series, and this time-varying distribution can significantly impact wind prediction ac…

Cited by 0SourceScholar
2025

Co-training with Progressive Distribution Alignment and Uncertainty-Interactive Relabeling for Semi-Supervised Domain Adaptive Semantic Segmentation

ICASSP 2025accepted

Self-training is a strong baseline for semi-supervised domain adaptive semantic segmentation. However, it inevitably introduces biased links between features and concepts in the prediction of certain "hard pixels", which may mislead the generalization of models. We consider these hard pixels to come…

Cited by 0SourceScholar
2025

DS-MHP: Improving Chain-of-Thought through Dynamic Subgraph-Guided Multi-Hop Path

EMNLP 2025

Large language models (LLMs) excel in natural language tasks, with Chain-of-Thought (CoT) prompting enhancing reasoning through step-by-step decomposition. However, CoT struggles in knowledge-intensive tasks with multiple entities and implicit multi-hop relations, failing to connect entities systema

2025

OCLNet: Obfuscation feature Contrastive Learning Network for Weakly Supervised Semantic Segmentation on Ultrasound Images

ICASSP 2025accepted

Deep learning-based semantic segmentation technology has become a critical tool in assisting doctors with automatic lesion segmentation in medical images. However, the high cost of acquiring large-scale, pixel-level annotations poses a significant challenge, limiting the scalability and application…

Cited by 0SourceScholar
2025

Popularity and Interest Signal Detection for Sequential Recommendation Denoising

ICASSP 2025accepted

Sequential recommender systems aim to learn user preferences through historical interaction sequences. User interactions are driven both by popular trends and personal interests, introducing two types of noise: popular choices triggered by conformist behavior and irrelevant terms that do not reflect…

Cited by 0SourceScholar
2024

Balanced And Discriminative Contrastive Learning For Class-Imbalanced Medical Images

ICASSP 2024accepted

The class imbalance problem, which is prevalent in medical image datasets, seriously affects the diagnostic effectiveness of deep learning-based network models. Recently, the method based on two-stage learning has produced promising results in solving class imbalance. In two-stage learning, the lear…

Cited by 0SourceScholar
2024

Debiasing Recommenders Through Personalized Popularity-Aware Margins

ICASSP 2024accepted

Recommender systems based on Matrix Factorization are widely used. However, they can easily suffer from the problem of overrecommendation of popular items, i.e., popularity bias. To mitigate popularity bias, current methods often uniformly model interactions' popularity bias degree considering user…

Cited by 0SourceScholar
2024

DualGCN-MIL: Whole Slide Image Classification Based on Double Relationship Graph Learning

ICASSP 2024accepted

The resolution of a whole slide image (WSI) is too large to process directly, but WSI can be segmented into patches and be classified through multiple instance learning (MIL). Some patches have either close distances or similar pathological morphology, indicating that there are at least two types of…

Cited by 0SourceScholar
2024

Multi-Level Augmentation Consistency Learning and Sample Selection for Semi-Supervised Domain Generalization

ICASSP 2024accepted

Semi-supervised domain generalization (SSDG) aims to build a domain-generalized model using partially labeled data from source domains. Mainstream SSDG methods follow the augmentation consistency in FixMatch. However, the extraction of domain-invariant features may be challenging due to the absence…

Cited by 0SourceScholar
2024

SphereDiffusion: Spherical Geometry-Aware Distortion Resilient Diffusion Model

AAAI 2024technical

Controllable spherical panoramic image generation holds substantial applicative potential across a variety of domains. However, it remains a challenging task due to the inherent spherical distortion and geometry characteristics, resulting in low-quality content generation. In this paper, we introduc…

Cited by 8SourcePDFScholar
2023

Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object Detection

ICCV 2023poster

Knowledge distillation (KD) has shown potential for learning compact models in dense object detection. However, the commonly used softmax-based distillation ignores the absolute classification scores for individual categories. Thus, the optimum of the distillation loss does not necessarily lead to t…

Cited by 31PDFcodeScholar
2023

LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image Generation

CVPR 2023poster

Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging…

2023

SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation

IJCAI 2023poster

As an important and challenging problem in computer vision, PAnoramic Semantic Segmentation (PASS) gives complete scene perception based on an ultra-wide angle of view. Usually, prevalent PASS methods with 2D panoramic image input focus on solving image distortions but lack consideration of the 3D p…

2023

Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory Inversion

ICASSP 2023accepted

Acoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly…

Cited by 0SourceScholar
2022

Residual-Guided Personalized Speech Synthesis based on Face Image

ICASSP 2022accepted

Previous works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to syn…

Cited by 0SourceScholar
2022

Transtl: Spatial-Temporal Localization Transformer for Multi-Label Video Classification

ICASSP 2022accepted

Multi-label video classification (MLVC) is a long-standing and challenging research problem in video signal analysis. Generally, there exist many complex action labels in real-world videos and these actions are with inherent dependencies at both spatial and temporal domains. Motivated by this observ…

Cited by 0SourceScholar
2022

W-ART: Action Relation Transformer for Weakly-Supervised Temporal Action Localization

ICASSP 2022accepted

Weakly-supervised temporal action localization (WTAL) is a long-standing and challenging research problem in video signal analysis. It is to localize the action segments in the video given only video-level labels. The key to this task is understanding how the diverse actions interact. In this paper,…

Cited by 0SourceScholar
2021

Self-Supervised Depth Estimation Via Implicit Cues from Videos

ICASSP 2021accepted

In self-supervised monocular depth estimation, the depth discontinuity and motion objects' artifacts are still challenging problems. Existing self-supervised methods usually utilize two views to train the depth estimation network and use one single view to make predictions. Compared with static view…

Cited by 0SourceScholar
Xuewei Li — accepted AI-conference papers · AIConfPaper