← Search

Xiaoqi Wang

13 accepted papers

2026

KSS-MoE: Knowledge Space Synergy Framework in Mixture of Experts for Continual Visual Instruction Tuning

AAAI 2026technical

Multimodal Large Language Models (MLLMs) employing the Mixture-of-Experts (MoE) structure exhibit encouraging results in visual language tasks. However, they struggle with catastrophic forgetting due to a lack of effective collaboration among experts and negative transfer across tasks. This happens

Cited by 0SourcePDFScholar
2025

AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?

ICCV 2025poster

Vision Language Models (VLMs) have exhibited remarkable generalization capabilities, yet their robustness in dynamic real-world scenarios remains largely unexplored. To systematically evaluate VLMs' robustness to real-world 3D variations, we propose AdvDreamer, the first framework capable of generat…

Cited by 0SourcePDFScholar
2025

MTGIB-UNet: A Multi-Task Graph Information Bottleneck and Uncertainty Weighted Network for ADMET Prediction

IJCAI 2025

Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) properties is crucial in drug development, as these properties directly impact a drug's efficacy and safety. However, existing multi-task learning models often face challenges related to noise interference a

Cited by 0SourcePDFScholar
2025

Multi-Scale Temporal Neural Network for Stock Trend Prediction Enhanced by Temporal Hyepredge Learning

IJCAI 2025

Existing research in Stock Trend Prediction (STP) focuses on temporal features extracted from a temporal sequence of stock data with a look-back window, which frequently leads to the omission of important periodic patterns, such as weekly and monthly variations in stock prices. Furthermore, these me

2025

ProSAM: Enhancing the Robustness of SAM-based Visual Reference Segmentation with Probabilistic Prompts

ICCV 2025poster

The recent advancements in large foundation models have driven the success of open-set image segmentation, a task focused on segmenting objects beyond predefined categories. Among various prompt types (such as points, boxes, texts, and visual references), visual reference segmentation stands out for…

Cited by 0SourcePDFScholar
2024

FedNE: Surrogate-Assisted Federated Neighbor Embedding for Dimensionality Reduction

NeurIPS 2024poster

Federated learning (FL) has rapidly evolved as a promising paradigm that enables collaborative model training across distributed participants without exchanging their local data. Despite its broad applications in fields such as computer vision, graph learning, and natural language processing, the de…

Cited by 0SourcePDFScholar
2024

GNNBoundary: Towards Explaining Graph Neural Networks through the Lens of Decision Boundaries

ICLR 2024poster

While Graph Neural Networks (GNNs) have achieved remarkable performance on various machine learning tasks on graph data, they also raised questions regarding their transparency and interpretability. Recently, there have been extensive research efforts to explain the decision-making process of GNNs.…

2024

USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation

CVPR 2024poster

The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating…

Cited by 5SourcePDFScholar
2023

GNNInterpreter: A Probabilistic Generative Model-Level Explanation for Graph Neural Networks

ICLR 2023poster

Recently, Graph Neural Networks (GNNs) have significantly advanced the performance of machine learning tasks on graphs. However, this technological breakthrough makes people wonder: how does a GNN make such decisions, and can we trust its prediction with high confidence? When it comes to some critic…

2023

Learning Hybrid Representations of Semantics and Distortion for Blind Image Quality Assessment

ICASSP 2023accepted

Recently, some studies have shown that semantic and distortion representations both benefit the evaluation of image quality. However, the images of existing synthetic distortion databases are annotated with subjective quality scores and distortion types, lacking labels with semantic objects. Therefo…

Cited by 0SourceScholar
2019

VTNFP: An Image-Based Virtual Try-On Network With Body and Clothing Feature Preservation

ICCV 2019poster

Image-based virtual try-on systems with the goal of transferring a desired clothing item onto the corresponding region of a person have made great strides recently, but challenges remain in generating realistic looking images that preserve both body and clothing details. Here we present a new virtua…

Cited by 203PDFScholar
2018

A Novel Learnable Dictionary Encoding Layer for End-to-End Language Identification

ICASSP 2018accepted

A novel learnable dictionary encoding layer is proposed in this paper for end-to-end language identification. It is inline with the conventional GMM i-vector approach both theoretically and practically. We imitate the mechanism of traditional GMM training and Supervector encoding procedure on the to…

Cited by 79SourceScholar
2018

Insights in-to-End Learning Scheme for Language Identification

ICASSP 2018accepted

A novel interpretable end-to-end learning scheme for language identification is proposed. It is in line with the classical GMM i-vector methods both theoretically and practically. In the end-to-end pipeline, a general encoding layer is employed on top of the frontend CNN, so that it can encode the v…

Cited by 20SourceScholar