← Search

Shuai Xiao

19 accepted papers

2026

Automatic Translational Correction of Multi-View Coronary Angiography Based on Auto-Annotation Data Generation

AAAI 2026technical

Multi-view automatic translational correction (ATC) in coronary angiography (CAG) is critical for intraoperative automatic diagnosis, in which deep learning playing a key role. However, heartbeat-induced soft matching errors and costly annotations make it difficult to build high-quality, large-scale

Cited by 2SourcePDFScholar
2026

Geometrically Constrained Stenosis Editing in Coronary Angiography via Entropic Optimal Transport

ICML 2026poster

The scarcity of high-quality imaging data for coronary angiography (CAG) stenosis limits the clinical translation of automated stenosis detection. Synthetic stenosis data provides a practical avenue to augment training sets, improving data quality, diversity, and distributional coverage, and enhanci…

Cited by 0SourceScholar
2026

MGT-Prism: Enhancing Domain Generalization for Machine-Generated Text Detection via Spectral Alignment

AAAI 2026technical

Large Language Models have shown growing ability to generate fluent and coherent texts that are highly similar to the writing style of humans. Current detectors for Machine-Generated Text (MGT) perform well when they are trained and tested in the same domain but generalize poorly to unseen domains,

Cited by 0SourcePDFScholar
2026

scChord: A Probabilistic Manifold Rectification Framework for RNA-to-Protein Translation

ICML 2026poster

Measuring single-cell protein abundance is essential for resolving biological mechanisms and disease progression with high resolution. However, due to the high costs and antibody throughput limitations of current proteomics, inferring protein levels from readily available RNA data has become a criti…

Cited by 0SourceScholar
2025

Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-training

CVPR 2025poster

In rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web d…

2025

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

ICML 2025poster

Large language models (LLMs) have demonstrated strong capabilities in handling long-context tasks, but processing such long contexts remains challenging due to the substantial memory requirements and inference latency. In this work, we discover that certain attention heads exhibit sequential consist…

Cited by 0SourcePDFScholar
2025

Contrast-Unity for Partially-Supervised Temporal Sentence Grounding

ICASSP 2025accepted

Temporal sentence grounding aims to detect event timestamps described by the natural language query from given untrimmed videos. The existing fully-supervised setting achieves great results but requires expensive annotation costs; while the weakly-supervised setting adopts cheap labels but performs…

Cited by 0SourceScholar
2025

FOLDER: Accelerating Multi-Modal Large Language Models with Enhanced Performance

ICCV 2025poster

Recently, Multi-modal Large Language Models (MLLMs) have shown remarkable effectiveness for multi-modal tasks due to their ability of cross-modal understanding. However, processing long sequences of visual tokens extracted from visual backbones poses challenges for deployment in real-time applicatio…

2025

Gesture Identification and Object Temperature Detection of a Robotic Hand Using a Wireless Flexible Sensing Feedback Control System

RA-L 2025

The ability of robotic hands to sense their environment and provide feedback is becoming increasingly vital for advanced robotic systems. Real-time tactile interaction is crucial for ensuring safe and effective human-machine collaboration. However, conventional sensors face significant challenges in

Cited by 0SourceScholar
2025

Inner Information Analysis Algorithm for Deep Neural Network based on Community

ICLR 2025poster

Deep learning has achieved advancements across a variety of forefront fields. However, its inherent 'black box' characteristic poses challenges to the comprehension and trustworthiness of the decision-making processes within neural networks. To mitigate these challenges, we introduce InnerSightNet,…

Cited by 0SourcePDFScholar
2025

TimePro: Efficient Multivariate Long-term Time Series Forecasting with Variable- and Time-Aware Hyper-state

ICML 2025poster

In long-term time series forecasting, different variables often influence the target variable over distinct time intervals, a challenge known as the multi-delay issue. Traditional models typically process all variables or time points uniformly, which limits their ability to capture complex variable…

2024

Generative Model Perception Rectification Algorithm for Trade-Off between Diversity and Quality

AAAI 2024technical

How to balance the diversity and quality of results from generative models through perception rectification poses a significant challenge. Abnormal perception in generative models is typically caused by two factors: inadequate model structure and imbalanced data distribution. In response to this iss…

Cited by 3SourcePDFScholar
2024

Pseudo Label Refinery for Unsupervised Domain Adaptation on Cross-dataset 3D Object Detection

CVPR 2024poster

Recent self-training techniques have shown notable improvements in unsupervised domain adaptation for 3D object detection (3D UDA). These techniques typically select pseudo labels i.e. 3D boxes to supervise models for the target domain. However this selection process inevitably introduces unreliable…

2024

Wear-Any-Way: Manipulable Virtual Try-on via Sparse Correspondence Alignment

ECCV 2024poster

"This paper introduces a novel framework for virtual try-on, termed . Different from previous methods, is a customizable solution. Besides generating high-fidelity results, our method supports users to precisely manipulate the wearing style. To achieve this goal, we first construct a strong pipeline…

2022

Tile Networks: Learning Optimal Geometric Layout for Whole-page Recommendation

AISTATS 2022poster

Finding optimal configurations in a geometric space is a key challenge in many technological disciplines. Current approaches either rely heavily on human domain expertise and are difficult to scale. In this paper we show it is possible to solve configuration optimization problems for whole-page reco…

2021

Efficient Multi-Stage Video Denoising With Recurrent Spatio-Temporal Fusion

CVPR 2021poster

In recent years, denoising methods based on deep learning have achieved unparalleled performance at the cost of large computational complexity. In this work, we propose an Efficient Multi-stage Video Denoising algorithm, called EMVD, to drastically reduce the complexity while maintaining or even imp…

Cited by 73PDFScholar
2018

Learning Temporal Point Processes via Reinforcement Learning

NeurIPS 2018spotlight

Social goods, such as healthcare, smart city, and information networks, often produce ordered event data in continuous time. The generative processes of these event data can be very complex, requiring flexible models to capture their dynamics. Temporal point processes offer an elegant framework for…

2017

Wasserstein Learning of Deep Generative Point Process Models

NeurIPS 2017poster

Point processes are becoming very popular in modeling asynchronous sequential data due to their sound mathematical foundation and strength in modeling a variety of real-world phenomena. Currently, they are often characterized via intensity function which limits model's expressiveness due to unrealis…