← Search

Xulei Yang

30 accepted papers

2026

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative Policy Optimization (GRPO) based approaches suffer from two

Cited by 0SourcePDFScholar
2026

GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection

CVPR 2026

In this paper, we study unsupervised multi-modal anomaly detection under challenging few-shot conditions, where only a few normal training samples are available for each class. To prevent the trivial reconstruction of anomalies, recent methods often rely on discrete prototypes to establish an inform

Cited by 0SourceScholar
2026

MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer

CVPR 2026

3D pose transfer aims to transfer the pose-style of a source mesh to a target character while preserving both the target's geometry and the source's pose characteristic. Existing methods are largely restricted to characters with similar structures and fail to generalize to category-free settings (e.

Cited by 0SourceScholar
2026

Next-Generation Metalens Vision System: Powered by AI and Applied to AI

AAAI 2026technical

Metalenses have been widely recognized as a key building block of next-generation optical systems, offering unprecedented advantages in compactness, lightweight design, and scalable manufacturing compared to traditional refractive optics. Despite this promise, practical use is limited by optical abe

Cited by 0SourcePDFScholar
2026

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

CVPR 2026

Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an intuitive interface, yet translating passenger open-ended instructions into control signals--without sacrificing interpretability and traceability--remai

Cited by 0SourceScholar
2026

PIRN: Prototypical-based Intra-modal Reconstruction with Normality Communication for Multi-modal Anomaly Detection.

ICLR 2026poster

Unsupervised Multimodal anomaly detection (MAD) — identifying defects by jointly analyzing RGB images and 3D data — is crucial for quality control in manufacturing. However, existing MAD methods struggle when only a few normal samples are available. Cross-modal alignment models fail to learn stable…

Cited by 0SourceScholar
2026

PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving

CVPR 2026

This paper presents the first study on Unsupervised Domain Adaptation (UDA) for multimodal 3D panoptic segmentation (mm-3DPS), aiming to improve generalization under domain shifts commonly encountered in real-world autonomous driving. A straightforward solution is to employ a pseudo-labeling strateg

Cited by 0SourceScholar
2026

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

ICLR 2026poster

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation.…

Cited by 0SourcecodeScholar
2026

Test-Time Optimization of 3D Point Cloud LLM via Manifold-Aware In-Context Guidance and Refinement

ICLR 2026poster

Multimidal Large Language Models (MLLMs) have demonstrated impressive capabilities in textual and 2D visual reasoning, yet their ability to understand and reason over 3D data remains limited. The issues become more challenging for understanding standalone 3D point cloud due to the high interclass co…

Cited by 0SourceScholar
2026

Towards Illumination-Aware Restoration of Metalens-Captured Images: A New Dataset and a Strong Baseline

AAAI 2026technical

Metalenses offer compelling advantages such as lightweight and ultra-thin design, making them promising alternatives to conventional lenses. However, their widespread adoption is hindered by image quality degradation caused by chromatic and angular aberrations. To mitigate this, restoration processe

Cited by 0SourcePDFScholar
2026

VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery

CVPR 2026

Human mesh recovery (HMR) from a single RGB image is inherently ambiguous, as multiple 3D poses can correspond to the same 2D observation. Recent diffusion-based methods tackle this by generating various hypotheses, but often sacrifice accuracy. They yield predictions that are either physically impl

Cited by 0SourceScholar
2025

Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation

ICASSP 2025accepted

Curvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for…

Cited by 0SourceScholar
2025

Evidential Learning-based Certainty Estimation for Robust Dense Feature Matching

ICLR 2025poster

Dense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether e…

Cited by 0SourcePDFScholar
2025

Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation

ICCV 2025poster

Out-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detection methods have shown progress for 2D image data, extending these to 3D environments involves unique obstacles. This pa…

2025

FIND: Few-Shot Anomaly Inspection with Normal-Only Multi-Modal Data

ICCV 2025poster

Multi-modal anomaly detection (MAD) improves industrial inspection by exploiting complementary 2D and 3D data. However, existing methods struggle in few-shot scenarios due to limited data and modality gaps. Current approaches either fuse multi-modal features or align cross-modal representations; how…

Cited by 0SourcePDFScholar
2025

Global-Aware Monocular Semantic Scene Completion with State Space Models

ICCV 2025poster

Monocular Semantic Scene Completion (MonoSSC) reconstructs and interprets 3D environments from a single image, enabling diverse real-world applications. However, existing methods are often constrained by the local receptive field of Convolutional Neural Networks (CNNs), making it challenging to hand…

Cited by 0SourcePDFScholar
2025

How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation

ICML 2025poster

LiDAR-based 3D panoptic segmentation often struggles with the inherent sparsity of data from LiDAR sensors, which makes it challenging to accurately recognize distant or small objects. Recently, a few studies have sought to overcome this challenge by integrating LiDAR inputs with camera images, leve…

2025

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

CVPR 2025poster

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model's ability to synthesize coherent and novel…

2025

On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning

ICLR 2025poster

Test-time adaptation (TTA) updates the model weights during the inference stage using testing data to enhance generalization. However, this practice exposes TTA to adversarial risks. Existing studies have shown that when TTA is updated with crafted adversarial test samples, also known as test-time p…

2025

Rectification-specific Supervision and Constrained Estimator for Online Stereo Rectification

CVPR 2025poster

Online stereo rectification is critical for autonomous vehicles and robots in dynamic environments, where factors such as vibration, temperature fluctuations, and mechanical stress can affect rectification accuracy and severely degrade downstream stereo depth estimation. Current dominant approaches…

Cited by 0SourcePDFScholar
2025

SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

CVPR 2025poster

3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome…

2025

Text-to-Image Rectified Flow as Plug-and-Play Priors

ICLR 2025poster

Large-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For instance, 2D diffusion models can serve as loss functions to optimize 3D implic…

2024

LPViT: Low-Power Semi-structured Pruning for Vision Transformers

ECCV 2024poster

"Vision transformers (ViTs) have emerged as a promising alternative to convolutional neural networks (CNNs) for various image analysis tasks, offering comparable or superior performance. However, one significant drawback of ViTs is their resource-intensive nature, leading to increased memory footpri…

2024

Learn to Optimize Denoising Scores: A Unified and Improved Diffusion Prior for 3D Generation

ECCV 2024poster

"In this paper, we propose a unified framework aimed at enhancing the diffusion priors for 3D generation tasks. Despite the critical importance of these tasks, existing methodologies often struggle to generate high-caliber results. We begin by examining the inherent limitations in previous diffusion…

2024

Learning Intra-view and Cross-view Geometric Knowledge for Stereo Matching

CVPR 2024poster

Geometric knowledge has been shown to be beneficial for the stereo matching task. However prior attempts to integrate geometric insights into stereo matching algorithms have largely focused on geometric knowledge from single images while crucial cross-view factors such as occlusion and matching uniq…

2024

Training Binary Neural Networks via Gaussian Variational Inference and Low-Rank Semidefinite Programming

NeurIPS 2024poster

Current methods for training Binarized Neural Networks (BNNs) heavily rely on the heuristic straight-through estimator (STE), which crucially enables the application of SGD-based optimizers to the combinatorial training problem. Although the STE heuristics and their variants have led to significant…

Cited by 0SourcePDFScholar
2023

SemiGNN-PPI: Self-Ensembling Multi-Graph Neural Network for Efficient and Generalizable Protein–Protein Interaction Prediction

IJCAI 2023poster

Protein-protein interactions (PPIs) are crucial in various biological processes and their study has significant implications for drug development and disease diagnosis. Existing deep learning methods suffer from significant performance degradation under complex real-world scenarios due to various fa…

Cited by 21SourcePDFScholar