← Search

Xiangyuan Lan

24 accepted papers

2026

Bolster Hallucination Detection via Prompt-Guided Data Augmentation

AAAI 2026technical

Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fabricated information, a phenomenon known as hallucinations. Consequently, hallucination detection has become critical to

Cited by 0SourcePDFScholar
2026

CATAL: Causally Disentangled Task Representation Learning for Offline Meta-Reinforcement Learning

AAAI 2026technical

Context-based Offline Meta Reinforcement Learning (COMRL) has shown promising results in improving the cross-task generalization ability of meta-policies. However, current methods often lead to entangled task representations, in which each latent dimension is influenced by multiple causal factors th

Cited by 0SourcePDFScholar
2026

Data Scaling Laws for Imitation Learning-Based End-To-End Autonomous Driving

ICRA 2026poster

The end-to-end autonomous driving paradigm has recently attracted lots of attention due to its scalability. However, existing methods are constrained by the limited scale of real-world data, which hinders a comprehensive exploration of the scaling laws associated with end-to-end autonomous driving. …

2026

X-SAM: From Segment Anything to Any Segmentation

AAAI 2026technical

Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exh

Cited by 0SourcePDFScholar
2025

AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment

CVPR 2025poster

Cross-modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-…

Cited by 1SourcePDFScholar
2025

DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval

AAAI 2025technical

Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges sti…

2025

EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment

ICLR 2025poster

Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanc…

2025

HYMAN: Hybrid Memory and Attention Network for Unsupervised Anomaly Detection

ICASSP 2025accepted

Detecting anomalies in unsupervised multivariate time series is challenging due to the intricate temporal patterns present in both local short-term and global long-term dependencies. Long short-term memory has achieved impressive results in this domain, yet it is gradually being supplemented by Tran…

Cited by 0SourceScholar
2025

Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL

NeurIPS 2025poster

Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem…

Cited by 0SourcecodeScholar
2025

Open-Det: An Efficient Learning Framework for Open-Ended Detection

ICML 2025poster

Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training,…

2025

Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era

NeurIPS 2025poster

Visual place recognition (VPR) is typically regarded as a specific image retrieval task, whose core lies in representing images as global descriptors. Over the past decade, dominant VPR methods (e.g., NetVLAD) have followed a paradigm that first extracts the patch features/tokens of the input image…

Cited by 0SourcecodeScholar
2025

Transferable Adversarial Face Attack with Text Controlled Attribute

AAAI 2025technical

Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically meaningful perturbations. Current unrestricted adversarial impersonation attacks exhibit limited control over adversaria…

2024

CricaVPR: Cross-image Correlation-aware Representation Learning for Visual Place Recognition

CVPR 2024poster

Over the past decade most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and neglect the cross-image variations (e.g. viewpoint and illumina…

2024

Deep Homography Estimation for Visual Place Recognition

AAAI 2024technical

Visual place recognition (VPR) is a fundamental task for many applications such as robot localization and augmented reality. Recently, the hierarchical VPR methods have received considerable attention due to the trade-off between accuracy and efficiency. They usually first use global features to ret…

2024

High-Resolution Image Harmonization with Adaptive-Interval Color Transformation

NeurIPS 2024poster

Existing high-resolution image harmonization methods typically rely on global color adjustments or the upsampling of parameter maps. However, these methods ignore local variations, leading to inharmonious appearances. To address this problem, we propose an Adaptive-Interval Color Transformation meth…

2024

MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection

IJCAI 2024poster

Popular transformer-based detectors detect objects in a one-to-one manner, where both the bounding box and category of each object are predicted only by the single query, leading to the box-sensitive category predictions. Additionally, the initialization of positional queries solely based on the pre…

2024

Revisiting Context Aggregation for Image Matting

ICML 2024poster

Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle…

2024

SuperVLAD: Compact and Robust Image Descriptors for Visual Place Recognition

NeurIPS 2024poster

Visual place recognition (VPR) is an essential task for multiple applications such as augmented reality and robot localization. Over the past decade, mainstream methods in the VPR area have been to use feature representation based on global aggregation, as exemplified by NetVLAD. These features are…

2024

Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition

ICLR 2024poster

Recent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts have been made to exploit pre-trained foundation models in visual place recogniti…

2023

Strip-MLP: Efficient Token Interaction for Vision MLP

ICCV 2023poster

Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the m…

Cited by 13PDFcodeScholar
2019

Multi-Adversarial Discriminative Deep Domain Generalization for Face Presentation Attack Detection

CVPR 2019poster

Face presentation attacks have become an increasingly critical issue in the face recognition community. Many face anti-spoofing methods have been proposed, but they cannot generalize well on "unseen" attacks. This work focuses on improving the generalization ability of face anti-spoofing methods fro…

Cited by 428PDFcodeScholar
2018

Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection

ECCV 2018poster

3D mask face presentation attack, as a new challenge in face recognition, has been attracting increasing attention. Recently, remote Photoplethysmography (rPPG) is employed as an intrinsic liveness cue which is independent of the mask appearance. Although existing rPPG-based methods achieve promisin…

Cited by 153SourcePDFScholar
2018

Robust Anchor Embedding for Unsupervised Video Person Re-Identification in the Wild

ECCV 2018poster

This paper addresses the scalability and robustness issues of estimating labels from imbalanced unlabeled data for unsupervised video-based person re-identification (re-ID). To achieve it, we propose a novel Robust AnChor Embedding (RACE) framework via deep feature representation learning for large-…

Cited by 127SourcePDFScholar