← Search

Xun Xu

27 accepted papers

2026

AD-FM: Multimodal LLMs for Anomaly Detection via Multi-Stage Reasoning and Fine-Grained Reward Optimization

AAAI 2026technical

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities across diverse domains, their application to specialized anomaly detection (AD) remains constrained by domain adaptation challenges. Existing Group Relative Policy Optimization (GRPO) based approaches suffer from two

Cited by 0SourcePDFScholar
2026

Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos

ICRA 2026poster

Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand–object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT)…

Cited by 0codeScholar
2026

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

AAAI 2026technical

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further

Cited by 0SourcePDFScholar
2026

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

ICLR 2026poster

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as generating coordinates as text for detection, which limits performance and prevents dense prediction tasks like segmentation.…

Cited by 0SourcecodeScholar
2026

Test-Time Optimization of 3D Point Cloud LLM via Manifold-Aware In-Context Guidance and Refinement

ICLR 2026poster

Multimidal Large Language Models (MLLMs) have demonstrated impressive capabilities in textual and 2D visual reasoning, yet their ability to understand and reason over 3D data remains limited. The issues become more challenging for understanding standalone 3D point cloud due to the high interclass co…

Cited by 0SourceScholar
2026

Zero-Shot Image Denoising via Hybrid Prior-Guided Pseudo Sample Generation

CVPR 2026

Zero-shot image denoising has gained prominence in recent years, as it inherently relies on the intrinsic priors of images rather than learning from external data. Nevertheless, most existing methods either fail to fully exploit global priors, or do not properly preserve the fine-grained details gov

Cited by 0SourceScholar
2025

Distribution Alignment Informed Thresholding for Semi-Supervised Curvilinear Structure Segmentation

ICASSP 2025accepted

Curvilinear structure segmentation using deep neural networks is often limited by the high cost of annotation. Semi-supervised learning (SSL) helps mitigate this dependency on extensive annotated data. State-of-the-art SSL approaches generate pseudo-labels for unlabeled data, which are then used for…

Cited by 0SourceScholar
2025

Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model

ICLR 2025poster

Vision-language models (VLMs) have revolutionized machine learning by leveraging large pre-trained models to tackle various downstream tasks. Although label, training, and data efficiency have improved, many state-of-the-art VLMs still require task-specific hyperparameter tuning and fail to fully ex…

2025

Evidential Learning-based Certainty Estimation for Robust Dense Feature Matching

ICLR 2025poster

Dense feature matching methods aim to estimate a dense correspondence field between images. Inaccurate correspondence can occur due to the presence of unmatchable region, necessitating the need for certainty measurement. This is typically addressed by training a binary classifier to decide whether e…

Cited by 0SourcePDFScholar
2025

Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation

ICCV 2025poster

Out-of-distribution (OOD) detection in 3D point cloud data remains a challenge, particularly in applications where safe and robust perception is critical. While existing OOD detection methods have shown progress for 2D image data, extending these to 3D environments involves unique obstacles. This pa…

2025

GS-EVT: Cross-Modal Event Camera Tracking Based on Gaussian Splatting

ICRA 2025

Reliable self-localization is a foundational skill for many intelligent mobile platforms. This paper explores the use of event cameras for motion tracking thereby providing a solution with inherent robustness under difficult dynamics and illumination. In order to circumvent the challenge of event ca

Cited by 3SourceScholar
2025

Global-Aware Monocular Semantic Scene Completion with State Space Models

ICCV 2025poster

Monocular Semantic Scene Completion (MonoSSC) reconstructs and interprets 3D environments from a single image, enabling diverse real-world applications. However, existing methods are often constrained by the local receptive field of Convolutional Neural Networks (CNNs), making it challenging to hand…

Cited by 0SourcePDFScholar
2025

On the Adversarial Risk of Test Time Adaptation: An Investigation into Realistic Test-Time Data Poisoning

ICLR 2025poster

Test-time adaptation (TTA) updates the model weights during the inference stage using testing data to enhance generalization. However, this practice exposes TTA to adversarial risks. Existing studies have shown that when TTA is updated with crafted adversarial test samples, also known as test-time p…

2024

DuCAS: a knowledge-enhanced dual-hand compositional action segmentation method for human-robot collaborative assembly

IROS 2024poster

Recognising and tracking human actions from videos is crucial for human-robot collaborative assembly (HRCA). However, traditional action segmentation methods suffer from limited scene adaptability, partly because they conceptualise actions as unified verb-object entities with complete semantics. To…

Cited by 1SourcecodeScholar
2024

Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised Adaptation

CVPR 2024poster

The success of large language models has inspired the computer vision community to explore image segmentation foundation model that is able to zero/few-shot generalize through prompt engineering. Segment-Anything (SAM) among others is the state-of-the-art image segmentation foundation model demonstr…

2024

Towards Real-World Test-Time Adaptation: Tri-net Self-Training with Balanced Normalization

AAAI 2024technical

Test-Time Adaptation aims to adapt source domain model to testing data at inference stage with success demonstrated in adapting to unseen corruptions. However, these attempts may fail under more challenging real-world scenarios. Existing works mainly consider real-world test-time adaptation under no…

2023

On the Robustness of Open-World Test-Time Training: Self-Training with Dynamic Prototype Expansion

ICCV 2023oral

Generalizing deep learning models to unknown target domain distribution with low latency has motivated research into test-time training/adaptation (TTT/TTA). Existing approaches often focus on improving test-time training performance under well-curated target domain data. As figured out in this work…

Cited by 23PDFcodeScholar
2022

Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored Clustering

NeurIPS 2022accept

Deploying models on target domain data subject to distribution shift requires adaptation. Test-time training (TTT) emerges as a solution to this adaptation under a realistic scenario where access to full source domain data is not available and instant inference on target domain is required. Despite…

2021

3D AffordanceNet: A Benchmark for Visual Object Affordance Understanding

CVPR 2021poster

The ability to understand the ways to interact with objects from visual cues, a.k.a. visual affordance, is essential to vision-guided robotic research. This involves categorizing, segmenting and reasoning of visual affordance. Relevant studies in 2D and 2.5D image domains have been made previously,…

Cited by 136PDFcodeScholar
2021

ARMOURED: Adversarially Robust MOdels using Unlabeled data by REgularizing Diversity

ICLR 2021poster

Adversarial attacks pose a major challenge for modern deep neural networks. Recent advancements show that adversarially robust generalization requires a large amount of labeled data for training. If annotation becomes a burden, can unlabeled data help bridge the gap? In this paper, we propose ARMOUR…

Cited by 3SourcePDFScholar
2021

Revisiting Superpixels for Active Learning in Semantic Segmentation With Realistic Annotation Costs

CVPR 2021poster

State-of-the-art methods for semantic segmentation are based on deep neural networks that are known to be data-hungry. Region-based active learning has shown to be a promising method for reducing data annotation costs. A key design choice for region-based AL is whether to use regularly-shaped region…

Cited by 75PDFScholar
2017

Hawkeye: Open source framework for field surveillance

IROS 2017poster

This paper introduces a generic framework for field surveillance using consumer rotorcrafts and ground vehicles. Building such an autonomous system comes with two key challenges in persistent perception and obstacle avoidance. We begin with explaining two core algorithms to solve the challenges: an…

Cited by 8SourceScholar