← Search

Min Xu

48 accepted papers

2026

ALIGNING WHAT YOU SEPARATE: DENOISED PATCH MIXING FOR SOURCE-FREE DOMAIN ADAPTATION IN MEDICAL IMAGE SEGMENTATION

ICASSP 2026poster

Source-Free Domain Adaptation (SFDA) is emerging as a compelling solution for medical image segmentation under privacy constraints, yet current approaches often ignore sample difficulty and struggle with noisy supervision under domain shift. We present a new SFDA framework that leverages Hard Sample…

Cited by 0SourcePDFScholar
2026

AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation

ICML 2026poster

LLM agents are rapidly becoming the practical interface for task automation, yet the ecosystem lacks a principled way to \emph{choose} among an exploding space of deployable configurations. Existing LLM leaderboards and tool/agent benchmarks evaluate components in isolation and remain fragmented acr…

Cited by 0SourceScholar
2026

DOMAIN-INVARIANT MIXED-DOMAIN SEMI-SUPERVISED MEDICAL IMAGE SEGMENTATION WITH CLUSTERED MAXIMUM MEAN DISCREPANCY ALIGNMENT

ICASSP 2026poster

Deep learning has shown remarkable progress in medical image semantic segmentation, yet its success heavily depends on large-scale expert annotations and consistent data distributions. In practice, annotations are scarce, and images are collected from multiple scanners or centers, leading to mixed-d…

Cited by 0SourcePDFScholar
2026

DiLO: Disentangled Latent Optimization for Learning Shape and Deformation in Grouped Deforming 3D Objects

AAAI 2026technical

In this work, we propose a disentangled latent optimization-based method for parameterizing grouped deforming 3D objects into shape and deformation factors in an unsupervised manner. Our approach involves the joint optimization of a generator network along with the shape and deformation factors, sup

Cited by 0SourcePDFScholar
2026

Doubly Robust Distributionally Robust Offline Contextual Pricing

ICML 2026poster

Offline contextual pricing often relies on logged observational data, but faces challenges from distributional shifts between training and deployment environments. Distributionally robust optimization (DRO) provides a principled approach to off-policy evaluation and learning (OPE/L). However, existi…

Cited by 0SourceScholar
2026

Focal-General Diffusion Model with Semantic Consistent Guidance for Sign Language Production

CVPR 2026

Sign Language Production (SLP) aims to translate spoken language into sign sequences, where the main challenge lies in generating coherent and natural poses from discrete glosses (G2P). Existing G2P methods typically treat each pose as an indivisible unit, limiting their ability to capture fine-grai

Cited by 0SourcecodeScholar
2026

MedLIME: A Distribution-Aligned and Evidence-Supported Framework for Medical Saliency Explanations

CVPR 2026

Saliency-based explainability methods are widely used to interpret deep learning models in medical imaging, yet many existing approaches rely on white box access of models, which is not always possible due to privacy concerns. In this work, we introduce **MedLIME**, a novel, model-agnostic explanati

Cited by 0SourceScholar
2026

MicroFM: Physics-guided Flow Matching for Isotropic Microscopy Reconstruction

CVPR 2026

Isotropic microscopy reconstruction remains challenging because the anisotropic point spread function in optical systems yields much poorer axial resolution and hampers accurate 3D analysis. Hardware strategies can approach isotropy, yet they are complex, costly, susceptible to sidelobes, and introd

Cited by 0SourceScholar
2026

Multi-Object System Identification from Videos

ICLR 2026poster

We introduce the challenging problem of multi-object system identification from videos, for which prior methods are ill-suited due to their focus on single-object scenes or discrete material classification with a fixed set of material prototypes. To address this, we propose MOSIV, a new framework th…

Cited by 0SourceScholar
2026

Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model

CVPR 2026

We introduce an unsupervised approach for segmenting multiscale subcellular objects in 3D volumetric cryo-electron tomography (cryo-ET) images. To this end, we address key challenges such as lack of annotated data, large data volumes, high heterogeneity of subcellular shapes and sizes, and high inte

Cited by 0SourceScholar
2025

Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

ACL 2025long

We introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents. Agentic Reasoning dynamically leverages web search, code execution, and structured memory to address complex problems requiring deep research. A key innovation in…

Cited by 0SourcePDFScholar
2025

Answering Narrative-Driven Recommendation Queries via a Retrieve–Rank Paradigm and the OCG-Agent

EMNLP 2025

Narrative-driven recommendation queries are common in question-answering platforms, AI search engines, social forums, and some domain-specific vertical applications. Users typically submit free-form text requests for recommendations, e.g., “Any mind-bending thrillers like Shutter Island you’d recomm

2025

BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram Alignment

CVPR 2025poster

Subtomogram alignment is a critical task in cryo-electron tomography (cryo-ET) analysis, essential for achieving high-resolution reconstructions of macromolecular complexes. However, learning effective positional representations remains challenging due to limited labels and high noise levels inheren…

Cited by 0SourcePDFScholar
2025

DiffCAM: Data-Driven Saliency Maps by Capturing Feature Differences

CVPR 2025highlight

In recent years, the interpretability of Deep Neural Networks (DNNs) has garnered significant attention, particularly due to their widespread deployment in critical domains like healthcare, finance, and autonomous systems. To address the challenge of understanding how DNNs make decisions, Explainabl…

Cited by 0SourcePDFScholar
2025

Hierarchical Spatial-Temporal Enhancement Network For Continuous Sign Language Recognition

ICASSP 2025accepted

In continuous sign language recognition (CSLR), 2D-CNN-based extractors are often insufficiently trained for spatial capture and struggle with temporal modeling. This leads to incomplete spatial discrimination, hindering the understanding actions across frames. To address these limitations, we propo…

Cited by 0SourceScholar
2025

Improved Calibration for Panoramic Annular Lens Systems with Angular Modulation

IROS 2025

This paper addresses the challenges of calibrating Panoramic Annular Lens (PAL) systems, which exhibit unique projection characteristics due to their imaging relationship designed to compress blind zones. Traditional camera calibration methods often fail to accurately capture these properties. To re

Cited by 0SourcecodeScholar
2025

Improving Continuous Sign Language Recognition via Cross-Frame Interactions in Expanded Contextual Spaces

ICASSP 2025accepted

Current continuous sign language recognition (CSLR) methods typically rely on single or adjacent frames for calculations, which can overlook broader contextual information and result in lower accuracy. To address this issue, we introduce CVSign, which constructs an extended contextual space frame by…

Cited by 0SourceScholar
2025

Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation

ACL 2025long

We introduce MedGraphRAG, a novel graph-based Retrieval-Augmented Generation (RAG) framework designed to enhance LLMs in generating evidence-based medical responses, improving safety and reliability with private medical data. We introduce Triple Graph Construction and U-Retrieval to enhance GraphRAG…

2025

OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language Recognition

AAAI 2025technical

The primary challenge in continuous sign language recognition (CSLR) mainly stems from the presence of multi-orientational and long-term motions. However, current research overlooks these crucial aspects, significantly impacting accuracy. To tackle these issues, we propose a novel CSLR framework: Or…

Cited by 0SourcePDFScholar
2025

PersonaX: A Recommendation Agent-Oriented User Modeling Framework for Long Behavior Sequence

ACL 2025finding

User profile embedded in the prompt template of personalized recommendation agents play a crucial role in shaping their decision-making process. High-quality user profiles are essential for aligning agent behavior with real user interests. Typically, these profiles are constructed by leveraging LLMs…

2025

TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection

ICASSP 2025accepted

Object detection has witnessed remarkable advancements over the past decade, largely driven by breakthroughs in deep learning and the proliferation of large-scale datasets. However, the domain of road damage detection remains relatively underexplored, despite its critical significance for applicatio…

Cited by 0SourceScholar
2025

Toward Material-Agnostic System Identification from Videos

ICCV 2025poster

System identification from videos aims to recover object geometry and governing physical laws. Existing methods integrate differentiable rendering with simulation but rely on predefined material priors, limiting their ability to handle unknown ones. We introduce MASIV, the first vision-based framewo…

2025

Unsupervised Identification of Protein Compositions and Conformations via Implicit Content-Transformation Disentanglement

ICCV 2025poster

Identifying different protein compositions and conformations from microscopic images of protein mixtures is a challenging open problem. We address this through disentangled representation learning, where separating protein compositions and conformations in an intermediate latent space enables accura…

Cited by 0SourcePDFScholar
2025

Vox-UDA: Voxel-wise Unsupervised Domain Adaptation for Cryo-Electron Subtomogram Segmentation with Denoised Pseudo-Labeling

AAAI 2025technical

Cryo-Electron Tomography (cryo-ET) is a 3D imaging technology that facilitates the study of macromolecular structures at near-atomic resolution. Recent volumetric segmentation approaches on cryo-ET images have drawn widespread interest in the biological sector. However, existing methods heavily rely…

2025

iAgent: LLM Agent as a Shield between User and Recommender Systems

ACL 2025finding

Traditional recommender systems usually take the user-platform paradigm, where users are directly exposed under the control of the platform’s recommendation algorithms. However, the defect of recommendation algorithms may put users in very vulnerable positions under this paradigm. First, many sophis…

2024

Deep Active Learning with Noise Stability

AAAI 2024technical

Uncertainty estimation for unlabeled data is crucial to active learning. With a deep neural network employed as the backbone model, the data selection process is highly challenging due to the potential over-confidence of the model inference. Existing methods resort to special learning fashions (e.g.…

Cited by 19SourcePDFScholar
2024

MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with Transformer

AAAI 2024technical

The Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the c…

2024

Metric from Human: Zero-shot Monocular Metric Depth Estimation via Test-time Adaptation

NeurIPS 2024poster

Monocular depth estimation (MDE) is fundamental for deriving 3D scene structures from 2D images. While state-of-the-art monocular relative depth estimation (MRDE) excels in estimating relative depths for in-the-wild images, current monocular metric depth estimation (MMDE) approaches still face chall…

Cited by 4SourcePDFScholar
2024

RealNet: A Feature Selection Network with Realistic Synthetic Anomaly for Anomaly Detection

CVPR 2024poster

Self-supervised feature reconstruction methods have shown promising advances in industrial image anomaly detection and localization. Despite this progress these methods still face challenges in synthesizing realistic and diverse anomaly samples as well as addressing the feature redundancy and pre-tr…

2024

Synergistic Global-space Camera and Human Reconstruction from Videos

CVPR 2024poster

Remarkable strides have been made in reconstructing static scenes or human bodies from monocular videos. Yet the two problems have largely been approached independently without much synergy. Most visual SLAM methods can only reconstruct camera trajectories and scene structures up to scale while most…

Cited by 3SourcePDFScholar
2023

BiCro: Noisy Correspondence Rectification for Multi-Modality Data via Bi-Directional Cross-Modal Similarity Consistency

CVPR 2023poster

As one of the most fundamental techniques in multimodal learning, cross-modal matching aims to project various sensory modalities into a shared feature space. To achieve this, massive and correctly aligned data pairs are required for model training. However, unlike unimodal datasets, multimodal data…

2023

Dataset Pruning: Reducing Training Data by Examining Generalization Influence

ICLR 2023poster

The great success of deep learning heavily relies on increasingly larger training data, which comes at a price of huge computational and infrastructural costs. This poses crucial questions that, do all training data contribute to model's performance? How much does each individual training sample or…

Cited by 135SourcePDFScholar
2022

Boosting Active Learning via Improving Test Performance

AAAI 2022technical

Central to active learning (AL) is what data should be selected for annotation. Existing works attempt to select highly uncertain or informative data for annotation. Nevertheless, it remains unclear how selected data impacts the test performance of the task model used in AL. In this work, we explore…

2022

Estimating Instance-dependent Bayes-label Transition Matrix using a Deep Neural Network

ICML 2022spotlight

In label-noise learning, estimating the transition matrix is a hot topic as the matrix plays an important role in building statistically consistent classifiers. Traditionally, the transition from clean labels to noisy labels (i.e., clean-label transition matrix (CLTM)) has been widely exploited to l…

Cited by 64SourcePDFScholar
2022

Harmony: A Generic Unsupervised Approach for Disentangling Semantic Content From Parameterized Transformations

CVPR 2022poster

In many real-life image analysis applications, particularly in biomedical research domains, the objects of interest undergo multiple transformations that alters their visual properties while keeping the semantic content unchanged. Disentangling images into semantic content factors and transformation…

Cited by 8PDFScholar
2022

Sparse Local Patch Transformer for Robust Face Alignment and Landmarks Inherent Relation Learning

CVPR 2022poster

Heatmap regression methods have dominated face alignment area in recent years while they ignore the inherent relation between different landmarks. In this paper, we propose a Sparse Local Patch Transformer (SLPT) for learning the inherent relation. The SLPT generates the representation of each singl…

Cited by 62PDFcodeScholar
2021

Weakly Supervised 3D Semantic Segmentation Using Cross-Image Consensus and Inter-Voxel Affinity Relations

ICCV 2021poster

We propose a novel weakly supervised approach for 3D semantic segmentation on volumetric images. Unlike most existing methods that require voxel-wise densely labeled training data, our weakly-supervised CIVA-Net is the first model that only needs image-level class labels as guidance to learn accurat…

Cited by 20PDFcodeScholar
2020

Gum-Net: Unsupervised Geometric Matching for Fast and Accurate 3D Subtomogram Image Alignment and Averaging

CVPR 2020poster

We propose a Geometric unsupervised matching Net-work (Gum-Net) for finding the geometric correspondence between two images with application to 3D subtomogram alignment and averaging. Subtomogram alignment is the most important task in cryo-electron tomography (cryo-ET), a revolutionary 3D imaging t…

Cited by 23PDFcodeScholar
2019

Structured Modeling of Joint Deep Feature and Prediction Refinement for Salient Object Detection

ICCV 2019poster

Recent saliency models extensively explore to incorporate multi-scale contextual information from Convolutional Neural Networks (CNNs). Besides direct fusion strategies, many approaches introduce message-passing to enhance CNN features or predictions. However, the messages are mainly transmitted in…

Cited by 59PDFScholar
2017

Chromatic surface microstructures on bionic soft robots for non-contact deformation measurement

ICRA 2017poster

This paper presents a bionic soft robot with chromatic surface micro-structure (CSM), as a new approach for the measurement of body deformation of the soft robots. Firstly, the CSM films are fabricated by diffraction gratings mold using material polydimethylsiloxane (PDMS). Then, the CSM films are a…

Cited by 6SourceScholar
2016

Locomotion and gait analysis of multi-limb soft robots driven by smart actuators

IROS 2016poster

Animals provide inspiration for developing soft machines in bionics, robotics research as well as potential applications. This paper presents an integrated development of locomotive soft robot platforms: starfish-inspired robots with multi-limb bodies actuated by shape memory alloys (SMAs). The desi…

Cited by 25SourceScholar