← Search

Yuan Xu

12 accepted papers

2026

Adaptor: Advancing Assistive Teleoperation with Few-Shot Learning and Cross-Operator Generalization

ICRA 2026poster

Assistive teleoperation enhances efficiency via shared control, yet inter-operator variability, stemming from diverse habits and expertise, induces highly heterogeneous trajectory distributions that undermine intent recognition stability. We present Adaptor, a few-shot framework for robust cross-ope…

2026

FrontierCS: Evolving Challenges for Evolving Intelligence

ICML 2026poster

We introduce FrontierCS, a benchmark of 240 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competitive programming participants and problem setters. Unlike existing benchmarks that focus on tasks with known optimal solut…

Cited by 0SourceScholar
2025

Detecting Perception-Based Attacks using Visual Odometry: Inconsistency Modeling and Checking on Robotic States

ICRA 2025

Perception systems in robotic vehicles are crucial for safe and efficient operation, providing key state estimates necessary for planning and control. However, these systems are increasingly vulnerable to perception-based attacks, such as odometry spoofing, position spoofing, obstacle hiding, and ob

Cited by 1SourceScholar
2025

Latent Embedding Adaptation for Human Preference Alignment in Diffusion Planners

ICRA 2025

This work addresses the challenge of personalizing trajectories generated in automated decision-making systems by introducing a resource-efficient approach that enables rapid adaptation to individual users' preferences. Our method leverages a pretrained conditional diffusion model with Preference La

Cited by 1SourceScholar
2024

FedCDA: Federated Learning with Cross-rounds Divergence-aware Aggregation

ICLR 2024poster

In Federated Learning (FL), model aggregation is pivotal. It involves a global server iteratively aggregating client local trained models in successive rounds without accessing private data. Traditional methods typically aggregate the local models from the current round alone. However, due to the st…

Cited by 31SourcePDFScholar
2024

Investigating and Mitigating the Multimodal Hallucination Snowballing in Large Vision-Language Models

ACL 2024long

Though advanced in understanding visual information with human languages, Large Vision-Language Models (LVLMs) still suffer from multimodal hallucinations. A natural concern is that during multimodal interaction, the generated hallucinations could influence the LVLMs’ subsequent generation. Thus, we…

2024

MCM-CSD: Multi-Granularity Context Modeling with Contrastive Speaker Detection for Emotion Recognition in Real-Time Conversation

ICASSP 2024accepted

Emotion recognition in conversation (ERC) has received extensive attention for its wide applications in recent years. Considering the actual situation, we focus on the real-time conversation scenarios, in which how to model the conversation emotion with only the historical contextual information and…

Cited by 0SourceScholar
2024

Probing Synergistic High-Order Interaction in Infrared and Visible Image Fusion

CVPR 2024poster

Infrared and visible image fusion aims to generate a fused image by integrating and distinguishing complementary information from multiple sources. While the cross-attention mechanism with global spatial interactions appears promising it only capture second-order spatial interactions neglecting high…

2024

ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis Scoring

CVPR 2024poster

Monocular 3D human mesh estimation is an ill-posed problem characterized by inherent ambiguity and occlusion. While recent probabilistic methods propose generating multiple solutions little attention is paid to obtaining high-quality estimates from them. To address this limitation we introduce Score…

2023

FouriDown: Factoring Down-Sampling into Shuffling and Superposing

NeurIPS 2023poster

Spatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employ…

2023

Training Your Image Restoration Network Better with Random Weight Network as Optimization Function

NeurIPS 2023poster

The blooming progress made in deep learning-based image restoration has been largely attributed to the availability of high-quality, large-scale datasets and advanced network structures. However, optimization functions such as L_1 and L_2 are still de facto. In this study, we propose to investigate…

Cited by 1SourcePDFScholar
2022

Bidirectional Sim-to-Real Transfer for GelSight Tactile Sensors With CycleGAN

RA-L 2022

GelSight optical tactile sensors have high-resolution and low-cost advantages and have witnessed growing adoption in various contact-rich robotic applications. Sim2Real for GelSight sensors can reduce the time cost and sensor damage during data collection and is crucial for learning-based tactile pe

Cited by 46SourcecodeScholar