← Search

Xiangyu Wu

16 accepted papers

2026

Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation

ICLR 2026poster

Mainstream Test-Time Adaptation (TTA) methods for adapting vision-language models, e.g., CLIP, typically rely on Shannon Entropy (SE) at test time to measure prediction uncertainty and inconsistency. However, since CLIP has a built-in bias from pretraining on highly imbalanced web-crawled data, SE i…

Cited by 0SourcecodeScholar
2026

Beyond Tokens: Dynamic Latent Reasoning via Semantic Residual Refinement

AAAI 2026technical

Chain-of-Thought prompting has remarkably advanced LLM reasoning by generating explicit step-by-step tokens, yet its discrete nature inherently limits expressiveness and efficiency, struggling with abstract, ambiguous, or semantically divergent cognition beyond linguistic tokens. Latent reasoning of

Cited by 0SourcePDFScholar
2026

Hierarchical Enhancement of Semantic Priors for Disentangled Text-Driven Motion Generation

CVPR 2026

Text-to-motion generation aims to synthesize realistic and semantically aligned 3D human motions from natural language descriptions. Existing diffusion-based methods often rely on isotropic latent priors and shallow cross-modal supervision, which lead to semantic entanglement, limited controllabilit

Cited by 0SourceScholar
2026

MetaphorVU: Towards Metaphorical Video Understanding

ICML 2026spotlight

Metaphorical videos are prevalent across various real-world scenarios to convey complex ideas, and understanding them typically requires high-order cognitive capabilities. The lack of systematic studies on metaphorical video understanding not only constrains the real-world applicability of MLLMs but…

Cited by 0SourceScholar
2026

Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding

AAAI 2026technical

Multimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering f

Cited by 0SourcePDFScholar
2026

Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?

ICML 2026poster

Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in visual understanding, yet their performance degrades significantly under real-world visual corruptions. While existing robustness enhancement approaches exist, they are limited: black-box feature alignment lacks interpr…

Cited by 0SourceScholar
2025

Multi-Label Test-Time Adaptation with Bound Entropy Minimization

ICLR 2025poster

Mainstream test-time adaptation (TTA) techniques endeavor to mitigate distribution shifts via entropy minimization for multi-class classification, inherently increasing the probability of the most confident class. However, when encountering multi-label instances, the primary challenge stems from the…

2024

TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt

IJCAI 2024poster

The recent introduction of prompt tuning based on pre-trained vision-language models has dramatically improved the performance of multi-label image classification. However, some existing strategies that have been explored still have drawbacks, i.e., either exploiting massive labeled visual data at a…

2023

Design, Modeling and Control of a Top-Loading Fully-Actuated Cargo Transportation Multirotor

RA-L 2023

Existing multirotor-based cargo transportation does not maintain a constant cargo attitude due to underactuation; however, fragile payloads may require a consistent posture. The conventional method is also cumbersome when loading cargo, and the size of the cargo to be loaded is limited. To overcome

Cited by 7SourceScholar
2023

Learning a Single Near-hover Position Controller for Vastly Different Quadcopters

ICRA 2023poster

This paper proposes an adaptive near-hover position controller for quadcopters, which can be deployed to quadcopters of very different mass, size and motor constants, and also shows rapid adaptation to unknown disturbances during runtime. The core algorithmic idea is to learn a single policy that ca…

Cited by 23SourceScholar
2021

Real-time Geo-localization Using Satellite Imagery and Topography for Unmanned Aerial Vehicles

IROS 2021poster

The capabilities of autonomous flight with unmanned aerial vehicles (UAVs) have significantly increased in recent times. However, basic problems such as fast and robust geo-localization in GPS-denied environments still remain unsolved. Existing research has primarily concentrated on improving the ac…

Cited by 30SourceScholar
2020

A collision-resilient aerial vehicle with icosahedron tensegrity structure

IROS 2020poster

Aerial vehicles with collision resilience can operate with more confidence in environments with obstacles that are hard to detect and avoid. This paper presents the methodology used to design a collision resilient aerial vehicle with icosahedron tensegrity structure. A simplified stress analysis of…

Cited by 56SourceScholar
2020

In-flight range optimization of multicopters using multivariable extremum seeking with adaptive step size

IROS 2020poster

Limited flight range is a common problem for multicopters. To alleviate this problem, we propose a method for finding the optimal speed and heading of a multicopter when flying a given path to achieve the longest flight range. Based on a novel multivariable extremum seeking controller with adaptive…

Cited by 12SourceScholar
2019

Model-free Online Motion Adaptation for Optimal Range and Endurance of Multicopters

ICRA 2019poster

In this work we introduce an approach that allows a quadcopter to find the velocity which maximizes its flight time (endurance) or flight distance (range) while moving along a given path, using on-board power measurement. The proposed strategy is based on Extremum Seeking control and (a) does not re…

Cited by 29SourceScholar