← Search

Feng Yu

21 accepted papers

2026

AI-IO: An Aerodynamics-Inspired Real-Time Inertial Odometry for Quadrotors

ICRA 2026poster

Inertial Odometry (IO) has gained attention in quadrotor applications due to its sole reliance on inertial measurement units (IMUs), attributed to its lightweight design, low cost, and robust performance across diverse environments. However, most existing learning-based inertial odometry systems for…

2026

Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation

ICLR 2026poster

Mainstream Test-Time Adaptation (TTA) methods for adapting vision-language models, e.g., CLIP, typically rely on Shannon Entropy (SE) at test time to measure prediction uncertainty and inconsistency. However, since CLIP has a built-in bias from pretraining on highly imbalanced web-crawled data, SE i…

Cited by 0SourcecodeScholar
2026

Mastering Diverse, Unknown, and Cluttered Tracks for Robust Vision-Based Drone Racing

RA-L 2026

Most reinforcement learning (RL)-based methods for drone racing target fixed, obstacle-free tracks, leaving the generalization to unknown, cluttered environments largely unaddressed. This challenge stems from the need to balance racing speed and collision avoidance, limited feasible space causing po

Cited by 3SourceScholar
2026

Vector Field Augmented Differentiable Policy Learning for Vision-Based Drone Racing

RA-L 2026

Autonomous drone racing in complex environments requires agile, high-speed flight while maintaining reliable obstacle avoidance. Differentiable-physics-based policy learning has recently demonstrated high sample efficiency and remarkable performance across various tasks, including agile drone flight

Cited by 0SourceScholar
2026

Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation

AAAI 2026technical

Video-to-Music generation seeks to generate musically appropriate background music that enhances audiovisual immersion for videos. However, current approaches suffer from two critical limitations: 1) incomplete representation of video details, leading to weak alignment, and 2) inadequate temporal an

Cited by 0SourcePDFScholar
2026

Vision-Based End-to-End Learning for UAV Traversal of Irregular Gaps via Differentiable Simulation

RA-L 2026

Navigation through narrow and irregular gaps is an essential skill in autonomous drones for applications such as inspection, search-and-rescue, and disaster response. However, traditional planning and control methods rely on explicit gap extraction and measurement, while recent end-to-end approaches

Cited by 0SourceScholar
2025

CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models

NAACL 2025findings

Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible…

2025

CLaMP 3: Universal Music Information Retrieval Across Unaligned Modalities and Unseen Languages

ACL 2025finding

CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities–including sheet music, performance signals, and audio recordings–with multilingual text in a…

2025

GarFast: Realistic and Fast Garment Transfer with a Simplified Parser-Free Approach

AAAI 2025technical

A good garment try-on model should learn the transfer between different types of garments while satisfying: 1) high fidelity and 2) low inference speed. Existing methods address either of these two issues, limited processing speed or low generation quality. We directly use a lightweight encoder-deco…

Cited by 0SourcePDFScholar
2025

Latent Diffusion-Enhanced Virtual Try-On via Optimized Pseudo-Label Generation

AAAI 2025technical

Efficiently applying fully supervised learning to virtual try-on tasks is challenging due to the lack of paired ground truth in available training samples. Recent works have achieved virtual try-ons by employing self-supervised learning-based inpainting paradigms. However, this approach is heavily d…

Cited by 0SourcePDFScholar
2025

Mapless Collision-Free Flight via MPC using Dual KD-Trees in Cluttered Environments

IROS 2025

Collision-free flight in cluttered environments is a critical capability for autonomous quadrotors. Traditional methods often rely on detailed 3D map construction, trajectory generation, and tracking. However, this cascade pipeline can introduce accumulated errors and computational delays, limiting

Cited by 3SourcecodeScholar
2025

Multi-Label Test-Time Adaptation with Bound Entropy Minimization

ICLR 2025poster

Mainstream test-time adaptation (TTA) techniques endeavor to mitigate distribution shifts via entropy minimization for multi-class classification, inherently increasing the probability of the most confident class. However, when encountering multi-label instances, the primary challenge stems from the…

2025

NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms

IJCAI 2025

We introduce NotaGen, a symbolic music generation model aiming to explore the potential of producing high-quality classical sheet music. Inspired by the success of Large Language Models (LLMs), NotaGen adopts pre-training, fine-tuning, and reinforcement learning paradigms (henceforth referred to as

2025

SuperLightNet: Lightweight Parameter Aggregation Network for Multimodal Brain Tumor Segmentation

CVPR 2025poster

Multimodal 3D segmentation involves a significant number of 3D convolution operations, which requires substantial computational resources and high-performance computing devices in MRI multimodal brain tumor segmentation. The key challenge in multimodal 3D segmentation is how to minimize network comp…

2024

A Bi-Pyramid Multimodal Fusion Method for the Diagnosis Of Bipolar Disorders

ICASSP 2024accepted

Previous research on the diagnosis of Bipolar disorder has mainly focused on resting-state functional magnetic resonance imaging. However, their accuracy can not meet the requirements of clinical diagnosis. Efficient multimodal fusion strategies have great potential for applications in multimodal da…

Cited by 0SourceScholar
2024

A Subspace-Constrained Tyler's Estimator and its Applications to Structure from Motion

CVPR 2024poster

We present the subspace-constrained Tyler's estimator (STE) designed for recovering a low-dimensional subspace within a dataset that may be highly corrupted with outliers. STE is a fusion of the Tyler's M-estimator (TME) and a variant of the fast median subspace. Our theoretical analysis suggests th…

2024

S-MolSearch: 3D Semi-supervised Contrastive Learning for Bioactive Molecule Search

NeurIPS 2024poster

Virtual Screening is an essential technique in the early phases of drug discovery, aimed at identifying promising drug candidates from vast molecular libraries. Recently, ligand-based virtual screening has garnered significant attention due to its efficacy in conducting extensive database screening…

Cited by 1SourcePDFScholar
2024

SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection

NeurIPS 2024poster

Detection of face forgery videos remains a formidable challenge in the field of digital forensics, especially the generalization to unseen datasets and common perturbations. In this paper, we tackle this issue by leveraging the synergy between audio and visual speech elements, embarking on a novel a…

Cited by 1SourcePDFScholar
2022

Multi-Pose Virtual Try-On Via Self-Adaptive Feature Filtering

ICASSP 2022accepted

With the growing trend of virtual try-on, multi-pose tasks attract researchers due to their higher commercial value. Prior methods lack an effective geometric deformation to maintain the original image details resulting in many details loss in the head and garment. To address this problem, we propos…

Cited by 0SourceScholar
2022

Realistic Monocular-To-3d Virtual Try-On Via Multi-Scale Characteristics Capture

ICASSP 2022accepted

3D virtual try-on receives widespread attention from scholars due to its great practical and commercial values. In prior methods, the fundamental problems lie in the limitations on texture retention during garment deformation and the lack of feature context capture during depth estimation. To addres…

Cited by 0SourceScholar