← Search

YU HU

55 accepted papers

2026

AI-IO: An Aerodynamics-Inspired Real-Time Inertial Odometry for Quadrotors

ICRA 2026poster

Inertial Odometry (IO) has gained attention in quadrotor applications due to its sole reliance on inertial measurement units (IMUs), attributed to its lightweight design, low cost, and robust performance across diverse environments. However, most existing learning-based inertial odometry systems for…

2026

AVION: Aerial Vision-Language Instruction from Offline Teacher to Prompt-Tuned Network

CVPR 2026

Adapting vision-language models to remote sensing imagery remains challenging due to two key factors: limited semantic coverage in textual representations and insufficient adaptability of visual features. These issues are particularly significant in aerial scenes, which involve various visual appear

Cited by 0SourcecodeScholar
2026

Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks

ICRA 2026poster

A major bottleneck in off-road autonomous driving research lies in the scarcity of large-scale, high-quality datasets and benchmarks. To bridge this gap, we present ORAD-3D, which, to the best of our knowledge, is the largest dataset specifically curated for off-road autonomous driving. ORAD-3D cove…

2026

Beyond Endpoints: Path-Centric Reasoning for Vectorized Off-Road Network Extraction

CVPR 2026

Deep learning has advanced vectorized road extraction in urban settings, yet off-road environments remain underexplored and challenging. A significant domain gap causes advanced models to fail in wild terrains due to two key issues: lack of large-scale vectorized datasets and structural weakness in

Cited by 0SourcecodeScholar
2026

Curriculum Reinforcement Learning for Quadrotor Racing with Random Obstacles

ICRA 2026poster

Autonomous drone racing has attracted increasing interest as a research topic for exploring the limits of agile flight. However, existing studies primarily focus on obstacle free racetracks, while the perception and dynamic challenges introduced by obstacles remain underexplored, often resulting in …

2026

Mastering Diverse, Unknown, and Cluttered Tracks for Robust Vision-Based Drone Racing

RA-L 2026

Most reinforcement learning (RL)-based methods for drone racing target fixed, obstacle-free tracks, leaving the generalization to unknown, cluttered environments largely unaddressed. This challenge stems from the need to balance racing speed and collision avoidance, limited feasible space causing po

Cited by 3SourceScholar
2026

Vector Field Augmented Differentiable Policy Learning for Vision-Based Drone Racing

RA-L 2026

Autonomous drone racing in complex environments requires agile, high-speed flight while maintaining reliable obstacle avoidance. Differentiable-physics-based policy learning has recently demonstrated high sample efficiency and remarkable performance across various tasks, including agile drone flight

Cited by 0SourceScholar
2026

Vision-Based End-to-End Learning for UAV Traversal of Irregular Gaps via Differentiable Simulation

RA-L 2026

Navigation through narrow and irregular gaps is an essential skill in autonomous drones for applications such as inspection, search-and-rescue, and disaster response. However, traditional planning and control methods rely on explicit gap extraction and measurement, while recent end-to-end approaches

Cited by 0SourceScholar
2025

CA2Point: Learning Keypoint Detection and Description with Context Aggregation and Cross Augmentation

IROS 2025

Keypoint detection and description are fundamental tasks for a variety of computer vision applications. Due to the limited receptive field of convolutional neural networks, most existing methods based on deep learning mainly focus on the local features, instead of taking into account the global cont

Cited by 0SourcecodeScholar
2025

CORENet: Cross-Modal 4D Radar Denoising Network with LiDAR Supervision for Autonomous Driving

IROS 2025

4D radar-based object detection has garnered great attention for its robustness in adverse weather conditions and capacity to deliver rich spatial information across diverse driving scenarios. Nevertheless, the sparse and noisy nature of 4D radar point clouds poses substantial challenges for effecti

Cited by 0SourcecodeScholar
2025

Efficient Dynamic Ensembling for Multiple LLM Experts

IJCAI 2025

LLMs have demonstrated impressive performance across various language tasks. However, the strengths of LLMs can vary due to different architectures, model sizes, areas of training data, etc. Therefore, ensemble reasoning for the strengths of different LLM experts is critical to achieving consistent

2025

Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance

IJCAI 2025

Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding a

2025

Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement

NAACL 2025long

Long-form story generation task aims to produce coherent and sufficiently lengthy text, essential for applications such as novel writingand interactive storytelling. However, existing methods, including LLMs, rely on rigid outlines or lack macro-level planning, making it difficult to achieve both co…

2025

Indoor Geomagnetic Matching Location Based on Iterative Local Search and Improved Particle Swarm Fusion

RA-L 2025

In the complex indoor environment, geomagnetic matching is an effective way to realize indoor positioning of mobile robots. Aiming at the problem that the application of Particle Swarm Optimization (PSO) algorithm leads to the decline of geomagnetic matching accuracy, stability and convergence speed

Cited by 5SourceScholar
2025

Mapless Collision-Free Flight via MPC using Dual KD-Trees in Cluttered Environments

IROS 2025

Collision-free flight in cluttered environments is a critical capability for autonomous quadrotors. Traditional methods often rely on detailed 3D map construction, trajectory generation, and tracking. However, this cascade pipeline can introduce accumulated errors and computational delays, limiting

Cited by 3SourcecodeScholar
2025

MedCite: Can Language Models Generate Verifiable Text for Medicine?

ACL 2025finding

Existing LLM-based medical question answering systems lack citation generation and evaluation capabilities, raising concerns about their adoption in practice. In this work, we introduce MedCite, the first end-to-end framework that facilitates the design and evaluation of LLM citations for medical ta…

2025

ROD: RGB-Only Fast and Efficient Off-Road Freespace Detection

ICRA 2025

Off-road freespace detection is more challenging than on-road scenarios because of the blurred boundaries of traversable areas. Previous state-of-the-art (SOTA) methods employ multi-modal fusion of RGB images and LiDAR data. However, due to the significant increase in inference time when calculating

Cited by 2SourcecodeScholar
2025

RegGS: Unposed Sparse Views Gaussian Splatting with 3DGS Registration

ICCV 2025poster

3D Gaussian Splatting (3DGS) has demonstrated its potential in reconstructing scenes from unposed images. However, optimization-based 3DGS methods struggle with sparse views due to limited prior knowledge. Meanwhile, feed-forward Gaussian approaches are constrained by input formats, making it challe…

Cited by 0SourcePDFScholar
2025

Vibration-Aware Trajectory Optimization for Mobile Robots in Wild Environments via Physics-Informed Neural Network

IROS 2025

The suspension system, through effective damping of vibrations and shocks, can enhance the stability of wheeled robots traversing challenging terrain. Because the suspension system decouples the rigid correspondence between terrain changes and robot vibrations, considering suspension modeling in tra

Cited by 0SourceScholar
2024

A Safe and Efficient Timed-Elastic-Band Planner for Unstructured Environments

IROS 2024poster

In unstructured environments with complex obstacles and obscure road boundaries, the local planner faces more severe challenges in terms of safety and real-time performance. In order to fulfill these emerging requirements, we propose a novel Timed-Elastic-Band approach for unstructured environments,…

Cited by 2SourceScholar
2024

Preliminary Result of Cury: A Backdrivable Leg Design Using Linear Actuators

IROS 2024poster

This paper reports the design, simulation, and experiment of a robotic leg prototype named Cury, which has the potential to achieve minimal clearance and excellent backdrivability. Inspired by human walking data, the actuator design incorporates four-bar linkages and ball screws and is further optim…

Cited by 1SourceScholar
2024

SCOML: Trajectory Planning Based on Self-Correcting Meta-Reinforcement Learning in Hybrid Terrain for Mobile Robot

IROS 2024

Trajectory planning is important for ground robots to achieve safe and efficient autonomous navigation in unstructured off-road environments. Most existing methods treat each terrain as a single type. However, in the real world, a ground usually consists of hybrid terrains. In this work, we propose

Cited by 2SourceScholar
2024

SRFUND: A Multi-Granularity Hierarchical Structure Reconstruction Benchmark in Form Understanding

NeurIPS 2024poster

Accurately identifying and organizing textual content is crucial for the automation of document processing in the field of form understanding. Existing datasets, such as FUNSD and XFUND, support entity classification and relationship prediction tasks but are typically limited to local and entity-lev…

2024

TeFF: Tracking-enhanced Forgetting-free Few-shot 3D LiDAR Semantic Segmentation

IROS 2024

In autonomous driving, 3D LiDAR plays a crucial role in understanding the vehicle’s surroundings. However, the newly emerged, unannotated objects presents few-shot learning problem for semantic segmentation. This paper addresses the limitations of current few-shot semantic segmentation by exploiting

Cited by 2SourcecodeScholar
2024

UniTabNet: Bridging Vision and Language Models for Enhanced Table Structure Recognition

EMNLP 2024finding

In the digital era, table structure recognition technology is a critical tool for processing and analyzing large volumes of tabular data. Previous methods primarily focus on visual aspects of table structure recovery but often fail to effectively comprehend the textual semantics within tables, parti…

2023

FeatureBooster: Boosting Feature Descriptors With a Lightweight Neural Network

CVPR 2023poster

We introduce a lightweight network to improve descriptors of keypoints within the same image. The network takes the original descriptors and the geometric properties of keypoints as the input, and uses an MLP-based self-boosting stage and a Transformer-based cross-boosting stage to enhance the descr…

2023

PA&DA: Jointly Sampling Path and Data for Consistent NAS

CVPR 2023poster

Based on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during tra…

2023

PINAT: A Permutation INvariance Augmented Transformer for NAS Predictor

AAAI 2023technical

Time-consuming performance evaluation is the bottleneck of traditional Neural Architecture Search (NAS) methods. Predictor-based NAS can speed up performance evaluation by directly predicting performance, rather than training a large number of sub-models and then validating their performance. Most p…

2023

Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NAS

ICCV 2023poster

Neural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate architectures without any training. However, we have observed that most existing pr…

Cited by 6PDFcodeScholar
2022

AGNAS: Attention-Guided Micro and Macro-Architecture Search

ICML 2022spotlight

Micro- and macro-architecture search have emerged as two popular NAS paradigms recently. Existing methods leverage different search strategies for searching micro- and macro- architectures. When using architecture parameters to search for micro-structure such as normal cell and reduction cell, the a…

2022

Model-Based Contact Detection and Accommodation for Soft Bending Actuators: An Integrated Direct/Indirect Adaptive Robust Approach

RA-L 2022

Soft robots have intrinsic advantages in interaction with humans or complex environments for actual applications, during which various external disturbances (e.g., external contact or collision) are inevitable. They show remarkable abilities in complicated tasks due to their easily deformable bodies

Cited by 6SourceScholar
2022

SMS-MPC: Adversarial Learning-based Simultaneous Prediction Control with Single Model for Mobile Robots

IROS 2022poster

Model predictive control is a promising method in robot control tasks. How to design an effective model structure and efficient prediction framework for model predictive control is still an open challenge. To reduce the time consumption and avoid compounding-error of the multi-step prediction proces…

Cited by 3SourceScholar
2022

Searching for BurgerFormer with Micro-Meso-Macro Space Design

ICML 2022spotlight

With the success of Transformers in the computer vision field, the automated design of vision Transformers has attracted significant attention. Recently, MetaFormer found that simple average pooling can achieve impressive performance, which naturally raises the question of how to design a search spa…

2021

Savable but Lost Lives when ICU Is Overloaded: a Model from 733 Patients in Epicenter Wuhan, China

AAAI 2021technical

Coronavirus Disease 2019 (COVID-19) causes a sudden turnover to bad at some checkpoints and thus needs the intervention of intensive care unit (ICU). This resulted in urgent and large needs of ICUs posed great risks to the medical system. Estimating the mortality of critical in-patients who were not…

Cited by 1SourcePDFScholar
2021

Tracking Interaction States for Multi-Turn Text-to-SQL Semantic Parsing

AAAI 2021technical

The task of multi-turn text-to-SQL semantic parsing aims to translate natural language utterances in an interaction into SQL queries in order to answer them using a database which normally contains multiple table schemas. Previous studies on this task usually utilized contextual information to enric…

2020

Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction

CVPR 2020poster

Video prediction is a pixel-wise dense prediction task to infer future frames based on past frames. Missing appearance details and motion blur are still two major problems for current models, leading to image distortion and temporal inconsistency. We point out the necessity of exploring multi-freque…

Cited by 134PDFcodeScholar
2018

Quantization of Fully Convolutional Networks for Accurate Biomedical Image Segmentation

CVPR 2018poster

With pervasive applications of medical imaging in healthcare, biomedical image segmentation plays a central role in quantitative analysis, clinical diagnosis, and medical intervention. Since manual annotation suffers limited reproducibility, arduous efforts, and excessive time, automatic segmentatio…

Cited by 122SourcePDFScholar
2018

RT3D: Real-Time 3-D Vehicle Detection in LiDAR Point Cloud for Autonomous Driving

RA-L 2018

For autonomous driving, vehicle detection is the prerequisite for many tasks like collision avoidance and path planning. In this letter, we present a real-time three-dimensional (RT3D) vehicle detection method that utilizes pure LiDAR point cloud to predict the location, orientation, and size of veh

Cited by 175SourceScholar
2018

See and Think: Disentangling Semantic Scene Completion

NeurIPS 2018poster

Semantic scene completion predicts volumetric occupancy and object category of a 3D scene, which helps intelligent agents to understand and interact with the surroundings. In this work, we propose a disentangled framework, sequentially carrying out 2D semantic segmentation, 2D-3D reprojection and 3D…

2018

VarNet: Exploring Variations for Unsupervised Video Prediction

IROS 2018poster

Unsupervised video prediction is a very challenging task due to the complexity and diversity in natural scenes. Prior works directly predicting pixels or optical flows either have the blurring problem or require additional assumptions. We highlight that the crux for video frame prediction lies in pr…

Cited by 39SourcecodeScholar
2017

GeoCueDepth: Exploiting geometric structure cues to estimate depth from a single image

IROS 2017poster

Depth estimation from a single image is very challenging due to the inherent ambiguity of mapping a color image to a depth map. Previous work tackles this problem by exploiting various levels of features with multi-scale deep convolutional neural networks. However, most of the local geometric struct…

Cited by 7SourceScholar
2016

Modulation spectrum compensation for HMM-based speech synthesis using line spectral pairs

ICASSP 2016accepted

In previous work, a method to compensate the divergence between the distributions of natural and generated modulation spectra (MS) has been proposed for hidden Markov model (HMM) based speech synthesis. This method can alleviate the over-smoothing effect of parameter generation when Mel-cepstral coe…

Cited by 0SourceScholar