← Search

Yaonan Wang

61 accepted papers

2026

CRAFT: Adapting VLA Models to Contact-Rich Manipulation Via Force-Aware Curriculum Fine-Tuning

ICRA 2026poster

Vision-Language-Action (VLA) models have shown a strong capability in enabling robots to execute general instructions, yet they struggle with contact-rich manipulation tasks, where success requires precise alignment, stable contact maintenance,and effective handling of deformable objects. A fundamen…

2026

Cooperative Informed Tree (CoIT*): Cooperative Bi-Directional Multi-Resolution Motion Planning with Adaptive Edge Screening

ICRA 2026poster

In informed search-based path planning, heuristic functions that incorporate problem knowledge are essential for guiding the search and improving efficiency. The accuracy and computational cost of these heuristics are therefore critical to performance. However, accuracy and computational efficiency …

Cited by 0Scholar
2026

Design of an Adaptive Modular Anthropomorphic Dexterous Hand for Human-Like Manipulation

ICRA 2026poster

Biological synergies have emerged as a widely adopted paradigm for dexterous hand design, enabling human-like manipulation with a small number of actuators. Nonetheless, excessive coupling tends to diminish the dexterity of hands. This paper tackles the trade-off between actuation complexity and dex…

2026

Efficient Robotic 3D Measurement Through Multi-DoF Reinforcement Learning for Continuous Viewpoint Planning

RA-L 2026

Three-dimensional (3D) measurement is essential for quality control in manufacturing, especially for components with complex geometries. Conventional viewpoint planning methods based on fixed spherical coordinates often fail to capture intricate surfaces, leading to suboptimal reconstructions. To ad

Cited by 0SourceScholar
2026

GroupMIL: Semantic Group Based Multiple Instance Learning for Whole Slide Image Analysing

IJCAI 2026

Whole Slide Image (WSI) analysis faces challenges due to gigapixel resolutions and slide-level weak supervision. Multiple Instance Learning (MIL) serves as a pivotal method for this task. However, existing MIL frameworks often fail to exploit the inherent redundancy of tissue patterns or the semanti

Cited by 0Scholar
2026

Improving Explicit Dynamic Gaussian Splatting Optimization via Update Mixture

ICML 2026poster

3D Gaussian Splatting (3DGS) enables real-time, high-fidelity view synthesis via explicit scene representations and has recently been extended to dynamic scene modeling. In spite of excellent quality and interpretability, we find explicit Dynamic GS often exhibits generalization degradation in large…

Cited by 0SourceScholar
2026

Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D Generation

CVPR 2026

3D shape generation has become increasingly important for graphics and vision applications. Current part-aware 3D generation usually overlooks hierarchical part relations or inefficiently encodes multi-level semantics in Euclidean space. Thus we propose a novel framework for hierarchical and efficie

Cited by 0SourceScholar
2026

MIRNet: Integrating Constrained Graph-Based Reasoning with Pre-training for Diagnostic Medical Imaging

AAAI 2026technical

Automated interpretation of medical images demands robust modeling of complex visual-semantic relationships while addressing annotation scarcity, label imbalance, and clinical plausibility constraints. We introduce MIRNet (Medical Image Reasoner Network), a novel framework that integrates self-super

Cited by 0SourcePDFScholar
2026

Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding

AAAI 2026technical

Monocular 3D Visual Grounding (Mono3DVG) is an emerging task that locates 3D objects in RGB images using text descriptions with geometric cues. However, existing methods face two key limitations. Firstly, they often over-rely on high-certainty keywords that explicitly identify the target object whil

Cited by 0SourcePDFScholar
2026

Robust Fault-Tolerant Control for Underwater Vehicles With Saturation-Triggered Thruster Health Estimation

RA-L 2026

This letter presents a saturation-triggered adaptive fault-tolerant control framework for underwater vehicles to handle concurrent thruster failures and environmental disturbances. Unlike existing methods that suffer from performance degradation when disturbances exceed predefined bounds, this paper

Cited by 0SourceScholar
2026

SGDE: Self-supervised Geometry Degradation Estimation Framework for Coded Aperture Compressive Spectral Imaging

CVPR 2026

Coded Aperture Snapshot Spectral Imaging (CASSI) has emerged as a prominent technique for efficient hyperspectral imaging. However, the tight coupling between physical encoding and computational decoding makes CASSI highly sensitive to slight hardware misalignments, which can significantly degrade r

Cited by 0SourcecodeScholar
2026

Simultaneous Arrival Control for Distributed Multi-Robot Systems with Curvature and Constant-Speed Constraints

ICRA 2026poster

The simultaneous arrival of multiple mobile robots at their respective target points is crucial for cooperative tasks such as encirclement, interception, and disaster relief. Although the problem of simultaneous arrival is inherently complex, it becomes even more challenging in multi-robot systems w…

Cited by 0Scholar
2026

UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands

RA-L 2026

Dexterous grasp datasets are vital for embodied intelligence, but mostly emphasize grasp stability, ignoring functional grasps needed for tasks like opening bottle caps or holding cup handles. Most rely on bulky, costly, and hard-to-control high-DOF Shadow Hands. Inspired by the human hand's underac

Cited by 1SourceScholar
2026

UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands

ICRA 2026poster

Dexterous grasp datasets are vital for embodied intelligence, but mostly emphasize grasp stability, ignoring functional grasps needed for tasks like opening bottle caps or holding cup handles. Most rely on bulky, costly, and hard-to-control high-DOF Shadow Hands. Inspired by the human hand’s underac…

2026

Weighted Online Koopman Learning for Model Predictive Control of Thrust-Vectored Underwater Vehicles

RA-L 2026

Underwater vehicles face control challenges including hydrodynamic parameter uncertainties and environmental disturbances. While Koopman-based model predictive control (Koopman-MPC) offers computational efficiency through linear lifted representations, existing methods rely on offline-collected data

Cited by 0SourceScholar
2025

A Method for Constructing Building Structure Grid Map Based on a Climbing Algorithm

ICRA 2025

Aerial-terrestrial amphibious robots excel in search and rescue tasks in unstructured terrains but face challenges in autonomous navigation indoors. Traditional full-mapping methods can degrade global path planning performance, especially when semi-static obstacles shift, leading to suboptimal paths

Cited by 0SourceScholar
2025

A Parallel Fuzzy Nonlinear ADRC Framework for Robotic Machining with Vibrations

IROS 2025

This paper proposes a Parallel Fuzzy Nonlinear Active Disturbance Rejection Control (FNLADRC) strategy to improve the precision and robustness of robotic manipulators in machining large, complex components. By decoupling the multi-degree-of-freedom dynamics and integrating Nonlinear ADRC with fuzzy

Cited by 0SourceScholar
2025

Coordinated Energy-Trajectory Economic Model Predictive Control for Autonomous Surface Vehicles under Disturbances

IROS 2025

The paper proposes a novel Economic Model Predictive Control (EMPC) scheme for Autonomous Surface Vehicles (ASVs) to simultaneously address path following accuracy and energy constraints under environmental disturbances. By formulating lateral deviations as energy-equivalent penalties in the cost fu

Cited by 0SourceScholar
2025

Cross-Modal Interactive Perception Network with Mamba for Lung Tumor Segmentation in PET-CT Images

CVPR 2025poster

Lung cancer is a leading cause of cancer-related deaths globally. PET-CT is crucial for imaging lung tumors, providing essential metabolic and anatomical information, while it faces challenges such as poor image quality, motion artifacts, and complex tumor morphology. Deep learning-based models are…

2025

Decentralized Multi-robot Navigation Policy with Enhanced Security Using Graph GRU Policy Network

IROS 2025

Formulating a multi-robot obstacle avoidance policy is essential for enabling safe and efficient navigation in multi-robot environments, forming a critical component of the effective operation of multi-robot systems. Recently, reinforcement learning has been applied to improve the performance of dec

Cited by 0SourceScholar
2025

Dual-Mode Passive Fault-Tolerant Control for Underwater Vehicles with Actuator Faults and Time-Varying Disturbances

IROS 2025

This paper investigates the control problem of underwater vehicles subject to time-varying external disturbances and actuator faults. A novel passive fault-tolerant control (PFTC) scheme is developed to address the coupled disturbance-fault dynamics inherent in underwater vehicle systems. The propos

Cited by 0SourceScholar
2025

Efficient Multimodal 3D Object Detector via Instance-Level Contrastive Distillation

IROS 2025

Multimodal 3D object detectors leverage the strengths of both geometry-aware LiDAR point clouds and semantically rich RGB images to enhance detection performance. However, the inherent heterogeneity between these modalities, including unbalanced convergence and modal misalignment, poses significant

Cited by 1SourcecodeScholar
2025

FFBGNet: Full-Flow Bidirectional Feature Fusion Grasp Detection Network Based on Hybrid Architecture

RA-L 2025

Effectively integrating the complementary information from RGB-D images presents a significant challenge in robotic grasping. In this letter, we propose a full-flow bidirectional feature fusion grasp detection network (FFBGNet) based on a hybrid architecture to generate accurate grasp poses from RGB

Cited by 3SourceScholar
2025

Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection

CVPR 2025poster

Tiny object detection remains challenging in spite of the success of generic detectors. The dramatic performance degradation of generic detectors on tiny objects is mainly due to the the weak representations of extremely limited pixels. To address this issue, we propose a plug-and-play architecture…

Cited by 0SourcePDFScholar
2025

Foggy-Aware Teacher: An Unsupervised Domain Adaptive Learning Framework for Object Detection in Foggy Scenes

RA-L 2025

Unsupervised domain adaptation (UDA) is an effective scheme to improve the performance of an object detector in foggy scenes by adapting labeled normal images (source domain) to unlabeled foggy images (target domain). Existing methods leverage the Teacher-Student mutual learning framework, <italic x

Cited by 0SourcecodeScholar
2025

GeoScene: Temporal 3D Semantic Scene Completion with Geometric Correlation between Images

IROS 2025

Semantic Scene Completion (SSC) aims to reconstruct the entire 3D scene in terms of both occupancy and semantics, serving as a fundamental task for autonomous driving and robotic systems. Camera-based methods have seen significant advancements due to their low cost and rich visual cues. However, pre

Cited by 0SourceScholar
2025

Hierarchical Gaussian Mixture Model Splatting for Efficient and Part Controllable 3D Generation

CVPR 2025poster

3D content creation has achieved significant progress in terms of both quality and speed. Although current Gaussian Splatting-based methods can produce 3D objects within seconds, they are still limited by complex preprocessing or low controllability. In this paper, we introduce a novel framework des…

Cited by 0SourcePDFScholar
2025

Joint Optimization of Multi-Agent Task Allocation and Path Planning for Continuous Pickup and Delivery Tasks

IROS 2025

The multi-agent pickup and delivery problem is central to coordinating multiple agents in real-world applications such as warehouse automation, urban logistics, and robotic delivery networks, where efficient task assignment and pathfinding are vital for maximizing production efficiency. However, exi

Cited by 0SourceScholar
2025

MRMT-PR: A Multi-Scale Reverse-View Mamba-Transformer for LiDAR Place Recognition

IROS 2025

Place recognition is a fundamental technology of high relevance for autonomous robot navigation. Existing methods encounter significant challenges arising from scene variations (e.g., illumination changes, dynamic objects), view-point shifts, and difficulties in data fusion and alignment. These fact

Cited by 1SourceScholar
2025

Multi-Keypoint Affordance Representation for Functional Dexterous Grasping

RA-L 2025

Functional dexterous grasping requires precise hand-object interaction, going beyond simple gripping. Existing affordance-based methods primarily predict coarse interaction regions and cannot directly constrain the grasping posture, leading to a disconnection between visual perception and manipulati

Cited by 3SourcecodeScholar
2025

Multi-range Adaptive Perception Transformer for Iterative Homography Estimation

ICASSP 2025accepted

Homography estimation is fundamental to various vision tasks. Iteration-based methods have recently achieved significant success in this field. However, errors introduced during iterations can lead to increased image deformation. Existing methods often focus on capturing local correspondences in the…

Cited by 0SourceScholar
2025

Optimization Based Human-Guided Variable-Stiffness Visual Impedance Control for Contact-Rich Tasks

IROS 2025

In contact-rich tasks such as polishing and drilling, inevitable physical interactions often lead to task deviations due to interference, typically resulting in excessive contact forces and eventual task failure. To tackle these challenges, we propose an innovative human-guided visual-impedance cont

Cited by 0SourceScholar
2025

Parameterized Motion Planning for Aerial Manipulators in Contact with Unstructured Surfaces

IROS 2025

Motion planning for continuous contact-based aerial manipulators on complex unstructured surfaces remains a substantial challenge due to the sophisticated topology of unstructured surfaces. While direct planning in the high-dimensional configuration space manifolds faces efficiency limitations, simp

Cited by 0SourceScholar
2025

Partially Matching Submap Helps: Uncertainty Modeling and Propagation for Text to Point Cloud Localization

ICCV 2025poster

Text to point cloud cross-modal localization is a crucial vision-language task for future human-robot collaboration. Existing coarse-to-fine frameworks assume that each query text precisely corresponds to the center area of a submap, limiting their applicability in real-world scenarios. This work re…

2025

Prompt-driven Transferable Adversarial Attack on Person Re-Identification with Attribute-aware Textual Inversion

ICCV 2025poster

Person re-identification (re-id) models are vital in security surveillance systems, requiring transferable adversarial attacks to explore the vulnerabilities of them. Recently, vision-language models (VLM) based attacks have shown superior transferability by attacking generalized image and textual f…

2025

Safety-Aware Geometric Force-Impedance Control for Manipulators

IROS 2025

Since its inception, impedance control has emerged as a fundamental framework for robotic interaction control. Recent advancements in geometric impedance control have demonstrated certain advantages over traditional Cartesian impedance control. However, existing geometric impedance control approache

Cited by 0SourceScholar
2025

Searching Efficient Semantic Segmentation Architectures via Dynamic Path Selection

NeurIPS 2025poster

Existing NAS methods for semantic segmentation typically apply uniform optimization to all candidate networks (paths) within a one-shot supernet. However, the concurrent existence of both promising and suboptimal paths often results in inefficient weight updates and gradient conflicts. This issue is…

Cited by 0SourceScholar
2025

Semantic Ambiguity Modeling and Propagation for Fine-Grained Visual Cross View Geo-Localization

AAAI 2025technical

Visual cross view geo-localization is generally approached within a joint retrieval-and-calibration framework. However, existing methods overlook semantic ambiguities arising from query and reference images characterized by low overlap, dynamic foregrounds, viewpoint changes, and perceptual aliasing…

2025

WLuav: An Air-Ground Robot with High Ground Adaptability and Trajectory Tracking Performance

IROS 2025

Air-ground robots have received more and more attention and applications due to their air-to-ground motion performance and excellent energy efficiency. However, airground robots have many gaps including complex structure mechanisms, low terrain adaptability and low-precision controllers to significa

Cited by 0SourceScholar
2024

A Fast and Accurate Visual Inertial Odometry Using Hybrid Point-Line Features

RA-L 2024

Mainstream visual-inertial SLAM systems use point features for motion estimation and localization. However, point features do not perform well in scenes such as weak texture and motion blur. Therefore, the introduction of line features has received a lot of attention. In this letter, we propose a po

Cited by 6SourceScholar
2024

Decentralized Multi-Robot Navigation Coupled with Spatial-Temporal RetNet Based on Deep Reinforcement Learning

IROS 2024poster

Navigating robots through dynamic multi-robot environments, avoiding collisions with both other robots and obstacles, has emerged as a central challenge in robotics. The existing approaches fall short in allowing the policy network to effectively capture spatial-temporal reciprocal collision avoidan…

Cited by 0SourceScholar
2024

Decentralized Trajectory Planning for Formation Flight in Unknown and Dense Environments

IROS 2024poster

For aerial swarms, formation flight has been applied in various scenes. However, most existing works do not consider balancing the conflicting requirements among keeping formation, keeping the smoothness of trajectories, and obstacle avoidance within the limited time. To address this issue, we propo…

Cited by 0SourceScholar
2024

Domain Adaptation in Visual Reinforcement Learning via Self-Expert Imitation with Purifying Latent Feature

IROS 2024poster

Generalizing visual reinforcement learning is fundamental to robot visual navigation, involving the acquisition of a policy from interactions with source environments to facilitate adaptation to analogous, yet unfamiliar target environments. Recent advancements capitalize on data augmentation techni…

Cited by 0SourceScholar
2024

External Knowledge Enhanced 3D Scene Generation from Sketch

ECCV 2024poster

"Generating realistic 3D scenes is challenging due to the complexity of room layouts and object geometries. We propose a sketch based knowledge enhanced diffusion architecture (SEK) for generating customized, diverse, and plausible 3D scenes. SEK conditions the denoising process with a hand-drawn sk…

Cited by 6SourcePDFScholar
2024

L4D-Track: Language-to-4D Modeling Towards 6-DoF Tracking and Shape Reconstruction in 3D Point Cloud Stream

CVPR 2024poster

3D visual language multi-modal modeling plays an important role in actual human-computer interaction. However the inaccessibility of large-scale 3D-language pairs restricts their applicability in real-world scenarios. In this paper we aim to handle a real-time multi-task for 6-DoF pose tracking of u…

Cited by 0SourcePDFScholar
2023

Sketch and Text Guided Diffusion Model for Colored Point Cloud Generation

ICCV 2023poster

Diffusion probabilistic models have achieved remarkable success in text guided image generation. However, generating 3D shapes is still challenging due to the lack of sufficient data containing 3D models along with their descriptions. Moreover, text based descriptions of 3D shapes are inherently amb…

Cited by 32PDFScholar
2023

VDBblox: Accurate and Efficient Distance Fields for Path Planning and Mesh Reconstruction

IROS 2023poster

Highly accurate and efficient map in unknown and complex environments is essential for robotics navigation. Traditionally, mobile robot platforms are often computationally constrained when using multiple sensors to process large amounts of input data. In previous works, some of them have been deploy…

Cited by 3SourcecodeScholar
2022

ICK-Track: A Category-Level 6-DoF Pose Tracker Using Inter-Frame Consistent Keypoints for Aerial Manipulation

IROS 2022poster

Robots that are supposed to interact with or manipulate objects in the world must be able to track the poses of objects in their sensor data. Thus, Detecting and tracking the 6-DoF poses of targeted objects is important for aerial manipulation and is still in the early stage due to the high dynamics…

Cited by 8SourcecodeScholar
2022

Learning From Pixel-Level Noisy Label: A New Perspective for Light Field Saliency Detection

CVPR 2022poster

Saliency detection with light field images is becoming attractive given the abundant cues available, however, this comes at the expense of large-scale pixel level annotated data which is expensive to generate. In this paper, we propose to learn light field saliency from pixel-level noisy labels obta…

Cited by 25PDFcodeScholar
2022

Negative Stiffness Analysis and Regulation of In-Hand Manipulation with Underactuated Compliant Hands

ICRA 2022poster

This paper addresses the generation mechanism and avoidance method of negative stiffness during in-Hand manipulation with underactuated compliant hands. Firstly, a planar hand with two three-jointed fingers manipulating a rectangular is set, and a quasi-static underactuated operation model is establ…

Cited by 0SourceScholar
2021

FlowDriveNet: An End-to-End Network for Learning Driving Policies from Image Optical Flow and LiDAR Point Flow

ICRA 2021poster

Learning driving policies using an end-to-end network has been proved a promising solution for autonomous driving. Due to the lack of a benchmark driver behavior dataset that contains both the visual and the LiDAR data, existing works solely focus on learning driving from visual sensors. Besides, mo…

Cited by 4SourceScholar
2021

Free-Form Description Guided 3D Visual Graph Network for Object Grounding in Point Cloud

ICCV 2021poster

3D object grounding aims to locate the most relevant target object in a raw point cloud scene based on a free-form language description. Understanding complex and diverse descriptions, and lifting them directly to a point cloud is a new and challenging topic due to the irregular and sparse nature of…

Cited by 100PDFcodeScholar
2021

Lane-free Autonomous Intersection Management: A Batch-processing Framework Integrating Reservation-based and Planning-based Methods

ICRA 2021poster

Autonomous intersection management (AIM) refers to planning the trajectories for multiple connected and automated vehicles (CAVs) when they traverse an unsignalized intersection cooperatively. As an extension of the conventional AIM, lane-free AIM allows the CAVs to adjust their velocities and paths…

Cited by 23SourceScholar
2021

Multi-Expert Adversarial Attack Detection in Person Re-Identification Using Context Inconsistency

ICCV 2021poster

The success of deep neural networks (DNNs) has promoted the widespread applications of person re-identification (ReID). However, ReID systems inherit the vulnerability of DNNs to malicious attacks of visually inconspicuous adversarial perturbations. Detection of adversarial attacks is, therefore, a…

Cited by 44PDFScholar
2020

Design and Analysis of a Synergy-Inspired Three-Fingered Hand

ICRA 2020poster

Hand synergy from neuroscience provides an effective tool for anthropomorphic hands to realize versatile grasping with simple planning and control. This paper aims to extend the synergy-inspired design from anthropomorphic hands to multi-fingered robot hands. The synergy-inspired hands are not neces…

Cited by 10SourceScholar
2018

3D Face Reconstruction from Light Field Images: A Model-free Approach

ECCV 2018poster

Reconstructing 3D facial geometry from a single RGB image has recently instigated wide research interest. However, it is still an ill-posed problem and most methods rely on prior models hence undermining the accuracy of the recovered 3D faces. In this paper, we exploit the Epipolar Plane Images (EPI…

Cited by 50SourcePDFScholar
2015

Minimum sweeping area motion planning for flexible serpentine surgical manipulator with kinematic constraints

IROS 2015poster

Flexible serpentine manipulators are widely used in surgical robots as it can be operated inside the patient's body cavity by backbone bending. However, during the bending the manipulator sweeps over a region, where sensitive organs may locate. This raises the safety concern. In this paper, a motion…

Cited by 13SourceScholar