← Search

Fan Shi

39 accepted papers

2026

Class-Guided Network with Rare-Class Amplification for Sea State Estimation Based on Ship Motion Data

ICRA 2026poster

Accurate, real-time Sea State Estimation (SSE) is crucial for the safety and operational efficiency of Autonomous Surface Vessels (ASVs). However, existing deep learning methods for this task commonly face three major challenges: the inherent class imbalance of marine environments, the ambiguous bou…

Cited by 0Scholar
2026

CodeMamba: Shifting from Target Semantics to Self-Supervised Background Manifold Learning for Singularity Detection in Infrared Sequences

ICML 2026poster

Multi-frame infrared small target detection suffers from extreme semantic paucity of targets and representation collapse due to overwhelming class imbalance, resulting in the persistent inability to accurately distinguish point-like targets from dynamic background clutter. To address these issues, w…

Cited by 0SourceScholar
2026

Decomposition of Concept-Level Rules in Visual Scenes

ICLR 2026poster

Human cognition is compositional, and one can parse a visual scene into independent concepts and the corresponding concept-changing rules. By contrast, many vision-language systems process images holistically, with limited support for explicit decomposition. And previous methods of decomposing conce…

Cited by 0SourceScholar
2026

Envision, Attend, Then Respond: Counterfactual Hallucination Mitigation in Large Vision-Language Models

CVPR 2026

Large Vision-Language Models (LVLMs) often hallucinate when visual evidence conflicts with world knowledge, i.e., in counterfactual scenarios. We propose Envision-Attend-Respond (EnAR), a training-free framework that leverages visual priors to steer the model's attention toward counterfactual elemen

Cited by 0SourcecodeScholar
2026

E²I-VRWKV: Explicit EPI-Representation and Interaction-Aware Vision-RWKV for Light Field Semantic Segmentation

ICML 2026poster

Pixel-level semantic segmentation of 4D light field (LF) data remains a considerable challenge, primarily due to the conflict between modeling complex spatial-angular dependencies and maintaining linear computational efficiency. Current linear models like VRWKV offer scalability but often fail to ca…

Cited by 0SourceScholar
2026

Few-Shot Neural Differentiable Simulator: Real-To-Sim Rigid-Contact Modeling

ICRA 2026poster

Accurate physics simulation is essential for robotic learning and control, yet analytical simulators often fail to capture complex contact dynamics, while learning-based simulators typically require large amounts of costly real-world data. To bridge this gap, we propose a few-shot real-to-sim approa…

2026

REBot: Reflexive Evasion Robot for Instantaneous Dynamic Obstacle Avoidance

RA-L 2026

Dynamic obstacle avoidance (DOA) is critical for quadrupedal robots operating in environments with moving obstacles or humans. Existing approaches typically rely on navigation-based trajectory replanning, which assumes sufficient reaction time and leading to fails when obstacles approach rapidly. In

Cited by 4SourcecodeScholar
2026

REBot: Reflexive Evasion Robot for Instantaneous Dynamic Obstacle Avoidance

ICRA 2026poster

Dynamic obstacle avoidance (DOA) is critical for quadrupedal robots operating in environments with moving obstacles or humans. Existing approaches typically rely on navigation-based trajectory replanning, which assumes sufficient reaction time and leading to fails when obstacles approach rapidly. In…

2026

SCRWKV: Ultra-Compact Structure-Calibrated Vision-RWKV for Topological Crack Segmentation

ICML 2026poster

Achieving pixel-level accurate segmentation of structural cracks across diverse scenarios remains a formidable challenge. Existing methods face significant bottlenecks in balancing crack topology modeling with computational efficiency, often failing to reconcile high segmentation quality with low re…

Cited by 0SourceScholar
2026

Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

CVPR 2026

Unaligned RGB-T salient object detection (SOD) remains challenging due to severe cross-modal spatial discrepancies and unreliable feature fusion. Existing methods often assume perfect alignment or rely on geometric registration, which is computationally demanding and sensitive to cross-modal inconsi

Cited by 0SourceScholar
2025

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning

ICML 2025poster

Abstract visual reasoning (AVR) enables humans to quickly discover and generalize abstract rules to new scenarios. Designing intelligent systems with human-like AVR abilities has been a long-standing topic in the artificial intelligence community. Deep AVR solvers have recently achieved remarkable s…

2025

Collaborative Association Network for Multi-view Multi-Human Association and Tracking using Constraint Optimization and Object Search

ICASSP 2025accepted

Multi-view multi-human association and tracking (MvMHAT) enhances scene perception using multiple cameras, crucial for applications such as surveillance and crowd analysis. Inherent feature disparities between views complicate similarity calculations. Recent works combine representation and motion i…

Cited by 0SourceScholar
2025

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation

ICCV 2025poster

Creating high-quality, generalizable speech-driven 3D talking heads remains a persistent challenge. Previous methods achieve satisfactory results for fixed viewpoints and small-scale audio variations, but they struggle with large head rotations and out-of-distribution (OOD) audio. Moreover, they are…

Cited by 0SourcePDFScholar
2025

GO-Flock: Goal-Oriented Flocking in 3D Unknown Environments with Depth Maps

IROS 2025

Artificial Potential Field (APF) methods are widely used for reactive flocking control, but they often suffer from challenges such as deadlocks and local minima, especially in the presence of obstacles. Existing solutions to address these issues are typically passive, leading to slow and inefficient

Cited by 0SourceScholar
2025

Learning Quiet Walking for a Small Home Robot

ICRA 2025

As home robotics gains traction, robots are increasingly integrated into households, offering companionship and assistance. Quadruped robots, particularly those resembling dogs, have emerged as popular alternatives for traditional pets. However, user feedback highlights concerns about the noise thes

Cited by 5SourceScholar
2025

Multi-Scale Convolutional Networks with Class-Normalized Logit Clipping for Robust Sea State Estimation from Noisy Ship Motion Data

ICRA 2025

Autonomous ships utilize automation systems to achieve unmanned navigation, driving innovation in maritime transportation. However, sea conditions, influenced by dynamic factors such as wave height, wind speed, and ocean currents, present a challenge in accurately assessing these conditions. Traditi

Cited by 0SourceScholar
2025

Residual Policy Learning for Perceptive Quadruped Control Using Differentiable Simulation

ICRA 2025

First-order Policy Gradient (FoPG) algorithms such as Backpropagation through Time and Analytical Policy Gradients leverage local simulation physics to accelerate policy search, significantly improving sample efficiency in robot control compared to standard model-free reinforcement learning. However

Cited by 16SourceScholar
2025

SCSegamba: Lightweight Structure-Aware Vision Mamba for Crack Segmentation in Structures

CVPR 2025poster

Pixel-level segmentation of structural cracks across various scenarios remains a considerable challenge. Current methods encounter challenges in effectively modeling crack morphology and texture, facing challenges in balancing segmentation quality with low computational resource usage. To overcome t…

2025

Serial Local Patterns and Irregular Dependencies Extract and Cascaded Fusion Network for Structural Crack Segmentation

ICASSP 2025accepted

Achieving pixel-level crack segmentation in complex scenarios is a major challenge, as current methods have difficulty effectively integrating both local features and irregular pixel dependencies. In this paper, we introduce a Cascaded Fusion Network (LICFN) specifically designed for crack segmentat…

Cited by 0SourceScholar
2024

HumanMimic: Learning Natural Locomotion and Transitions for Humanoid Robot via Wasserstein Adversarial Imitation

ICRA 2024poster

Transferring human motion skills to humanoid robots remains a significant challenge. In this study, we introduce a Wasserstein adversarial imitation learning system, allowing humanoid robots to replicate natural whole-body locomotion patterns and execute seamless transitions by mimicking human motio…

Cited by 33SourceScholar
2024

Rethinking Robustness Assessment: Adversarial Attacks on Learning-based Quadrupedal Locomotion Controllers

RSS 2024poster

Legged locomotion has recently achieved remarkable success with the progress of machine learning techniques, especially deep reinforcement learning (RL). Controllers employing neural networks have demonstrated empirical and qualitative robustness against real-world uncertainties, including sensor no…

Cited by 21SourcePDFScholar
2024

Towards Generative Abstract Reasoning: Completing Raven’s Progressive Matrix via Rule Abstraction and Selection

ICLR 2024poster

Endowing machines with abstract reasoning ability has been a long-term research topic in artificial intelligence. Raven's Progressive Matrix (RPM) is widely used to probe abstract visual reasoning in machine intelligence, where models will analyze the underlying rules and select one image from candi…

2022

Learning Agile Hybrid Whole-body Motor Skills for Thruster-Aided Humanoid Robots

IROS 2022poster

Humanoid robots are versatile platforms with the potential for multiple locomotion skills. However, this contact-switched system with only two contact feet is fragile to keep balance in many scenarios. Inspired by birds combining legs and wings, we propose the novel hybrid locomotion behavior for th…

Cited by 5SourceScholar
2021

Circus ANYmal: A Quadruped Learning Dexterous Manipulation with Its Limbs

ICRA 2021poster

Quadrupedal robots are skillful at locomotion tasks while lacking manipulation skills, not to mention dexterous manipulation abilities. Inspired by the animal behavior and the duality between multi-legged locomotion and multi-fingered manipulation, we showcase a circus ball challenge on a quadrupeda…

Cited by 58SourceScholar
2020

Aerial Regrasping: Pivoting with Transformable Multilink Aerial Robot

ICRA 2020poster

Regrasping is one of the most common and important manipulation skills used in our daily life. However, aerial regrasping has not been seriously investigated yet, since most of the aerial manipulator lacks dexterous manipulation abilities except for the basic pick-and-place. In this paper, we focus…

Cited by 38SourceScholar
2020

Model Reference Adaptive Control of Multirotor for Missions with Dynamic Change of Payloads During Flight

ICRA 2020poster

Carrying payloads in air is a major mission for multirotor aerial robot. However, the presence of payloads on multirotor aerial robot has a risk of degrading the performance of the flight controller. This concern becomes obvious especially when carrying objects not securely attached to the body or p…

Cited by 15SourceScholar
2020

Online Motion Planning for Deforming Maneuvering and Manipulation by Multilinked Aerial Robot Based on Differential Kinematics

RA-L 2020

State-of-the-art work on deformable multirotor aerial robots has developed a strong maneuvering ability in such robots, whereas there is no versatile aerial robot that can perform both deforming maneuvering and aerial manipulation yet. However, a novel multilinked aerial robot presented in our previ

Cited by 21SourceScholar
2020

Stable Control in Climbing and Descending Flight under Upper Walls using Ceiling Effect Model based on Aerodynamics

ICRA 2020poster

Stable flight control under ceilings is difficult for multirotor Unmanned Aerial Vehicles (UAVs). The wake interaction between rotors and upper walls, called the "ceiling effect", causes an increase of rotor thrust. As a result of the thrust increase, multi-rotors are drawn upward abruptly and colli…

Cited by 18SourceScholar
2019

Achievement of Online Agile Manipulation Task for Aerial Transformable Multilink Robot

IROS 2019poster

Transformable aerial robots are favorable in aerial manipulation tasks for their flexible ability to change configuration during the flight. By assuming robot keeping in the mild motion, the previous researches sacrifice aerial agility to simplify the complex non-linear system into a single rigid bo…

Cited by 6SourceScholar
2019

Design, Modeling and Control of Fully Actuated 2D Transformable Aerial Robot with 1 DoF Thrust Vectorable Link Module

IROS 2019poster

We present a novel transformable multilinked aerial robot which consists of link modules with 1 DoF thrust vectoring mechanism. Commonly used UAV is underactuated due to its simplicity and high flight duration, but can not control the position and orientation independently. To overcome this problem,…

Cited by 33SourceScholar
2019

External Wrench Estimation for Multilink Aerial Robot by Center of Mass Estimator Based on Distributed IMU System

ICRA 2019poster

External wrench estimation is very helpful for aerial exploration and manipulation tasks. During the exploration, there might be unseen obstacles to cause dangerous collisions. The estimation of the external force and torque is also beneficial in aerial manipulation tasks. In this paper, we present…

Cited by 20SourceScholar
2018

Aerial Grasping Based on Shape Adaptive Transformation by HALO: Horizontal Plane Transformable Aerial Robot with Closed-Loop Multilinks Structure

ICRA 2018poster

In this paper, we present the achievement of aerial grasping by shape adaptive transformation to the object shape, using a novel transformable aerial robot called HALO: Horizontal Plane Transformable Aerial Robot with Closed-loop Multilinks Structure. Aerial manipulation is an active research area a…

Cited by 41SourceScholar
2018

Design, Modeling, and Control of an Aerial Robot DRAGON: A Dual-Rotor-Embedded Multilink Robot With the Ability of Multi-Degree-of-Freedom Aerial Transformation

RA-L 2018

In this letter, we introduce a novel transformable aerial robot called DRAGON, which is a dual-rotor-embedded multilink robot with the ability of multi-degree-of-freedom (DoF) aerial transformation. The new aerial robot can control the full pose in SE(3) regarding the center of gravity (CoG) of mult

Cited by 160SourceScholar
2018

Flight Motion of Passing Through Small Opening by DRAGON: Transformable Multilinked Aerial Robot

IROS 2018poster

In this paper, we introduce the achievement of the flight motion to pass through small opening by the multilinked and transformable aerial robot. Previous works about such motion are based on under-actuated multirotors, indicating that aggressive maneuvering is necessary condition. This involves two…

Cited by 21SourceScholar
2017

Multilinked multirotor with internal communication system for multiple objects transportation based on form optimization method

IROS 2017poster

In this paper, we show the achievement of a transformable aerial robot with internal communication system for multiple objects transportation. As it is not easy to make the flight endurance of an aerial robot longer, we study the problem to transport multiple objects at the same time to improve the…

Cited by 33SourceScholar
2017

Robust real-time visual tracking using dual-frame deep comparison network integrated with correlation filters

IROS 2017poster

In recent years, applications of visual tracking algorithms has seen a substantial growth with deployments in intelligent robots such as drones for human tracking. The algorithms for such tasks has to be efficient in terms of computational cost while been robust, accurate and fast. Object tracking a…

Cited by 21SourceScholar