← Search

Kailun Yang

44 accepted papers

2026

CoBEVMoE: Heterogeneity-Aware Feature Fusion with Dynamic Mixture-Of-Experts for Collaborative Perception

ICRA 2026poster

Collaborative perception aims to extend sensing coverage and improve perception accuracy by sharing information among multiple agents. However, due to differences in viewpoints and spatial positions, agents often acquire heterogeneous observations. Existing intermediate fusion methods primarily focu…

2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

ICLR 2026poster

Despite substantial progress in video understanding, most existing datasets are limited to Earth’s gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-rob…

Cited by 0SourcecodeScholar
2026

Hallucinating 360°: Panoramic Street-View Generation Via Local Scenes Diffusion and Probabilistic Prompting

ICRA 2026poster

Panoramic perception holds significant potential for autonomous driving, enabling vehicles to acquire a comprehensive 360° surround view in a single shot. However, autonomous driving is a data-driven task. Complete panoramic data acquisition requires complex sampling systems and annotation pipelines…

2026

Learning Latent Transmission and Glare Maps for Lens Veiling Glare Removal

CVPR 2026

Beyond the commonly recognized optical aberrations, the imaging performance of simplified optical systems--including single-lens and metalens designs--is often further degraded by veiling glare caused by stray-light scattering from non-ideal optical surfaces and coatings, particularly in complex rea

Cited by 0SourcecodeScholar
2026

MICA: Multi-Agent Industrial Coordination Assistant

ICRA 2026poster

Industrial workflows demand adaptive and trustworthy assistance that can operate under limited computing, connectivity, and strict privacy constraints. In this work, we present MICA (Multi-Agent Industrial Coordination Assistant), a perception-grounded and speech-interactive system that delivers rea…

2026

OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera

CVPR 2026

Robust 3D semantic occupancy is essential for legged and humanoid robots, yet most Semantic Scene Completion (SSC) systems are built for wheeled platforms with forward-facing sensors. We present OneOcc, a vision-only panoramic SSC framework tailored to severe body jitter and 360deg continuity. OneOc

Cited by 0SourcecodeScholar
2026

ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction

CVPR 2026

3D semantic occupancy prediction is central to autonomous driving, yet current methods are vulnerable to long-tailed class bias and out-of-distribution (OOD) inputs, often overconfidently assigning anomalies to rare classes. We present ProOOD, a lightweight, plug-and-play method that couples prototy

Cited by 0SourcecodeScholar
2026

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

CVPR 2026

Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic panoramas and OpenStreetMap (OSM). To this end, we establish a

Cited by 0SourcecodeScholar
2026

Segment-To-Act: Label-Noise-Robust Action-Prompted Video Segmentation towards Embodied Intelligence

ICRA 2026poster

Embodied intelligence relies on accurately segmenting objects actively involved in interactions. Action-based video object segmentation addresses this by linking segmentation with action semantics, but it depends on large-scale annotations and prompts that are costly, inconsistent, and prone to mult…

2026

Towards Universal Computational Aberration Correction in Photographic Cameras: A Comprehensive Benchmark Analysis

CVPR 2026

Prevalent Computational Aberration Correction (CAC) methods are typically tailored to specific optical systems, leading to poor generalization and labor-intensive re-training for new lenses.Developing CAC paradigms capable of generalizing across diverse photographic lenses offers a promising solutio

Cited by 0SourcecodeScholar
2026

UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands

RA-L 2026

Dexterous grasp datasets are vital for embodied intelligence, but mostly emphasize grasp stability, ignoring functional grasps needed for tasks like opening bottle caps or holding cup handles. Most rely on bulky, costly, and hard-to-control high-DOF Shadow Hands. Inspired by the human hand's underac

Cited by 1SourceScholar
2026

UniFucGrasp: Human-Hand-Inspired Unified Functional Grasp Annotation Strategy and Dataset for Diverse Dexterous Hands

ICRA 2026poster

Dexterous grasp datasets are vital for embodied intelligence, but mostly emphasize grasp stability, ignoring functional grasps needed for tasks like opening bottle caps or holding cup handles. Most rely on bulky, costly, and hard-to-control high-DOF Shadow Hands. Inspired by the human hand’s underac…

2025

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

NeurIPS 2025spotlight

Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenar…

Cited by 0SourcecodeScholar
2025

Multi-Keypoint Affordance Representation for Functional Dexterous Grasping

RA-L 2025

Functional dexterous grasping requires precise hand-object interaction, going beyond simple gripping. Existing affordance-based methods primarily predict coarse interaction regions and cannot directly constrain the grasping posture, leading to a disconnection between visual perception and manipulati

Cited by 3SourcecodeScholar
2025

One-Shot Affordance Grounding of Deformable Objects in Egocentric Organizing Scenes

IROS 2025

Deformable object manipulation in robotics presents significant challenges due to uncertainties in component properties, diverse configurations, visual interference, and ambiguous prompts. These factors complicate both perception and control tasks. To address these challenges, we propose a novel met

Cited by 2SourcecodeScholar
2025

QuaDreamer: Controllable Panoramic Video Generation for Quadruped Robots

CoRL 2025poster

Panoramic cameras, capturing comprehensive 360-degree environmental data, are suitable for quadruped robots in surrounding perception and interaction with complex environments. However, the scarcity of high-quality panoramic training data — caused by inherent kinematic constraints and complex sensor…

Cited by 0SourceScholar
2025

Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts

IROS 2025

Affordance refers to the functional properties that an agent perceives and utilizes from its environment, and is key perceptual information required for robots to perform actions. This information is rich and multimodal in nature. Existing multimodal affordance methods face limitations in extracting

Cited by 0SourcecodeScholar
2025

SF-TIM: A Simple Framework for Enhancing Quadrupedal Robot Jumping Agility by Combining Terrain Imagination and Measurement

IROS 2025

Dynamic jumping on high platforms and over gaps differentiates legged robots from wheeled counterparts. Compared to walking on rough terrains, dynamic locomotion on abrupt surfaces requires fusing proprioceptive and exteroceptive perception for explosive movements. In this paper, we propose SF-TIM (

Cited by 3SourcecodeScholar
2025

Scene-agnostic Pose Regression for Visual Localization

CVPR 2025poster

Absolute Pose Regression (APR) predicts 6D camera poses but lacks the adaptability to unknown environments without retraining, while Relative Pose Regression (RPR) generalizes better yet requires a large image retrieval database. Visual Odometry (VO) generalizes well in unseen environments but suffe…

Cited by 0SourcePDFScholar
2025

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

NeurIPS 2025poster

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange,…

Cited by 0SourcecodeScholar
2025

Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation

ICCV 2025poster

Panoramic image processing is essential for omni-context perception, yet faces constraints like distortions, perspective occlusions, and limited annotations. Previous unsupervised domain adaptation methods transfer knowledge from labeled pinhole data to unlabeled panoramic images, but they require a…

2025

Unveiling the Potential of Segment Anything Model 2 for RGB-Thermal Semantic Segmentation with Language Guidance

IROS 2025

The perception capability of robotic systems relies on the richness of the dataset. Although Segment Anything Model 2 (SAM2), trained on large datasets, demonstrates strong perception potential in perception tasks, its inherent training paradigm prevents it from being suitable for RGB-T tasks. To ad

Cited by 11SourcecodeScholar
2025

mmWalk: Towards Multi-modal Multi-view Walking Assistance

NeurIPS 2025poster

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV community, we build mmWalk, a simulated multi-modal dataset tha…

Cited by 0SourceScholar
2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

NeurIPS 2024poster

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and ac…

2024

Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision

ICASSP 2024accepted

Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation o…

Cited by 0SourceScholar
2024

Label-efficient Semantic Scene Completion with Scribble Annotations

IJCAI 2024poster

Semantic scene completion aims to infer the 3D geometric structures with semantic classes from camera or LiDAR, which provide essential occupancy information in autonomous driving. Prior endeavors concentrate on constructing the network or benchmark in a fully supervised manner. While the dense occu…

2024

MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments

ICRA 2024poster

People with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system,…

Cited by 13SourcecodeScholar
2024

Navigating Open Set Scenarios for Skeleton-Based Action Recognition

AAAI 2024technical

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and t…

2024

Open Panoramic Segmentation

ECCV 2024poster

"Panoramic images, capturing a 360° field of view (FoV), encompass omnidirectional spatial information crucial for scene understanding. However, it is not only costly to obtain training-sufficient dense-annotated panoramas but also application-restricted when training models in a close-vocabulary se…

2024

Referring Atomic Video Action Recognition

ECCV 2024poster

"We introduce a new task called Referring Atomic Video Action Recognition (RAVAR), aimed at identifying atomic actions of a particular person based on a textual description and the video data of this person. This task differs from traditional action recognition and localization, where predictions ar…

2024

Skeleton-Based Human Action Recognition with Noisy Labels

IROS 2024poster

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are o…

Cited by 5SourcecodeScholar
2023

Bi-Mapper: Holistic BEV Semantic Mapping for Autonomous Driving

RA-L 2023

A semantic map of the road scene, covering fundamental road elements, is an essential ingredient in autonomous driving systems. It provides important perception foundations for positioning and planning when rendered in the Bird's-Eye-View (BEV). Currently, the prior knowledge of hypothetical depth c

Cited by 24SourcecodeScholar
2023

Delivering Arbitrary-Modal Semantic Segmentation

CVPR 2023poster

Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the DeLiVER arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we…

2022

Bending Reality: Distortion-Aware Transformers for Adapting to Panoramic Semantic Segmentation

CVPR 2022poster

Panoramic images with their 360deg directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations a…

Cited by 107PDFcodeScholar
2022

Event-Based Fusion for Motion Deblurring with Cross-Modal Attention

ECCV 2022poster

"Traditional frame-based cameras inevitably suffer from motion blur due to long exposure times. As a kind of bio-inspired camera, the event camera records the intensity changes in an asynchronous way with high temporal resolution, providing valid image degradation information within the exposure tim…

2022

LF-VIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras with Negative Plane

IROS 2022poster

Visual-inertial-odometry has attracted extensive attention in the field of autonomous driving and robotics. The size of Field of View (FoV) plays an important role in Visual-Odometry (VO) and Visual-Inertial-Odometry (VIO), as a large FoV enables to perceive a wide range of surrounding scene element…

Cited by 21SourcecodeScholar
2022

TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration

IROS 2022poster

Traditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced D…

Cited by 39SourcecodeScholar
2021

Capturing Omni-Range Context for Omnidirectional Segmentation

CVPR 2021poster

Convolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsi…

Cited by 91PDFcodeScholar
2021

ISSAFE: Improving Semantic Segmentation in Accidents by Fusing Event-based Data

IROS 2021poster

Ensuring the safety of all traffic participants is a prerequisite for bringing intelligent vehicles closer to practical applications. The assistance system should not only achieve high accuracy under normal conditions, but obtain robust perception against extreme situations. However, traffic acciden…

Cited by 60SourcecodeScholar
2020

Real-Time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-Driving Images

RA-L 2020

Semantic segmentation has made striking progress due to the success of deep convolutional neural networks. Considering the demands of autonomous driving, real-time semantic segmentation has become a research hotspot these years. However, few real-time RGB-D fusion semantic segmentation studies are c

Cited by 161SourcecodeScholar