← Search

Kunyu Peng

32 accepted papers

2026

Exploring Single Domain Generalization of LiDAR-Based Semantic Segmentation under Imperfect Labels

ICRA 2026poster

Accurate perception is critical for vehicle safety, with LiDAR as a key enabler in autonomous driving. To ensure robust performance across environments, sensor types, and weather conditions without costly re-annotation, domain generalization in LiDAR-based 3D semantic segmentation is essential. Howe…

2026

Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments

ICLR 2026poster

Despite substantial progress in video understanding, most existing datasets are limited to Earth’s gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for real-world vision systems. This presents a challenge for domain-rob…

Cited by 0SourcecodeScholar
2026

Hallucinating 360°: Panoramic Street-View Generation Via Local Scenes Diffusion and Probabilistic Prompting

ICRA 2026poster

Panoramic perception holds significant potential for autonomous driving, enabling vehicles to acquire a comprehensive 360° surround view in a single shot. However, autonomous driving is a data-driven task. Complete panoramic data acquisition requires complex sampling systems and annotation pipelines…

2026

HybriDLA: Hybrid Generation for Document Layout Analysis

AAAI 2026technical

Conventional document layout analysis (DLA) traditionally depends on empirical priors or a fixed set of learnable queries executed in a single forward pass. While sufficient for early-generation documents with a small, predetermined number of regions, this paradigm struggles with contemporary docume

Cited by 0SourcePDFScholar
2026

MICA: Multi-Agent Industrial Coordination Assistant

ICRA 2026poster

Industrial workflows demand adaptive and trustworthy assistance that can operate under limited computing, connectivity, and strict privacy constraints. In this work, we present MICA (Multi-Agent Industrial Coordination Assistant), a perception-grounded and speech-interactive system that delivers rea…

2026

ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction

CVPR 2026

3D semantic occupancy prediction is central to autonomous driving, yet current methods are vulnerable to long-tailed class bias and out-of-distribution (OOD) inputs, often overconfidently assigning anomalies to rare classes. We present ProOOD, a lightweight, plug-and-play method that couples prototy

Cited by 0SourcecodeScholar
2026

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

CVPR 2026

Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In this work, instead of pinhole and satellite images, we study robust MCVGL using holistic panoramas and OpenStreetMap (OSM). To this end, we establish a

Cited by 0SourcecodeScholar
2026

Segment-To-Act: Label-Noise-Robust Action-Prompted Video Segmentation towards Embodied Intelligence

ICRA 2026poster

Embodied intelligence relies on accurately segmenting objects actively involved in interactions. Action-based video object segmentation addresses this by linking segmentation with action semantics, but it depends on large-scale annotations and prompts that are costly, inconsistent, and prone to mult…

2025

CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG Decoding

NeurIPS 2025spotlight

Understanding and decoding human brain activity from electroencephalography (EEG) signals is a fundamental problem in neuroscience and artificial intelligence, with applications ranging from cognition and emotion recognition to clinical diagnosis and brain–computer interfaces. While recent EEG found…

Cited by 0SourceScholar
2025

Dense Semantic Bird-Eye-View Map Generation from Sparse LiDAR Point Clouds via Distribution-aware Feature Fusion

IROS 2025

Semantic scene understanding in bird-eye view (BEV) plays a crucial role in autonomous driving. A common approach to generating BEV maps from LiDAR point-cloud data involves constructing a pillar-level representation by projecting 3D point clouds onto a 2D plane. This process partially discards spat

Cited by 0SourceScholar
2025

Graph-based Document Structure Analysis

ICLR 2025poster

When reading a document, glancing at the spatial layout of a document is an initial step to understand it roughly. Traditional document layout analysis (DLA) methods, however, offer only a superficial parsing of documents, focusing on basic instance detection and often failing to capture the nuanced…

Cited by 0SourcePDFScholar
2025

HopaDIFF: Holistic-Partial Aware Fourier Conditioned Diffusion for Referring Human Action Segmentation in Multi-Person Scenarios

NeurIPS 2025spotlight

Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set. Existing methods primarily address single-person activities with fixed action sequences, overlooking multi-person scenar…

Cited by 0SourcecodeScholar
2025

MindAligner: Explicit Brain Functional Alignment for Cross-Subject Visual Decoding from Limited fMRI Data

ICML 2025poster

Brain decoding aims to reconstruct visual perception of human subject from fMRI signals, which is crucial for understanding brain's perception mechanisms. Existing methods are confined to the single-subject paradigm due to substantial brain variability, which leads to weak generalization across ind…

Cited by 0SourcePDFScholar
2025

Situat3DChange: Situated 3D Change Understanding Dataset for Multimodal Large Language Model

NeurIPS 2025poster

Physical environments and circumstances are fundamentally dynamic, yet current 3D datasets and evaluation benchmarks tend to concentrate on either dynamic scenarios or dynamic situations in isolation, resulting in incomplete comprehension. To overcome these constraints, we introduce Situat3DChange,…

Cited by 0SourcecodeScholar
2025

Unlocking Constraints: Source-Free Occlusion-Aware Seamless Segmentation

ICCV 2025poster

Panoramic image processing is essential for omni-context perception, yet faces constraints like distortions, perspective occlusions, and limited annotations. Previous unsupervised domain adaptation methods transfer knowledge from labeled pinhole data to unlabeled panoramic images, but they require a…

2025

VISO-Grasp: Vision-Language Informed Spatial Object-centric 6-DoF Active View Planning and Grasping in Clutter and Invisibility

IROS 2025

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active view planning, our framework constructs and updates an insta

Cited by 13SourcecodeScholar
2024

Advancing Open-Set Domain Generalization Using Evidential Bi-Level Hardest Domain Scheduler

NeurIPS 2024poster

In Open-Set Domain Generalization (OSDG), the model is exposed to both new variations of data appearance (domains) and open-set conditions, where both known and novel categories are present at test time. The challenges of this task arise from the dual need to generalize across diverse domains and ac…

2024

Elevating Skeleton-Based Action Recognition with Efficient Multi-Modality Self-Supervision

ICASSP 2024accepted

Self-supervised representation learning for human action recognition has developed rapidly in recent years. Most of the existing works are based on skeleton data while using a multi-modality setup. These works overlooked the differences in performance among modalities, which led to the propagation o…

Cited by 0SourceScholar
2024

MateRobot: Material Recognition in Wearable Robotics for People with Visual Impairments

ICRA 2024poster

People with Visual Impairments (PVI) typically recognize objects through haptic perception. Knowing objects and materials before touching is desired by the target users but under-explored in the field of human-centered robotics. To fill this gap, in this work, a wearable vision-based robotic system,…

Cited by 13SourcecodeScholar
2024

Muscles in Time: Learning to Understand Human Motion In-Depth by Simulating Muscle Activations

NeurIPS 2024poster

Exploring the intricate dynamics between muscular and skeletal structures is pivotal for understanding human motion. This domain presents substantial challenges, primarily attributed to the intensive resources required for acquiring ground truth muscle activation data, resulting in a scarcity of dat…

Cited by 3SourcePDFScholar
2024

Navigating Open Set Scenarios for Skeleton-Based Action Recognition

AAAI 2024technical

In real-world scenarios, human actions often fall outside the distribution of training data, making it crucial for models to recognize known actions and reject unknown ones. However, using pure skeleton data in such open-set conditions poses challenges due to the lack of visual background cues and t…

2024

Occlusion-Aware Seamless Segmentation

ECCV 2024poster

"Panoramic images can broaden the Field of View (FoV), occlusion-aware prediction can deepen the understanding of the scene, and domain adaptation can transfer across viewing domains. In this work, we introduce a novel task, Occlusion-Aware Seamless Segmentation (OASS), which simultaneously tackles…

2024

Open Panoramic Segmentation

ECCV 2024poster

"Panoramic images, capturing a 360° field of view (FoV), encompass omnidirectional spatial information crucial for scene understanding. However, it is not only costly to obtain training-sufficient dense-annotated panoramas but also application-restricted when training models in a close-vocabulary se…

2024

Position: Quo Vadis, Unsupervised Time Series Anomaly Detection?

ICML 2024poster

The current state of machine learning scholarship in Timeseries Anomaly Detection (TAD) is plagued by the persistent use of flawed evaluation metrics, inconsistent benchmarking practices, and a lack of proper justification for the choices made in novel deep learning-based model designs. Our paper pr…

2024

RoDLA: Benchmarking the Robustness of Document Layout Analysis Models

CVPR 2024poster

Before developing a Document Layout Analysis (DLA) model in real-world applications conducting comprehensive robustness testing is essential. However the robustness of DLA models remains underexplored in the literature. To address this we are the first to introduce a robustness benchmark for DLA mod…

Cited by 6SourcePDFScholar
2024

Skeleton-Based Human Action Recognition with Noisy Labels

IROS 2024poster

Understanding human actions from body poses is critical for assistive robots sharing space with humans in order to make informed and safe decisions about the next interaction. However, precise temporal localization and annotation of activity sequences is time-consuming and the resulting labels are o…

Cited by 5SourcecodeScholar
2024

SynthAct: Towards Generalizable Human Action Recognition based on Synthetic Data

ICRA 2024poster

Synthetic data generation is a proven method for augmenting training sets without the need for extensive setups, yet its application in human activity recognition is underexplored. This is particularly crucial for human-robot collaboration in household settings, where data collection is often privac…

Cited by 3SourceScholar
2023

Delivering Arbitrary-Modal Semantic Segmentation

CVPR 2023poster

Multimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the DeLiVER arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we…

2023

Quantized Distillation: Optimizing Driver Activity Recognition Models for Resource-Constrained Environments

IROS 2023poster

Deep learning-based models are at the top of most driver observation benchmarks due to their remarkable accuracies but come with a high computational cost, while the resources are often limited in real-world driving scenarios. This paper presents a lightweight framework for resource- efficient drive…

Cited by 1SourcecodeScholar
2022

Bending Reality: Distortion-Aware Transformers for Adapting to Panoramic Semantic Segmentation

CVPR 2022poster

Panoramic images with their 360deg directional view encompass exhaustive information about the surrounding space, providing a rich foundation for scene understanding. To unfold this potential in the form of robust panoramic segmentation models, large quantities of expensive, pixel-wise annotations a…

Cited by 107PDFcodeScholar
2022

TransDARC: Transformer-based Driver Activity Recognition with Latent Space Feature Calibration

IROS 2022poster

Traditional video-based human activity recognition has experienced remarkable progress linked to the rise of deep learning, but this effect was slower as it comes to the downstream task of driver behavior understanding. Understanding the situation inside the vehicle cabin is essential for Advanced D…

Cited by 39SourcecodeScholar