← Search

Yu Gao

23 accepted papers

2026

DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online Map Diffusion

RA-L 2026

Autonomous driving requires accurate scene understanding, including road geometry, traffic agents, and their semantic relationships. In online HD map generation scenarios, raster-based representations are well-suited to vision models but lack geometric precision, while graph-based representations re

Cited by 1SourcecodeScholar
2026

Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow

CVPR 2026

High-quality 3D scene representation in radiance fields relies on accurate camera poses which are often difficult to acquire in real-world scenarios. An effective solution is to use RGB images for the joint optimization of radiance fields and camera poses, an approach that has been well explored in

Cited by 0SourcecodeScholar
2026

FilterGS: Traversal-Free Parallel Filtering and Adaptive Shrinking for Large-Scale LoD 3D Gaussian Splatting

CVPR 2026

3D Gaussian Splatting has revolutionized neural rendering with real-time performance. However, scaling this approach to large scenes using Level-of-Detail methods faces critical challenges: inefficient serial traversal consuming over 60% of rendering time, and redundant Gaussian-tile pairs that incu

Cited by 0SourcecodeScholar
2026

HKAFER: Achieve Visual Parameter-Efficient Fine-Tuning via Heterogeneous Kronecker Adaptation for Facial Expression Recognition

AAAI 2026technical

Facial Expression Recognition (FER) seeks to classify affective states from facial images, which remains a challenging problem due to variations in real-world conditions. FER task becomes particularly complex when handling unconstrained environments characterized by partial occlusions, different hea

Cited by 0SourcePDFScholar
2026

LaDy: Lagrangian-Dynamic Informed Network for Skeleton-based Action Segmentation via Spatial-Temporal Modulation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) aims to densely parse untrimmed skeletal sequences into frame-level action categories. However, existing methods, while proficient at capturing spatio-temporal kinematics, neglect the underlying physical dynamics that govern human motion. This overs

Cited by 0SourcecodeScholar
2026

Leveraging Verifier-Based Reinforcement Learning in Image Editing

CVPR 2026

While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing remains largely unexplored. A key bottleneck is the lack of a robust general reward model for all editing tasks. Existing edit reward models usually gi

Cited by 0SourcecodeScholar
2026

Spectral Scalpel: Amplifying Adjacent Action Discrepancy via Frequency-Selective Filtering for Skeleton-Based Action Segmentation

CVPR 2026

Skeleton-based Temporal Action Segmentation (STAS) seeks to densely segment and classify diverse actions within long, untrimmed skeletal motion sequences. However, existing STAS methodologies face challenges of limited inter-class discriminability and blurred segmentation boundaries, primarily due t

Cited by 0SourcecodeScholar
2026

UniUncer: Unified Dynamic–Static Uncertainty for End-To-End Driving

ICRA 2026poster

End-to-end (E2E) driving has become a cornerstone of both industry deployment and academic research, offering a single learnable pipeline that maps multi-sensor inputs to actions while avoiding hand-engineered modules. However, the reliability of such pipelines strongly depends on how well they hand…

2026

Unified Map Prior Encoder for Mapping and Planning

ICRA 2026poster

Online mapping and end-to-end (E2E) planning in autonomous driving are still largely sensor-centric, leaving rich map priors—HD/SD vector maps, rasterized SD maps, and satellite imagery—underused due to heterogeneity, pose drift, and inconsistent availability at test time. We present emph{UMPE}, a U…

2025

Automated 3D-GS Registration and Fusion via Skeleton Alignment and Gaussian-Adaptive Features

IROS 2025

In recent years, 3D Gaussian Splatting (3D-GS)based scene representation demonstrates significant potential in real-time rendering and training efficiency. However, most existing methods primarily focus on single-map reconstruction, while the registration and fusion of multiple 3D-GS submaps remain

Cited by 2SourceScholar
2025

Diffusion Augmentation Sub-center Modeling for Unsupervised Anomalous Sound Detection with Partially Attribute-Unavailable Conditions

ICASSP 2025accepted

Current state-of-the-art unsupervised anomalous sound detection (ASD) methods typically rely on manually annotated attribute information as labels, employing auxiliary classification tasks to learn an embedding space for normal sounds, which helps detect anomalies deviating from this space. However,…

Cited by 0SourceScholar
2025

DroneSplat: 3D Gaussian Splatting for Robust 3D Reconstruction from In-the-Wild Drone Imagery

CVPR 2025highlight

Drones have become essential tools for reconstructing wild scenes due to their outstanding maneuverability. Recent advances in radiance field methods have achieved remarkable rendering quality, providing a new avenue for 3D reconstruction from drone imagery. However, dynamic distractors in wild env…

Cited by 2SourcePDFScholar
2025

GaussianGraph: 3D Gaussian-Based Scene Graph Generation for Open-World Scene Understanding

IROS 2025

Recent advancements in 3D Gaussian Splatting(3DGS) have significantly improved semantic scene understanding, enabling natural language queries to localize objects within a scene. However, existing methods primarily focus on embedding compressed CLIP features to 3D Gaussians, suffering from low objec

Cited by 6SourcecodeScholar
2025

OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting for Refined Object-Level Understanding

IROS 2025

Recent advancements in 3D scene understanding have made significant strides in enabling interaction with scenes using open-vocabulary queries, particularly for VR/AR and robotic applications. Nevertheless, existing methods are hindered by rigid offline pipelines and the inability to provide precise

Cited by 4SourcecodeScholar
2025

OpenGS-SLAM: Open-Set Dense Semantic SLAM with 3D Gaussian Splatting for Object-Level Scene Understanding

ICRA 2025

Recent advancements in 3D Gaussian Splatting have significantly improved the efficiency and quality of dense semantic SLAM. However, previous methods are generally constrained by limited-category pre-trained classifiers and implicit semantic representation, which hinder their performance in open-set

Cited by 15SourcecodeScholar
2025

SparseMeXt: Unlocking the Potential of Sparse Representations for HD Map Construction

IROS 2025

Recent advancements in high-definition (HD) map construction have demonstrated the effectiveness of dense representations, which heavily rely on computationally intensive bird’s-eye view (BEV) features. While sparse representations offer a more efficient alternative by avoiding dense BEV processing,

Cited by 4SourceScholar
2025

VOILA: Complexity-Aware Universal Segmentation of CT Images by Voxel Interacting with Language

AAAI 2025technical

Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of vision-language methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant im…

2024

Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition

ICASSP 2024accepted

The mainstream paradigm of speech emotion recognition (SER) is identifying the single emotion label of the entire utterance. This line of works neglect the emotion dynamics at fine temporal granularity and mostly fail to leverage linguistic information of speech signal explicitly. In this paper, we…

Cited by 0SourceScholar
2024

Fine-tuning the Diffusion Model and Distilling Informative Priors for Sparse-view 3D Reconstruction

IROS 2024poster

3D reconstruction methods such as Neural Radiance Fields (NeRFs) are capable of optimizing high-quality 3D representation from images. However, NeRF is limited by the requirement for a large number of multi-view images, making its application to real-world scenarios challenging. In this work, we pro…

Cited by 0SourcecodeScholar
2024

GroupTrack: Multi-Object Tracking by Using Group Motion Patterns

IROS 2024poster

The main challenge of Multi-Object Tracking (MOT) lies in maintaining a distinctive identity for each target in dense crowds or occluded scenarios. Although the existing methods have achieved significantly progress by using robust object detectors or complex association strategies, they cannot effec…

Cited by 0SourceScholar
2024

MLPER: Multi-Level Prompts for Adaptively Enhancing Vision-Language Emotion Recognition

IROS 2024poster

In the field of robotics, vision-based Emotion Recognition (ER) has achieved significant progress, but it still faces the challenge of poor generalization ability under unconstrained conditions (e.g., occlusions and pose variations). In this work, we propose MLPER model, which introduces Vision-Lang…

Cited by 1SourceScholar
2022

Fisheye object detection based on standard image datasets with 24-points regression strategy

IROS 2022poster

Fisheye object detection is a difficult task in robotics and autonomous driving. One of the reasons is that the fisheye datasets are inferior to standard image datasets in scale and quantity, which inspires the idea of using standard image datasets for fisheye object detection. However, the models t…

Cited by 4SourcecodeScholar
2021

TOOD: Task-Aligned One-Stage Object Detection

ICCV 2021poster

One-stage object detection is commonly implemented by optimizing two sub-tasks: object classification and localization, using heads with two parallel branches, which might lead to a certain level of spatial misalignment in predictions between the two tasks. In this work, we propose a Task-aligned On…

Cited by 1116PDFcodeScholar