← Search

Chao Sun

11 accepted papers

2026

AESF-LIO: Adaptive Error-State Fusion LiDAR-Inertial Odometry for Ground Vehicles in Structured Environments

RA-L 2026

In LiDAR-based Simultaneous Localization and Mapping (SLAM) systems for vehicles, the point-to-plane Iterative Closest Point (ICP) method is widely used for scan matching. This approach incorporates all planar points into a single objective function for optimization, yet does not explicitly distingu

Cited by 0SourceScholar
2026

Bootstrap Your Own AV-Proxies: Adaptive Contrastive and Prototype Learning for Audio-Visual Segmentation

CVPR 2026

Audio-Visual Segmentation (AVS) aims to accurately segment sounding objects in video frames by leveraging audio-visual correspondence cues. However, it remains challenging due to the intrinsic semantic incompleteness within a single modality and the semantic gap between audio and visual representati

Cited by 0SourceScholar
2026

OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination

AAAI 2026technical

Recently, Omni-modal large language models (OLLMs) have sparked a new wave of research, achieving impressive results in tasks such as audio-video understanding and real-time environment perception. However, hallucination issues still persist. Similar to the bimodal setting, the priors from the text

Cited by 0SourcePDFScholar
2025

Conditional Representation Learning for Customized Tasks

NeurIPS 2025spotlight

Conventional representation learning methods learn a universal representation that primarily captures dominant semantics, which may not always align with customized downstream tasks. For instance, in animal habitat analysis, researchers prioritize scene-related features, whereas universal embeddings…

Cited by 0SourcecodeScholar
2025

Enhancing Speech Emotion Recognition with Speech Dynamic Modeling and Multi-Modal Knowledge Distillation

ICASSP 2025accepted

Complementary semantic information from the text modality, obtained through runtime transcription, plays a crucial role in Speech Emotion Recognition (SER). However, it introduces additional computational overhead and potential errors. To address these issues, we propose the SDMMKD framework, which…

Cited by 0SourceScholar
2025

Integrating Expert Knowledge and Traffic Data for Lane-Changing Intention Prediction in Autonomous Vehicles

RA-L 2025

Accurate vehicle intention prediction is critical for autonomous driving safety in complex traffic environments. To address the interpretability limitations of data-driven methods while maintaining high accuracy, this letter proposes a knowledge-data co-learning framework featuring: (1) a knowledge-

Cited by 0SourceScholar
2025

Online Identification of Equivalent Roll Center for Underwater Gliders to Weaken the Repeated Yawing

RA-L 2025

Underwater gliders (UG) experience unexpected yaw during each dormancy stage of the controller, necessitating additional adjustments, which manifests as repeated yawing throughout the entire profiling time. Repeated yawing is frequently observed during the long-term deployment of the UGs, which can

Cited by 1SourceScholar
2024

SCE-LIO: An Enhanced LiDAR Inertial Odometry by Constructing Submap Constraints

RA-L 2024

In LiDAR-based Simultaneous Localization and Mapping (SLAM) systems, loop closure detection is crucial for enhancing the accuracy of odometry. However, constraints from loop closure detection are only provided when a loop is detected and can only enhance odometry accuracy at specific moments. Theref

Cited by 3SourceScholar
2023

Automatic Targetless Calibration for LiDAR and Camera Based on Instance Segmentation

RA-L 2023

In autonomous vehicles, accurate extrinsic calibration for LiDAR and camera is an essential prerequisite for multi-sensor information fusion. Automatic and targetless extrinsic calibration has become the mainstream of academic research in recent years. However, existing automatic calibration methods

Cited by 20SourceScholar
2023

MABe22: A Multi-Species Multi-Task Benchmark for Learned Representations of Behavior

ICML 2023poster

We introduce MABe22, a large-scale, multi-agent video and trajectory benchmark to assess the quality of learned behavior representations. This dataset is collected from a variety of biology experiments, and includes triplets of interacting mice (4.7 million frames video+pose tracking data, 10 millio…