← Search

Pengcheng Zhao

7 accepted papers

2026

DVMM: A Dual-View Combination Descriptor for Multi-Modal LiDARs Online Place Recognition

ICRA 2026poster

Existing place recognition descriptors developed for single-agent SLAM struggle with multi-modal LiDAR differences in collaborative SLAM. To overcome this, we propose an online place recognition method for multi-modal LiDARs. This method introduces a dual-view combination descriptor, termed DVMM, by…

Cited by 0Scholar
2025

DFCA: Disentangled Feature Contrastive Learning and Augmentation for Fairer Dermatological Diagnostics

IJCAI 2025

With the increasing integration of AI in medical research and applications, the issue of fairness has become as critical as diagnostic accuracy. In dermatology diagnosis, the challenge of class-imbalanced data, which is sometimes limited and contains demographic attributes, results in an imbalanced

Cited by 0SourcePDFScholar
2025

Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

AAAI 2025technical

The Audio-Visual Video Parsing task aims to recognize and temporally localize all events occurring in either the audio or visual stream, or both. Capturing accurate event semantics for each audio/visual segment is vital. Prior works directly utilize the extracted holistic audio and visual features f…

2025

Text-Infused Audio-Visual Video Parsing with Semantic-Aware Multimodal Contrastive Learning

ICASSP 2025accepted

The Audio-Visual Video Parsing task aims to recognize events occurring in video segments for each modality. Presently, the excellent performance in handling video parsing is shown by generating pseudo labels at the segment level. However, these approaches still suffer from adequate semantic learning…

Cited by 0SourceScholar
2020

Leveraging the Template and Anchor Framework for Safe, Online Robotic Gait Design

ICRA 2020poster

Online control design using a high-fidelity, full-order model for a bipedal robot can be challenging due to the size of the state space of the model. A commonly adopted solution to overcome this challenge is to approximate the fullorder model (anchor) with a simplified, reduced-order model (template…

Cited by 16SourcecodeScholar
2020

Spectrogram Analysis Via Self-Attention for Realizing Cross-Model Visual-Audio Generation

ICASSP 2020accepted

Human cognition is supported by the combination of multimodal information from different sources of perception. The two most important modalities are visual and audio. Cross-modal visual-audio generation enables the synthesis of data from one modality following the acquisition of data from another.…

Cited by 0SourceScholar
2019

TerrainFusion: Real-time Digital Surface Model Reconstruction based on Monocular SLAM

IROS 2019poster

This paper presents an algorithm which can generate live digtial surface model (DSM) during the flight based on simultaneous localization and mapping (SLAM). We process the keyframe which is output by a monocular SLAM system to generate a local DSM, and fuse the local DSM to the global tiled DSM inc…

Cited by 19SourceScholar