← Search

Xinru Zhang

4 accepted papers

2026

Native-Domain Cross-Attention for Camera-LiDAR Extrinsic Calibration Under Large Initial Perturbations

RA-L 2026

Accurate camera–LiDAR fusion relies on precise extrinsic calibration, which fundamentally depends on establishing reliable cross-modal correspondences under potentially large misalignments. Existing learning-based methods typically project LiDAR points into depth maps for feature fusion, which disto

Cited by 0SourceScholar
2025

Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysis

EMNLP 2025

Multimodal Sentiment Analysis (MSA) is the task of understanding human emotions by analyzing a combination of different data sources, such as text, audio, and visual inputs. Although recent advances have improved emotion modeling across modalities, existing methods still struggle with two fundamenta

2025

Task-Specific Information Decomposition for End-to-End Dense Video Captioning

ACL 2025long

Dense video captioning aims to localize events within input videos and generate concise descriptive texts for each event. Advanced end-to-end methods require both tasks to share the same intermediate features that serve as event queries, thereby enabling the mutual promotion of two tasks. However, r…