← Search

Zhen Sun

13 accepted papers

2026

6DAttack: Backdoor Attacks in the 6DoF Pose Estimation

AAAI 2026technical

Recent advances in deep learning have enabled highly accurate six-degree-of-freedom (6DoF) object pose estimation, leading to its widespread use in real-world applications such as robotics, augmented reality, virtual reality, and autonomous systems. However, backdoor attacks pose a major security ri

Cited by 2SourcePDFScholar
2026

JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) integrate the audio modality directly into the model, rather than converting speech into text and inputting text to Large Language Models (LLMs). While jailbreak attacks on LLMs have been extensively studied, the security of LALMs with audio modalities remains lar…

Cited by 0SourcecodeScholar
2026

SFCo-Nav: Efficient Zero-Shot Visual Language Navigation Via Collaboration of Slow LLM and Fast Attributed Graph Alignment

ICRA 2026poster

Recent advances in large vision-language models (VLMs) and large language models (LLMs) have enabled zero-shot approaches to Visual Language Navigation (VLN), where an agent follows natural language instructions using only ego perception and reasoning. However, existing zero‑shot methods typically c…

2025

A2I-Calib: An Anti-Noise Active Multi-IMU Spatial-Temporal Calibration Framework for Legged Robots

IROS 2025

Recently, multi-node inertial measurement unit (IMU)-based odometry for legged robots has gained attention due to its cost-effectiveness, power efficiency, and high accuracy. However, the spatial and temporal misalignment between foot-end motion derived from forward kinematics and foot IMU measureme

Cited by 1SourceScholar
2025

Are We in the AI-Generated Text World Already? Quantifying and Monitoring AIGT on Social Media

ACL 2025long

Social media platforms are experiencing a growing presence of AI-Generated Texts (AIGTs). However, the misuse of AIGTs could have profound implications for public opinion, such as spreading misinformation and manipulating narratives. Despite its importance, it remains unclear how prevalent AIGTs are…

2025

CHASM: Unveiling Covert Advertisements on Chinese Social Media

NeurIPS 2025poster

Current benchmarks for evaluating large language models (LLMs) in social media moderation completely overlook a serious threat: covert advertisements, which disguise themselves as regular posts to deceive and mislead consumers into making purchases, leading to significant ethical and legal concerns.…

Cited by 0SourceScholar
2025

DS-VLM: Diffusion Supervision Vision Language Model

ICML 2025poster

Vision-Language Models (VLMs) face two critical limitations in visual representation learning: degraded supervision due to information loss during gradient propagation, and the inherent semantic sparsity of textual supervision compared to visual data. We propose the Diffusion Supervision Vision-Lang…

Cited by 0SourcePDFScholar
2025

FC-Attack: Jailbreaking Multimodal Large Language Models via Auto-Generated Flowcharts

EMNLP 2025

Multimodal Large Language Models (MLLMs) have become powerful and widely adopted in some practical applications.However, recent research has revealed their vulnerability to multimodal jailbreak attacks, whereby the model can be induced to generate harmful content, leading to safety risks. Although m

2025

FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification

ICML 2025poster

Multimodal person re-identification (Re-ID) aims to match pedestrian images across different modalities. However, most existing methods focus on limited cross-modal settings and fail to support arbitrary query-retrieval combinations, hindering practical deployment. We propose FlexiReID, a flexible f…

Cited by 0SourcePDFScholar
2025

IMOST: Incremental Memory Mechanism with Online Self-Supervision for Continual Traversability Learning

ICRA 2025

Traversability estimation is the foundation of path planning for a general navigation system. However, complex and dynamic environments pose challenges for the latest methods using self-supervised learning (SSL) technique. Firstly, existing SSL-based methods generate sparse annotations lacking detai

Cited by 7SourcecodeScholar
2025

THE-SEAN: A Heart Rate Variation-Inspired Temporally High-Order Event-Based Visual Odometry with Self-Supervised Spiking Event Accumulation Networks

IROS 2025

Event-based visual odometry has recently gained attention for its high accuracy and real-time performance in fast-motion systems. Unlike traditional synchronous estimators that rely on constant-frequency (zero-order) triggers, event-based visual odometry can actively accumulate information to genera

Cited by 2SourceScholar
2023

NeRF-LOAM: Neural Implicit Representation for Large-Scale Incremental LiDAR Odometry and Mapping

ICCV 2023poster

Simultaneously odometry and mapping using LiDAR data is an important task for mobile systems to achieve full autonomy in large-scale environments. However, most existing LiDAR-based methods prioritize tracking quality over reconstruction quality. Although the recently developed neural radiance field…

Cited by 75PDFcodeScholar