← Search

Zeyu Dong

6 accepted papers

2026

IBMA: Information Bottleneck-Based Multimodal Alignment

ICML 2026poster

Multimodal learning aims to integrate information from heterogeneous data sources to improve representation quality and downstream task performance. A key challenge lies in aligning modality-specific representations while suppressing modality-dependent noise and redundancy. The Information Bottlenec…

Cited by 0SourceScholar
2026

SepPrune: Structured Pruning for Efficient Deep Speech Separation

AAAI 2026technical

Although deep learning has substantially advanced speech separation in recent years, most existing studies continue to prioritize separation quality while overlooking computational efficiency, an essential factor for low-latency speech processing in real-time applications. In this paper, we propose

Cited by 0SourcePDFScholar
2025

Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting

ICCV 2025poster

Spatiotemporal forecasting tasks, such as traffic flow, combustion dynamics, and weather forecasting, often require complex models that suffer from low training efficiency and high memory consumption. This paper proposes a lightweight framework, Spectral Decoupled Knowledge Distillation, which trans…

2025

Prototype-Driven Multi-Feature Generation for Visible-Infrared Person Re-identification

ICASSP 2025accepted

The primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by varying viewpoints and irregular movements. Existing methods of…

Cited by 0SourceScholar
2024

Generalizing End-To-End Autonomous Driving In Real-World Environments Using Zero-Shot LLMs

CoRL 2024poster

Traditional autonomous driving methods adopt modular design, decomposing tasks into sub-tasks, including perception, prediction, planning, and control. In contrast, end-to-end autonomous driving directly outputs actions from raw sensor data, avoiding error accumulation. However, training an end-to-e…

Cited by 5SourceScholar