← Search

Depeng Wang

2 accepted papers

2026

EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

ICML 2026poster

Audio-Visual Large Language Models (AV-LLMs) grapple with the prohibitive computational costs of processing massive, redundant audio and video tokens. Existing unimodal compression techniques fail to capture the heterogeneous and mutually influential information density of joint audio-visual signals…

Cited by 0SourceScholar
2025

Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge Editing

ICCV 2025poster

Multimodal Large Language Models (MLLMs) contain a substantial amount of factual knowledge, which may become outdated or inaccurate over time. Consequently, various knowledge editing techniques have been proposed to update the knowledge encoded within these models. Previous approaches maintain modal…

Cited by 0SourcePDFScholar