2026
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
ICML 2026poster
Audio-Visual Large Language Models (AV-LLMs) grapple with the prohibitive computational costs of processing massive, redundant audio and video tokens. Existing unimodal compression techniques fail to capture the heterogeneous and mutually influential information density of joint audio-visual signals…