← Search

S. Joe Qin

5 accepted papers

2026

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

ICML 2026poster

Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down a…

Cited by 0SourceScholar
2025

CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering

EMNLP 2025

Users often assume that large language models (LLMs) share their cognitive alignment of context and intent, leading them to omit critical information in question-answering (QA) and produce ambiguous queries. Responses based on misaligned assumptions may be perceived as hallucinations. Therefore, ide