← Search

Xiangdong Wang

10 accepted papers

2026

Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models

AAAI 2026technical

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perce

Cited by 0SourcePDFScholar
2024

Audio Generation with Multiple Conditional Diffusion Model

AAAI 2024technical

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the controllability of existing pre-trained text-to-audio models…

2024

Semi-Supervised Sound Event Detection with Local and Global Consistency Regularization

ICASSP 2024accepted

Learning meaningful frame-wise features on a partially labeled dataset is crucial to semi-supervised sound event detection. Prior works either maintain consistency on frame-level predictions or seek feature-level similarity among neighboring frames, which cannot exploit the potential of unlabeled da…

Cited by 0SourceScholar
2023

Inferential Knowledge-Enhanced Integrated Reasoning for Video Question Answering

AAAI 2023technical

Recently, video question answering has attracted growing attention. It involves answering a question based on a fine-grained understanding of video multi-modal information. Most existing methods have successfully explored the deep understanding of visual modality. We argue that a deep understanding…

Cited by 1SourcePDFScholar
2022

Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering

NAACL 2022long

Existing video question answering (video QA) models lack the capacity for deep video understanding and flexible multistep reasoning. We propose for video QA a novel model which performs dynamic multistep reasoning between questions and videos. It creates video semantic representation based on the vi…

Cited by 12SourcePDFScholar
2022

Explainable Question Answering based on Semantic Graph by Global Differentiable Learning and Dynamic Adaptive Reasoning

EMNLP 2022main

Multi-hop Question Answering is an agent task for testing the reasoning ability. With the development of pre-trained models, the implicit reasoning ability has been surprisingly improved and can even surpass human performance. However, the nature of the black box hinders the construction of explaina…

Cited by 3SourcePDFScholar
2022

Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question Answering

COLING 2022main

Recently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by computing confidence among questions, options and evidence sentences, which is insuff…

2020

Guided Learning for Weakly-Labeled Semi-Supervised Sound Event Detection

ICASSP 2020accepted

We propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detection. Instead of designing a single model by considering a trade-off between the tw…

Cited by 0SourceScholar
2020

Multi-Branch Learning for Weakly-Labeled Sound Event Detection

ICASSP 2020accepted

There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SE…

Cited by 0SourceScholar