← Search

Guangyao Li

11 accepted papers

2026

AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs

ICML 2026poster

Recent advances in Omni-Multimodal Large Language Models (Omni-MLLMs) have enabled strong integration of vision, audio, and language. However, their audio-visual intelligence (AVI) remains insufficiently evaluated due to the lack of systematic and comprehensive benchmarks. We introduce AVI-Bench, a …

Cited by 0SourceScholar
2026

SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking

AAAI 2026technical

Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in

Cited by 0SourcePDFScholar
2025

Audio-Visual Instance Segmentation

CVPR 2025poster

In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, contai…

2025

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

CVPR 2025poster

In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal localization, spatial localization, spatio-temporal reasoning, and pixel-level understanding. Instead, human possesses a unif…

2025

Language Decoupling with Fine-grained Knowledge Guidance for Referring Multi-object Tracking

ICCV 2025poster

Referring Multi-Object Tracking (RMOT) aims to detect and track specific objects based on natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, often failing to exploit fine-grained linguistic cues that are crucial for distinguishing objects with…

2025

PEDE: Enhance Multi-modal Sarcasm Detection in Videos via Prompted Emotion Distributions

ICASSP 2025accepted

Multi-modal sarcasm detection is crucial for understanding human communications. A key aspect of multi-modal sarcasm detection is the analysis of emotion incongruity. However, the advancement of emotion analysis in video is hindered by the scarcity of labeled datasets, which are limited in both scal…

Cited by 0SourceScholar
2024

CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing

ICASSP 2024accepted

Audio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities.…

Cited by 0SourceScholar
2024

Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

AAAI 2024technical

Never having seen an object and heard its sound simultaneously, can the model still accurately localize its visual position from the input audio? In this work, we concentrate on the Audio-Visual Localization and Segmentation tasks but under the demanding zero-shot and few-shot scenarios. To achieve…

2024

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

ECCV 2024poster

"Traditional reference segmentation tasks have predominantly focused on silent visual scenes, neglecting the integral role of multimodal perception and interaction in human experiences. In this work, we introduce a novel task called Reference Audio-Visual Segmentation (Ref-AVS), which seeks to segme…

2023

Enabling Abductive Learning to Exploit Knowledge Graph

IJCAI 2023poster

Most systems integrating data-driven machine learning with knowledge-driven reasoning usually rely on a specifically designed knowledge base to enable efficient symbolic inference. However, it could be cumbersome for the nonexpert end-users to prepare such a knowledge base in real tasks. Recent year…

2022

Learning To Answer Questions in Dynamic Audio-Visual Scenarios

CVPR 2022oral

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal understanding and spatio-temporal reasoning over audio-visual scenes.…

Cited by 157PDFcodeScholar