← Search

Zhangbin Li

3 accepted papers

2026

Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

CVPR 2026

Point-level weakly-supervised temporal sentiment localization (P-WTSL) aims to detect sentiment-relevant segments in untrimmed multimodal videos using timestamp sentiment annotations, which greatly reduces the costly frame-level labeling. To tackle the intrinsic challenges of imprecise sentiment bou

Cited by 0SourcecodeScholar
2025

Patch-level Sounding Object Tracking for Audio-Visual Question Answering

AAAI 2025technical

Answering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PS…

Cited by 6SourcePDFScholar
2024

Object-Aware Adaptive-Positivity Learning for Audio-Visual Question Answering

AAAI 2024technical

This paper focuses on the Audio-Visual Question Answering (AVQA) task that aims to answer questions derived from untrimmed audible videos. To generate accurate answers, an AVQA model is expected to find the most informative audio-visual clues relevant to the given questions. In this paper, we propos…