← Search

Jianquan Liu

4 accepted papers

2026

Object-Centric Framework for Video Moment Retrieval

AAAI 2026technical

Most existing video moment retrieval methods rely on temporal sequences of frame- or clip-level features that primarily encode global visual and semantic information. However, such representations often fail to capture fine-grained object semantics and appearance, which are crucial for localizing mo

Cited by 0SourcePDFScholar
2026

SPEAKING CLEARLY: A SIMPLIFIED WHISPER-BASED CODEC FOR LOW-BITRATE SPEECH CODING

ICASSP 2026poster

Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic codecs with complex semantic supervision. We explore the oppos…

Cited by 0SourcePDFScholar
2024

Visual Pivoting Unsupervised Multimodal Machine Translation in Low-Resource Distant Language Pairs

EMNLP 2024finding

Unsupervised multimodal machine translation (UMMT) aims to leverage vision information as a pivot between two languages to achieve better performance on low-resource language pairs. However, there is presently a challenge: how to handle alignment between distant language pairs (DLPs) in UMMT. To thi…

2021

COOPNet: Multi-Modal Cooperative Gender Prediction in Social Media User Profiling

ICASSP 2021accepted

The principal way of performing user profiling is to investigate accumulated social media data. However, the problem of information asymmetry generally exists in user generated contents since users post multi-modal contents in social media freely. In this paper, we propose a novel text-image coopera…

Cited by 0SourceScholar