← Search

Xirong Li

20 accepted papers

2026

SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval

CVPR 2026

For video-text retrieval, the use of CLIP has been a de facto standard. However, as CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have been made to reintroduce audio -- typically by inco

Cited by 0SourcecodeScholar
2025

D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX Matching

AAAI 2025technical

Videos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decorati…

Cited by 0SourcePDFScholar
2025

Hybrid-Tower: Fine-grained Pseudo-query Interaction and Generation for Text-to-Video Retrieval

ICCV 2025accepted

The Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers fr…

Cited by 0SourcePDFScholar
2025

Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization

ACL 2025finding

Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications. Recent works have attempted to apply Direct Preference Optimization (DPO) to enhance the performance of MLLMs, but have shown inconsistent improvements in mitigating hallucinations. To addre…

Cited by 0SourcePDFScholar
2025

Multi-Object Sketch Animation by Scene Decomposition and Motion Planning

ICCV 2025poster

Sketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in mult…

Cited by 0SourcePDFScholar
2025

PhD: A ChatGPT-Prompted Visual Hallucination Evaluation Dataset

CVPR 2025highlight

Multimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE). This paper contributes a ChatGPT-Prompted visual hallucination evaluation Dataset (PhD) for objective VHE at a large scale. The essence of VHE is to ask an MLLM questions ab…

2024

Cliprerank: An Extremely Simple Method For Improving Ad-Hoc Video Search

ICASSP 2024accepted

Ad-hoc Video Search (AVS) enables users to search for unlabeled video content using on-the-fly textual queries. Current deep learning-based models for AVS are trained to optimize holistic similarity between short videos and their associated descriptions. However, due to the diversity of ad-hoc queri…

Cited by 0SourceScholar
2024

Holistic Features are almost Sufficient for Text-to-Video Retrieval

CVPR 2024poster

For text-to-video retrieval (T2VR) which aims to retrieve unlabeled videos by ad-hoc textual queries CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature…

2024

Tackling Long Code Search with Splitting, Encoding, and Aggregating

COLING 2024main

Code search with natural language helps us reuse existing code snippets. Thanks to the Transformer-based pretraining models, the performance of code search has been improved significantly. However, due to the quadratic complexity of multi-head self-attention, there is a limit on the input token leng…

2023

SAFL-Net: Semantic-Agnostic Feature Learning Network with Auxiliary Plugins for Image Manipulation Detection

ICCV 2023poster

Since image editing methods in real world scenarios cannot be exhausted, generalization is a core challenge for image manipulation detection, which could be severely weakened by semantically related features. In this paper we propose SAFL-Net, which constrains a feature extractor to learn semantic-a…

Cited by 35PDFScholar
2022

Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval

ECCV 2022poster

"In this paper we revisit feature fusion, an old-fashioned topic, in the new context of text-to-video retrieval. Different from previous research that considers feature fusion only at one end, let it be video or text, we aim for feature fusion for both ends within a unified framework. We hypothesize…

2022

Semi-Supervised Keypoint Detector and Descriptor for Retinal Image Matching

ECCV 2022poster

"For retinal image matching (RIM), we propose SuperRetina, the first end-to-end method with jointly trainable keypoint detector and descriptor. SuperRetina is trained in a novel semi-supervised manner. A small set of (nearly 100) images are incompletely labeled and used to supervise the network to d…

2021

Article Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked Claims

ACL 2021long

False claims that have been previously fact-checked can still spread on social media. To mitigate their continual spread, detecting previously fact-checked claims is indispensable. Given a claim, existing works focus on providing evidence for detection by reranking candidate fact-checking articles (…

2021

Image Manipulation Detection by Multi-View Multi-Scale Supervision

ICCV 2021poster

The key challenge of image manipulation detection is how to learn generalizable features that are sensitive to manipulations in novel data, whilst specific to prevent false alarms on authentic images. Current research emphasizes the sensitivity, with the specificity overlooked. In this paper we addr…

Cited by 231PDFcodeScholar
2019

Dual Encoding for Zero-Example Video Retrieval

CVPR 2019poster

This paper attacks the challenging problem of zero-example video retrieval. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described in natural language text with no visual example provided. Given videos as sequences of frames and queries as sequences of wo…

Cited by 324PDFcodeScholar
2015

Detecting semantic concepts in consumer videos using audio

ICASSP 2015accepted

With the increasing use of audio sensors in user generated content collection, how to detect semantic concepts using audio streams has become an important research problem. In this paper, we present a semantic concept annotation system using soundtracks/ audio of the video. We investigate three diff…

Cited by 0SourceScholar