← Search

Xuelin Zhu

6 accepted papers

2026

Shot-Conditioned Vision-Language Adaptation for Effective Harmful Content Detection from Online Short Videos

IJCAI 2026

Short video harmful content detection aims to automatically identify diverse anomalies from user-generated media. This task presents unique challenges due to frequent editing cuts and highly variable anomaly densities, limiting the effectiveness of traditional surveillance-based approaches. Moreover

Cited by 0Scholar
2025

Denoise-then-Retrieve: Text-Conditioned Video Denoising for Video Moment Retrieval

IJCAI 2025

Current text-driven Video Moment Retrieval (VMR) methods encode all video clips, including irrelevant ones, disrupting multimodal alignment and hindering optimization. To this end, we propose a denoise-then-retrieve paradigm that explicitly filters text-irrelevant clips from videos and then retrieve

Cited by 0SourcePDFScholar
2025

MSCI: Addressing CLIP's Inherent Limitations for Compositional Zero-Shot Learning

IJCAI 2025

Compositional Zero-Shot Learning (CZSL) aims to recognize unseen state-object combinations by leveraging known combinations. Existing studies basically rely on the cross-modal alignment capabilities of CLIP but tend to overlook its limitations in capturing fine-grained local features, which arise fr

2025

MambaML: Exploring State Space Models for Multi-Label Image Classification

ICCV 2025poster

Mamba, a selective state-space model, has recently seen widespread application across various visual tasks due to its exceptional ability to capture long-range dependencies. While promising results have been demonstrated in image classification, its potential in multi-label image classification rema…

Cited by 0SourcePDFScholar
2023

Exploring Visual Pre-training for Robot Manipulation: Datasets, Models and Methods

IROS 2023poster

Visual pre-training with large-scale real-world data has made great progress in recent years, showing great potential in robot learning with pixel observations. However, the recipes of visual pre-training for robot manipulation tasks are yet to be built. In this paper, we thoroughly investigate the…

Cited by 16SourcecodeScholar
2023

Scene-Aware Label Graph Learning for Multi-Label Image Classification

ICCV 2023poster

Multi-label image classification refers to assigning a set of labels for an image. One of the main challenges of this task is how to effectively capture the correlation among labels. Existing studies on this issue mostly rely on the statistical label co-occurrence or semantic similarity of labels. H…

Cited by 31PDFScholar