← Search

Xiujun Shu

8 accepted papers

2026

All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning

ICLR 2026poster

The rapid proliferation of AI-generated images (AIGIs) highlights the pressing demand for generalizable detection methods. In this paper, we establish two key principles for AIGI detection task through systematic analysis: **(1) All Patches Matter**, since the uniform generation process ensures that…

Cited by 0SourceScholar
2026

MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models

ICML 2026poster

Recently, multimodal large language models (MLLMs) have been widely applied to reasoning tasks. However, they suffer from limited multi-rationale semantic modeling, insufficient logical robustness, and susceptibility to misleading cues. Therefore, we propose a Multi-rationale INtegrated Discriminati…

Cited by 0SourceScholar
2023

Collaborative Noisy Label Cleaner: Learning Scene-Aware Trailers for Multi-Modal Highlight Detection in Movies

CVPR 2023poster

Movie highlights stand out of the screenplay for efficient browsing and play a crucial role on social media platforms. Based on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations…

2023

D3G: Exploring Gaussian Prior for Temporal Sentence Grounding with Glance Annotation

ICCV 2023poster

Temporal sentence grounding (TSG) aims to locate a specific moment from an untrimmed video with a given natural language query. Recently, weakly supervised methods still have a large performance gap compared to fully supervised ones, while the latter requires laborious timestamp annotations. In this…

Cited by 15PDFcodeScholar
2023

NewsNet: A Novel Dataset for Hierarchical Temporal Segmentation

CVPR 2023poster

Temporal video segmentation is the get-to-go automatic video analysis, which decomposes a long-form video into smaller components for the following-up understanding tasks. Recent works have studied several levels of granularity to segment a video, such as shot, event, and scene. Those segmentations…

2023

Open-Vocabulary Multi-Label Classification via Multi-Modal Knowledge Transfer

AAAI 2023technical

Real-world recognition system often encounters the challenge of unseen labels. To identify such unseen labels, multi-label zero-shot learning (ML-ZSL) focuses on transferring knowledge by a pre-trained textual label embedding (e.g., GloVe). However, such methods only exploit single-modal knowledge f…

2022

Hyperspherical Learning in Multi-Label Classification

ECCV 2022poster

"Learning from online data with noisy web labels is gaining more attention due to the increasing cost of fully annotated datasets in large-scale multi-label classification tasks. Partial (positive) annotated data, as a particular case of data with noisy labels, are economically accessible. And they…

2021

Towards More Flexible and Accurate Object Tracking With Natural Language: Algorithms and Benchmark

CVPR 2021poster

Tracking by natural language specification is a new rising research topic that aims at locating the target object in the video sequence based on its language description. Compared with traditional bounding box (BBox) based tracking, this setting guides object tracking with high-level semantic inform…

Cited by 220PDFScholar