← Search

Chi Xie

6 accepted papers

2026

FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection

AAAI 2026technical

The remarkable realism of images generated by diffusion models poses critical detection challenges. Current methods utilize reconstruction error as a discriminative feature, exploiting the observation that real images exhibit higher reconstruction errors when processed through diffusion models. Howe

Cited by 0SourcePDFScholar
2025

Enhancing Document Understanding with Group Position Embedding: A Novel Approach to Incorporate Layout Information

ICLR 2025poster

Recent advancements in document understanding have been dominated by leveraging large language models (LLMs) and multimodal large models. However, enabling LLMs to comprehend complex document layouts and structural information often necessitates intricate network modifications or costly pre-training…

2025

Multi-Modal Object Re-identification via Sparse Mixture-of-Experts

ICML 2025poster

We present MFRNet, a novel network for multi-modal object re-identification that integrates multi-modal data features to effectively retrieve specific objects across different modalities. Current methods suffer from two principal limitations: (1) insufficient interaction between pixel-level semantic…

Cited by 0SourcePDFScholar
2023

Category Query Learning for Human-Object Interaction Classification

CVPR 2023poster

Unlike most previous HOI methods that focus on learning better human-object features, we propose a novel and complementary approach called category query learning. Such queries are explicitly associated to interaction categories, converted to image specific category representation via a transformer…

2023

Described Object Detection: Liberating Object Detection with Flexible Expressions

NeurIPS 2023poster

Detecting objects based on language information is a popular task that includes Open-Vocabulary object Detection (OVD) and Referring Expression Comprehension (REC). In this paper, we advance them to a more practical setting called *Described Object Detection* (DOD) by expanding category names to fle…