← Search

Lei Tan

18 accepted papers

2026

Aggregating Diverse Cue Experts for AI-Generated Image Detection

AAAI 2026technical

The rapid emergence of image synthesis models poses challenges to the generalization of AI-generated image detectors. However, existing methods often rely on model-specific features, leading to overfitting and poor generalization. In this paper, we introduce the Multi-Cue Aggregation Network (MCAN),

Cited by 0SourcePDFScholar
2026

Bridging Day and Night: Target-Class Hallucination Suppression in Unpaired Image Translation

AAAI 2026technical

Day-to-night unpaired image translation is important to downstream tasks but remains challenging due to large appearance shifts and the lack of direct pixel-level supervision. Existing methods often introduce semantic hallucinations, where objects from target classes such as traffic signs and vehicl

Cited by 7SourcePDFScholar
2026

FIND: A Simple Yet Effective Baseline for Diffusion-Generated Image Detection

AAAI 2026technical

The remarkable realism of images generated by diffusion models poses critical detection challenges. Current methods utilize reconstruction error as a discriminative feature, exploiting the observation that real images exhibit higher reconstruction errors when processed through diffusion models. Howe

Cited by 0SourcePDFScholar
2026

FUSE: Frequency-domain Unification and Spectral Energy Alignment for Multi-modal Object Re-Identification

ICML 2026poster

Despite significant progress in multi-modal Re-Identification (ReID), existing methods tend to emphasize low-frequency cues. Consequently, they focus on attributes such as color, illumination, and coarse appearance, while overlooking mid- and high-frequency structures that encode geometric, textural…

Cited by 0SourceScholar
2026

Joint Implicit and Explicit Language Learning for Pedestrian Attribute Recognition

AAAI 2026technical

Pedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by converting attributes into language descriptions to facilitate interactive learning between attributes an

Cited by 0SourcePDFScholar
2025

Embedding Robust Watermarking into Pattern to Protect the Copyright of Ceramic Artifacts

AAAI 2025technical

Ceramic artworks with elegant patterns present enormous collectible value and profits. To claim the copyright, the builder usually pastes their conspicuous stamp on the bottom or side of the ceramic artworks, which inevitably affects the external image of the artwork. In addition, the stamp is weak…

Cited by 0SourcePDFScholar
2025

FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification

ICML 2025poster

Multimodal person re-identification (Re-ID) aims to match pedestrian images across different modalities. However, most existing methods focus on limited cross-modal settings and fail to support arbitrary query-retrieval combinations, hindering practical deployment. We propose FlexiReID, a flexible f…

Cited by 0SourcePDFScholar
2025

GSAlign: Geometric and Semantic Alignment Network for Aerial-Ground Person Re-Identification

NeurIPS 2025poster

Aerial-Ground person re-identification (AG-ReID) is an emerging yet challenging task that aims to match pedestrian images captured from drastically different viewpoints, typically from unmanned aerial vehicles (UAVs) and ground-based surveillance cameras. The task poses significant challenges due to…

Cited by 0SourceScholar
2025

MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification

NeurIPS 2025spotlight

The challenge of inconsistent modalities in real-world applications presents significant obstacles to effective object re-identification (ReID). However, most existing approaches assume modality-matched conditions, significantly limiting their effectiveness in modality-mismatched scenarios. To overc…

Cited by 0SourceScholar
2025

Multi-Modal Object Re-identification via Sparse Mixture-of-Experts

ICML 2025poster

We present MFRNet, a novel network for multi-modal object re-identification that integrates multi-modal data features to effectively retrieve specific objects across different modalities. Current methods suffer from two principal limitations: (1) insufficient interaction between pixel-level semantic…

Cited by 0SourcePDFScholar
2025

Physical Marker: Revealing Invisible Hyperlinks Hidden in Printed Trademarks

AAAI 2025technical

Embedding links in brand logos is a promising technology, which allows consumers to access the online information of products by capturing physical logo images. Previous physical data hiding methods primarily embed data within cover media in a global manner, making them ineffective for processing br…

Cited by 0SourcePDFScholar
2024

Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification

AAAI 2024technical

Occluded person re-identification (Re-ID) aims to address the potential occlusion problem when matching occluded or holistic pedestrians from different camera views. Many methods use the background as artificial occlusion and rely on attention networks to exclude noisy interference. However, the si…

Cited by 10SourcePDFScholar
2024

Occluded Person Re-identification via Saliency-Guided Patch Transfer

AAAI 2024technical

While generic person re-identification has made remarkable improvement in recent years, these methods are designed under the assumption that the entire body of the person is available. This assumption brings about a significant performance degradation when suffering from occlusion caused by various…

Cited by 16SourcePDFScholar
2024

RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification

NeurIPS 2024poster

This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this…

2023

SiMFy: A Simple Yet Effective Approach for Temporal Knowledge Graph Reasoning

EMNLP 2023long findings

Temporal Knowledge Graph (TKG) reasoning, which focuses on leveraging temporal information to infer future facts in knowledge graphs, plays a vital role in knowledge graph completion. Typically, existing works for this task design graph neural networks and recurrent neural networks to respectively c…

Cited by 0SourceScholar
2020

PlugNet: Degradation Aware Scene Text Recognition Supervised by a Pluggable Super-Resolution Unit

ECCV 2020poster

In this paper, we address the problem of recognizing degradation images that are suffering from high blur or low-resolution. We propose a novel degradation aware scene text recognizer with a pluggable super-resolution unit (PlugNet) to recognize low-quality scene text to solve this task from the fea…

Cited by 107SourcePDFScholar
2015

Panoptic Studio: A Massively Multiview System for Social Motion Capture

ICCV 2015oral

We present an approach to capture the 3D structure and motion of a group of people engaged in a social interaction. The core challenges in capturing social interactions are: (1) occlusion is functional and frequent; (2) subtle motion needs to be measured over a space large enough to host a social gr…

Cited by 1063PDFScholar