← Search

Minlong Lu

5 accepted papers

2025

Tracing Copied Pixels and Regularizing Patch Affinity in Copy Detection

ICCV 2025poster

Image Copy Detection (ICD) aims to identify manipulated content between image pairs through robust feature representation learning. While self-supervised learning (SSL) has advanced ICD systems, existing view-level contrastive methods struggle with sophisticated edits due to insufficient fine-graine…

Cited by 0SourcePDFScholar
2023

TransVCL: Attention-Enhanced Video Copy Localization Network with Flexible Supervision

AAAI 2023technical

Video copy localization aims to precisely localize all the copied segments within a pair of untrimmed videos in video retrieval applications. Previous methods typically start from frame-to-frame similarity matrix generated by cosine similarity between frame-level features of the input video pair, an…

2022

Bridging the Gap between Reality and Ideality of Entity Matching: A Revisting and Benchmark Re-Constrcution

IJCAI 2022poster

Entity matching (EM) is the most critical step for entity resolution (ER). While current deep learning-based methods achieve very impressive performance on standard EM benchmarks, their real-world application performance is much frustrating. In this paper, we highlight that such the gap between real…

2022

M5Product: Self-Harmonized Contrastive Learning for E-Commercial Multi-Modal Pretraining

CVPR 2022poster

Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets. By leveraging the natural suitability of E-commerce, where different mod…

Cited by 44PDFcodeScholar
2021

Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-Modal Pretraining

ICCV 2021poster

Nowadays, customer's demands for E-commerce are more diversified, which introduces more complications to the product retrieval industry. Previous methods are either subject to single-modal input or perform supervised image-level product retrieval, thus fail to accommodate real-life scenarios where e…

Cited by 75PDFcodeScholar