← Search

Pengwen Dai

11 accepted papers

2026

CollectiveKV: Decoupling and Sharing Collaborative Information in Sequential Recommendation

ICLR 2026poster

Sequential recommendation models are widely used in applications, yet they face stringent latency requirements. Mainstream models leverage the Transformer attention mechanism to improve performance, but its computational complexity grows with the sequence length, leading to a latency challenge for…

Cited by 0SourceScholar
2026

Discretization Is Not Always Better: Rethinking Deep Quantization for Asymmetric Image Retrieval

AAAI 2026technical

Asymmetric image retrieval (AIR), which typically employs a compact model for the query side and a large model for the database server, has garnered significant attention in resource-constrained environments. While deep hashing methods have shown great potential in large-scale image retrieval, curre

Cited by 0SourcePDFScholar
2026

DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection

CVPR 2026

Multimodal tiny object detection plays a critical role in real-world applications, yet remains highly challenging due to weak target representations and complex cross-modal interference. Existing frequency-domain methods for tiny object detection are still largely limited to the visible modality and

Cited by 0SourceScholar
2026

EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval

CVPR 2026

Text-video retrieval tasks have seen significant improvements due to the recent development of large-scale vision-language pre-trained models. Traditional methods primarily focus on video representations or cross-modal alignment, while recent works shift toward enriching text expressiveness to bette

Cited by 0SourcecodeScholar
2026

One2Seq: One-Token Wise Decoder for Efficient Scene Text Recognition

AAAI 2026technical

Auto-regressive (AR)-based decoders, owing to their flexibility in handling variable-length outputs and their strong capability in modeling character-level dependencies, have emerged as the predominant decoding paradigm in the field of scene text recognition (STR). However, AR-based decoders suffer

Cited by 0SourcePDFScholar
2025

CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR

CVPR 2025poster

Scene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we r…

Cited by 0SourcePDFScholar
2025

Decoupled Graph Energy-based Model for Node Out-of-Distribution Detection on Heterophilic Graphs

ICLR 2025poster

Despite extensive research efforts focused on Out-of-Distribution (OOD) detection on images, OOD detection on nodes in graph learning remains underexplored. The dependence among graph nodes hinders the trivial adaptation of existing approaches on images that assume inputs to be i.i.d. sampled, since…

2025

Endogenous Recovery via Within-modality Prototypes for Incomplete Multimodal Hashing

IJCAI 2025

Multimodal hashing projects multimodal data into compact binary codes, enabling rapid and storage-efficient retrieval of large-scale multimedia content. In practical scenarios, the issue of missing modality frequently arises when dealing with multimodal data. Existing incomplete multimodal hashing t

2025

Segmentation-Guided Sparse Transformer for Under-Display Camera Image Restoration

ICASSP 2025accepted

Under-display Camera is an emerging technology for full-screen display with a camera under the display. However, the current implementation of UDC causes serious image degradation. Incident light required for camera imaging undergoes attenuation and diffraction when passing through the display. Curr…

Cited by 0SourceScholar
2025

Towards Irreversible Attack: Fooling Scene Text Recognition via Multi-Population Coevolution Search

NeurIPS 2025poster

Recent work has shown that scene text recognition (STR) models are vulnerable to adversarial examples. Different from non-sequential vision tasks, the output sequence of STR models contains rich information. However, existing adversarial attacks against STR models can only lead to a few incorrect c…

Cited by 0SourcecodeScholar
2021

Progressive Contour Regression for Arbitrary-Shape Scene Text Detection

CVPR 2021poster

State-of-the-art scene text detection methods usually model the text instance with local pixels or components from the bottom-up perspective and, therefore, are sensitive to noises and dependent on the complicated heuristic post-processing especially for arbitrary-shape texts. To relieve these two i…

Cited by 141PDFcodeScholar