← Search

Congpei Qiu

6 accepted papers

2026

UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

CVPR 2026

Representation learning with Vision Transformers (ViTs) has advanced rapidly, yet the utility of large-scale models in spatially sensitive tasks is hindered by spurious tokens. Prior efforts to mitigate this have been limited, often defining these artifacts narrowly, for example, as simple high-norm

Cited by 0SourceScholar
2025

Generating Multimodal Driving Scenes via Next-Scene Prediction

CVPR 2025poster

Generative models in Autonomous Driving (AD) enable diverse scenario creation, yet existing methods fall short by only capturing a limited range of modalities, restricting the capability of generating controllable scenes for comprehensive evaluation of AD systems. In this paper, we introduce a multi…

2025

Refining CLIP's Spatial Awareness: A Visual-Centric Perspective

ICLR 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in global alignment with language but exhibits limited sensitivity to spatial information, leading to strong performance in zero-shot classification tasks but underperformance in tasks requiring precise spatial understanding. Recent approaches ha…

Cited by 0SourcePDFScholar
2024

Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised Learning

ICLR 2024poster

Dense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions…

Cited by 2SourcePDFScholar
2024

Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange

CVPR 2024poster

In the realm of point cloud scene understanding particularly in indoor scenes objects are arranged following human habits resulting in objects of certain semantics being closely positioned and displaying notable inter-object correlations. This can create a tendency for neural networks to exploit the…

2022

Leverage Your Local and Global Representations: A New Self-Supervised Learning Strategy

CVPR 2022poster

Self-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similarity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain d…

Cited by 40PDFcodeScholar