← Search

Wenjie Pei

29 accepted papers

2026

Beyond Heuristics: Learnable Density Control for 3D Gaussian Splatting

ICML 2026poster

While 3D Gaussian Splatting (3DGS) has demonstrated impressive real-time rendering performance, its efficacy remains constrained by a reliance on heuristic density control. Despite numerous refinements to these handcrafted rules, such methods inherently lack the flexibility to adapt to diverse scene…

Cited by 0SourceScholar
2026

DiffTrans: Differentiable Geometry-Materials Decomposition for Reconstructing Transparent Objects

ICLR 2026poster

Reconstructing transparent objects from a set of multi-view images is a challenging task due to the complicated nature and indeterminate behavior of light propagation. Typical methods are primarily tailored to specific scenarios, such as objects following a uniform topology, exhibiting ideal transpa…

Cited by 0SourceScholar
2026

PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D Tracking

ICLR 2026poster

3D single object tracking (SOT) remains a highly challenging task due to the inherent crux in learning representations from point clouds to effectively capture both spatial shape features and temporal motion features. Most existing methods employ a category-specific optimization paradigm, training t…

Cited by 0SourceScholar
2026

Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity

ICLR 2026poster

Vision-language models (VLMs) face significant computational inefficiencies caused by excessive generation of visual tokens. While prior work shows that a large fraction of visual tokens are redundant, existing compression methods struggle to balance \textit{importance preservation} and \textit{info…

Cited by 0SourcecodeScholar
2026

Too Vivid to Be Real? Benchmarking and Calibrating Generative Color Fidelity

CVPR 2026

Recent advances in text-to-image (T2I) generation have greatly improved visual quality, yet producing images that appear visually authentic to real-world photography remains challenging. This is partly due to biases in existing evaluation paradigms: human ratings and preference-trained metrics often

Cited by 0SourcecodeScholar
2025

D2ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-shot Action Recognition

ICCV 2025poster

Adapting pre-trained image models to video modality has proven to be an effective strategy for robust few-shot action recognition. In this work, we explore the potential of adapter tuning in image-to-video model adaptation and propose a novel video adapter tuning framework, called Disentangled-and-D…

2025

EditInfinity: Image Editing with Binary-Quantized Generative Models

NeurIPS 2025poster

Adapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the generative trajectory inversely for a given source image by image…

Cited by 0SourcecodeScholar
2025

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

ICCV 2025poster

Recent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial re…

Cited by 0SourcePDFScholar
2025

Learning Compatible Multi-Prize Subnetworks for Asymmetric Retrieval

CVPR 2025poster

Asymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for p…

2025

MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text-to-Video Generation

ICCV 2025poster

Existing text-to-video methods struggle to transfer motion smoothly from a reference object to a target object with significant differences in appearance or structure between them. To address this challenge, we introduce MotionShot, a training-free framework capable of parsing reference-target corre…

2024

AnyControl: Create Your Artwork with Versatile Control on Text-to-Image Generation

ECCV 2024poster

"The field of text-to-image (T2I) generation has made significant progress in recent years, largely driven by advancements in diffusion models. Linguistic control enables effective content creation, but struggles with fine-grained control over image generation. This challenge has been explored, to a…

2024

Domain-Rectifying Adapter for Cross-Domain Few-Shot Segmentation

CVPR 2024poster

Few-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes supported by only a few annotated samples. However existing FSS methods often underperform in the presence of domain shifts especially when encountering new domain styles that are unseen during tra…

2024

Robust 3D Tracking with Quality-Aware Shape Completion

AAAI 2024technical

3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless,…

Cited by 6SourcePDFScholar
2024

SA²VP: Spatially Aligned-and-Adapted Visual Prompt

AAAI 2024technical

As a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embe…

2023

Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption Detection

ICCV 2023accepted

Effective image restoration with large-size corruptions, such as blind image inpainting, entails precise detection of corruption region masks which remains extremely challenging due to diverse shapes and patterns of corruptions. In this work, we present a novel method for automatic corruption detect…

2022

Alleviating the Sample Selection Bias in Few-shot Learning by Removing Projection to the Centroid

NeurIPS 2022accept

Few-shot learning (FSL) targets at generalization of vision models towards unseen tasks without sufficient annotations. Despite the emergence of a number of few-shot learning methods, the sample selection bias problem, i.e., the sensitivity to the limited amount of support data, has not been well un…

2022

Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations

ECCV 2022poster

"While fine-tuning based methods for few-shot object detection have achieved remarkable progress, a crucial challenge that has not been addressed well is the potential class-specific overfitting on base classes and sample-specific overfitting on novel classes. In this work we design a novel knowledg…

Cited by 18SourcePDFScholar
2022

Global Tracking via Ensemble of Local Trackers

CVPR 2022poster

The crux of long-term tracking lies in the difficulty of tracking the target with discontinuous moving caused by out-of-view or occlusion. Existing long-term tracking methods follow two typical strategies. The first strategy employs a local tracker to perform smooth tracking and uses another re-dete…

Cited by 44PDFcodeScholar
2022

Multi-Modal Sarcasm Detection via Cross-Modal Graph Convolutional Network

ACL 2022long

With the increasing popularity of posting multimodal messages online, many recent studies have been carried out utilizing both textual and visual information for multi-modal sarcasm detection. In this paper, we investigate multi-modal sarcasm detection from a novel perspective by constructing a cros…

2022

Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection

ECCV 2022poster

"Most of existing methods for few-shot object detection follow the fine-tuning paradigm, which potentially assumes that the class-agnostic generalizable knowledge can be learned and transferred implicitly from base classes with abundant samples to novel classes with limited samples via such a two-st…

2021

Audio2Gestures: Generating Diverse Gestures From Speech Audio With Conditional Variational Autoencoders

ICCV 2021poster

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, resulting in plain/boring motions during…

Cited by 130PDFcodeScholar
2020

CPGAN: Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis

ECCV 2020poster

Typical methods for text-to-image synthesis seek to design effective generative architecture to model the text-to-image mapping directly. It is fairly arduous due to the cross-modality translation. In this paper we circumvent this problem by focusing on parsing the content of both the input text and…

Cited by 89SourcePDFScholar
2020

Commonality-Parsing Network across Shape and Appearance for Partially Supervised Instance Segmentation

ECCV 2020poster

Partially supervised instance segmentation aims to perform learning on limited mask-annotated categories of data thus eliminating expensive and exhaustive mask annotation. The learned models are expected to be generalizable to novel categories. Existing methods either learn a transfer function from…

2019

Memory-Attended Recurrent Network for Video Captioning

CVPR 2019poster

Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in tr…

Cited by 294PDFScholar
2019

Non-Local Recurrent Neural Memory for Supervised Sequence Modeling

ICCV 2019oral

Typical methods for supervised sequence modeling are built upon the recurrent neural networks to capture temporal dependencies. One potential limitation of these methods is that they only model explicitly information interactions between adjacent time steps in a sequence, hence the high-order intera…

Cited by 13PDFcodeScholar
2017

Temporal Attention-Gated Model for Robust Sequence Classification

CVPR 2017poster

Typical techniques for sequence classification are designed for well-segmented sequences which have been edited to remove noisy or irrelevant parts. Therefore, such methods cannot be easily applied on noisy sequences expected in real-world applications. In this paper, we present the Temporal Attent…

Cited by 103PDFcodeScholar