← Search

Pichao WANG

31 accepted papers

2026

CTRL&SHIFT: High-quality Geometry-Aware Object Manipulation in Visual Generation

ICLR 2026poster

Object-level manipulation—relocating or reorienting objects in images or videos while preserving scene realism—is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core goals: background preservation, geometric consistency under viewpoi…

Cited by 0SourceScholar
2025

Beyond Speaker Identity: Text Guided Target Speech Extraction

ICASSP 2025accepted

Target Speech Extraction (TSE) traditionally relies on explicit clues about the speaker’s identity like enrollment audio, face images, or videos, which may not always be available. In this paper, we propose a text-guided TSE model StyleTSE that uses natural language descriptions of speaking style in…

Cited by 0SourceScholar
2025

Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach

ICLR 2025poster

As online video content rapidly grows, the task of text-video retrieval (TVR) becomes increasingly important. A key challenge in TVR is the information asymmetry between video and text: videos are inherently richer in information, while their textual descriptions often capture only fragments of this…

Cited by 0SourcePDFScholar
2025

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation

ACL 2025finding

Large language models (LLMs) have shown great potential in natural language processing tasks, but their application to machine translation (MT) remains challenging due to pretraining on English-centric data and the complexity of reinforcement learning from human feedback (RLHF). Direct Preference Op…

Cited by 0SourcePDFScholar
2025

Learning Rich Speech Representations with Acoustic-Semantic Factorization

ICASSP 2025accepted

Self-supervised pretraining has transformed speech representation learning, enabling models to generalize across various downstream tasks. However, empirical studies have highlighted two notable gaps. First, different speech tasks require varying levels of acoustic and semantic information, which ar…

Cited by 0SourceScholar
2025

SparseDiT: Token Sparsification for Efficient Diffusion Transformer

NeurIPS 2025poster

Diffusion Transformers (DiT) are renowned for their impressive generative performance; however, they are significantly constrained by considerable computational costs due to the quadratic complexity in self-attention and the extensive sampling steps required. While advancements have been made in exp…

Cited by 0SourcecodeScholar
2025

Training-Free Text-Guided Image Editing with Visual Autoregressive Model

ICCV 2025poster

Text-guided image editing is an essential task, enabling users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input…

2024

Diffusion-Inspired Truncated Sampler for Text-Video Retrieval

NeurIPS 2024poster

Prevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substa…

Cited by 2SourcePDFScholar
2024

Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning

NeurIPS 2024poster

Despite advancements in Text-to-Video (T2V) generation, producing videos with realistic motion remains challenging. Current models often yield static or minimally dynamic outputs, failing to capture complex motions described by text. This issue stems from the internal biases in text encoding which o…

2024

Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation

CVPR 2024highlight

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper we present a plug-and-play pruning-and-recovering framew…

2024

One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

NeurIPS 2024poster

We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA gen…

2024

Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval

CVPR 2024highlight

The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video relying on consistent embedding representations to compute similarity. However the text content in existing datasets is gene…

Cited by 40SourcePDFScholar
2023

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

ICCV 2023oral

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advance…

Cited by 20PDFScholar
2023

Frequency Domain Disentanglement for Arbitrary Neural Style Transfer

AAAI 2023technical

Arbitrary neural style transfer has been a popular research topic due to its rich application scenarios. Effective disentanglement of content and style is the critical factor for synthesizing an image with arbitrary style. The existing methods focus on disentangling feature representations of conten…

Cited by 5SourcePDFScholar
2023

Making Vision Transformers Efficient From a Token Sparsification View

CVPR 2023poster

The quadratic computational complexity to the number of tokens limits the practical applications of Vision Transformers (ViTs). Several works propose to prune redundant tokens to achieve efficient ViTs. However, these methods generally suffer from (i) dramatic accuracy drops, (ii) application diffic…

2023

PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation

CVPR 2023poster

Recently, transformer-based methods have gained significant success in sequential 2D-to-3D lifting human pose estimation. As a pioneering work, PoseFormer captures spatial relations of human joints in each video frame and human dynamics across frames with cascaded transformer layers and has achieved…

2023

Revisiting Vision Transformer from the View of Path Ensemble

ICCV 2023oral

Vision Transformers (ViTs) are normally regarded as a stack of transformer layers. In this work, we propose a novel view of ViTs showing that they can be seen as ensemble networks containing multiple parallel paths with different lengths. Specifically, we equivalently transform the traditional casca…

Cited by 6PDFcodeScholar
2023

Selective Structured State-Spaces for Long-Form Video Understanding

CVPR 2023poster

Effective modeling of complex spatiotemporal dependencies in long-form videos remains an open problem. The recently proposed Structured State-Space Sequence (S4) model with its linear complexity offers a promising direction in this space. However, we demonstrate that treating all image-tokens equall…

Cited by 126SourcePDFScholar
2022

CDTrans: Cross-domain Transformer for Unsupervised Domain Adaptation

ICLR 2022poster

Unsupervised domain adaptation (UDA) aims to transfer knowledge learned from a labeled source domain to a different unlabeled target domain. Most existing UDA methods focus on learning domain-invariant feature representation, either from the domain level or category level, using convolution neural n…

2022

Decoupling and Recoupling Spatiotemporal Representation for RGB-D-Based Motion Recognition

CVPR 2022poster

Decoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, the…

Cited by 46PDFcodeScholar
2022

EPro-PnP: Generalized End-to-End Probabilistic Perspective-N-Points for Monocular Object Pose Estimation

CVPR 2022oral

Locating 3D objects from a single RGB image via Perspective-n-Points (PnP) is a long-standing problem in computer vision. Driven by end-to-end deep learning, recent studies suggest interpreting PnP as a differentiable layer, so that 2D-3D point correspondences can be partly learned by backpropagatin…

Cited by 196PDFcodeScholar
2022

KVT: k-NN Attention for Boosting Vision Transformers

ECCV 2022poster

"Convolutional Neural Networks (CNNs) have dominated computer vision for years, due to its ability in capturing locality and translation invariance. Recently, many vision transformer architectures have been proposed and they show promising performance. A key component in vision transformers is the f…

2022

MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

CVPR 2022poster

Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasib…

Cited by 415PDFcodeScholar
2022

Scaled ReLU Matters for Training Vision Transformers

AAAI 2022technical

Vision transformers (ViTs) have been an alternative design paradigm to convolutional neural networks (CNNs). However, the training of ViTs is much harder than CNNs, as it is sensitive to the training parameters, such as learning rate, optimizer and warmup epoch. The reasons for training difficulty a…

Cited by 46SourcePDFScholar
2022

TransFGU: A Top-down Approach to Fine-Grained Unsupervised Semantic Segmentation

ECCV 2022poster

"Unsupervised semantic segmentation aims to obtain high-level semantic representation on low-level visual features without manual annotations. Most existing methods are bottom-up approaches that try to group pixels into regions based on their visual cues or certain predefined rules. As a result, it…

2022

VTC-LFC: Vision Transformer Compression with Low-Frequency Components

NeurIPS 2022accept

Although Vision transformers (ViTs) have recently dominated many vision tasks, deploying ViT models on resource-limited devices remains a challenging problem. To address such a challenge, several methods have been proposed to compress ViTs. Most of them borrow experience in convolutional neural netw…

Cited by 38SourcePDFScholar
2021

TransReID: Transformer-Based Object Re-Identification

ICCV 2021poster

Extracting robust feature representation is one of the key challenges in object re-identification (ReID). Although convolution neural network (CNN)-based methods have achieved great success, they only process one local neighborhood at a time and suffer from information loss on details caused by conv…

Cited by 1156PDFcodeScholar
2021

Zen-NAS: A Zero-Shot NAS for High-Performance Image Recognition

ICCV 2021poster

Accuracy predictor is a key component in Neural Architecture Search (NAS) for ranking architectures. Building a high-quality accuracy predictor usually costs enormous computation. To address this issue, instead of using an accuracy predictor, we propose a novel zero-shot index dubbed Zen-Score to ra…

Cited by 184PDFcodeScholar
2017

Scene Flow to Action Map: A New Representation for RGB-D Based Action Recognition With Convolutional Neural Networks

CVPR 2017poster

Scene flow describes the motion of 3D objects in real world and potentially could be the basis of a good feature for 3D action recognition. However, its use for action recognition, especially in the context of convolutional neural networks (ConvNets), has not been previously studied. In this paper,…

Cited by 176PDFScholar