← Search

Qingyi Tao

11 accepted papers

2025

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

ICCV 2025poster

Unifying visual understanding and generation within a single multimodal framework remains a significant challenge, as the two inherently heterogeneous tasks require representations at different levels of granularity. Current approaches that utilize vector quantization (VQ) or variational autoencoder…

2025

MatAnyone: Stable Video Matting with Consistent Memory Propagation

CVPR 2025poster

Auxiliary-free human video matting methods, which rely solely on input frames, often struggle with complex or ambiguous backgrounds. To tackle this, we propose MatAnyone, a practical framework designed for target-assigned video matting. Specifically, building on a memory-based framework, we introduc…

Cited by 2SourcePDFScholar
2025

SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer

ICCV 2025poster

Photorealistic style transfer (PST) enables real-world color grading by adapting reference image colors while preserving content structure.Existing methods mainly follow either approaches: generation-based methods that prioritize stylistic fidelity at the cost of content integrity and efficiency, or…

2024

Modeling Continuous Motion for 3D Point Cloud Object Tracking

AAAI 2024technical

The task of 3D single object tracking (SOT) with LiDAR point clouds is crucial for various applications, such as autonomous driving and robotics. However, existing approaches have primarily relied on appearance matching or motion modeling within only two successive frames, thereby overlooking the lo…

Cited by 6SourcePDFScholar
2023

PGDiff: Guiding Diffusion Models for Versatile Face Restoration via Partial Guidance

NeurIPS 2023poster

Exploiting pre-trained diffusion models for restoration has recently become a favored alternative to the traditional task-specific training approach. Previous works have achieved noteworthy success by limiting the solution space using explicit degradation models. However, these methods often fall sh…

2023

Towards Robust and Expressive Whole-body Human Pose and Shape Estimation

NeurIPS 2023poster

Whole-body pose and shape estimation aims to jointly predict different behaviors (e.g., pose, hand gesture, facial expression) of the entire human body from a monocular image. Existing methods often exhibit suboptimal performance due to the complexity of in-the-wild scenarios. We argue that the pred…

2020

Exploring Bottom-Up and Top-Down Cues With Attentive Learning for Webly Supervised Object Detection

CVPR 2020poster

Fully supervised object detection has achieved great success in recent years. However, abundant bounding boxes annotations are needed for training a detector for novel classes. To reduce the human labeling effort, we propose a novel webly supervised object detection (WebSOD) method for novel classes…

Cited by 13PDFScholar
2018

VQA-E: Explaining, Elaborating, and Enhancing Your Answers for Visual Questions

ECCV 2018poster

Most existing works in visual question answering (VQA) are dedicated to improving the accuracy of predicted answers, while disregarding the explanations. We argue that the explanation for an answer is of the same or even more importance compared with the answer itself, since it makes the question an…

Cited by 138SourcePDFScholar