← Search

Kwan-Yee K. Wong

33 accepted papers

2026

DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution

CVPR 2026

Diffusion-based video super-resolution (VSR) achieves remarkable fidelity but suffers from prohibitive sampling cost. While distribution matching distillation (DMD) accelerates diffusion models to one-step generation, directly applying it to VSR leads to training instability and degraded, insufficie

Cited by 0SourcecodeScholar
2025

Affordances-Oriented Planning Using Foundation Models for Continuous Vision-Language Navigation

AAAI 2025technical

LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) task. However, existing LLM-based methods often focus only on solving high-level task planning by selecting nodes in predefined navigation graphs for movements, overlooking low-level control in na…

Cited by 8SourcePDFScholar
2025

ArtiFade: Learning to Generate High-quality Subject from Blemished Images

CVPR 2025poster

Subject-driven text-to-image generation has demonstrated remarkable advancements in its ability to learn and capture characteristics of a subject using only a limited number of images. However, existing methods commonly rely on high-quality images for training and often struggle to generate reasonab…

Cited by 1SourcePDFScholar
2025

AvatarGO: Zero-shot 4D Human-Object Interaction Generation and Animation

ICLR 2025poster

Recent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation, which is limited by the scarcity of realistic large-scale inte…

Cited by 6SourcePDFScholar
2025

BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities

ICLR 2025poster

We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same fr…

2025

DriveGPT4-V2: Harnessing Large Language Model Capabilities for Enhanced Closed-Loop Autonomous Driving

CVPR 2025highlight

Multimodal large language models (MLLMs) possess the ability to comprehend visual images or videos, and show impressive reasoning ability thanks to the vast amounts of pretrained knowledge, making them highly suitable for autonomous driving applications. Unlike the previous work, DriveGPT4-V1, which…

Cited by 0SourcePDFScholar
2025

Dual-Expert Consistency Model for Efficient and High-Quality Video Generation

ICCV 2025poster

Diffusion Models have achieved remarkable results in video synthesis but require iterative denoising steps, leading to substantial computational overhead. Consistency Models have made significant progress in accelerating diffusion models. However, directly applying them to video diffusion models oft…

2025

FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality

ICLR 2025poster

In this paper, we present \textbf{\textit{FasterCache}}, a novel training-free strategy designed to accelerate the inference of video diffusion models with high-quality generation. By analyzing existing cache-based methods, we observe that \textit{directly reusing adjacent-step features degrades vid…

Cited by 6SourcePDFScholar
2025

Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers

ICCV 2025poster

Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-driven visual generation. However, even state-of-the-art MM-DiT models like FLUX struggle with achieving precise alignment between text prompts and generated content. We identify two key issues in the attention mec…

2025

SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning

NeurIPS 2025poster

Evaluating the step-by-step reliability of large language model (LLM) reasoning, such as Chain-of-Thought, remains challenging due to the difficulty and cost of obtaining high-quality step-level supervision. In this paper, we introduce Self-Play Critic (SPC), a novel approach where a critic model ev…

Cited by 0SourceScholar
2024

DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion Models

CVPR 2024poster

We present DreamAvatar a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common object generation generating high-quality human avatars remains an open challenge due…

2024

DriveGPT4: Interpretable End-to-End Autonomous Driving Via Large Language Model

RA-L 2024

Multimodallarge language models (MLLMs) have emerged as a prominent area of interest within the research community, given their proficiency in handling and reasoning with non-textual data, including images and videos. This study seeks to extend the application of MLLMs to the realm of autonomous dri

Cited by 603SourceScholar
2024

InsMapper: Exploring Inner-instance Information for Vectorized HD Mapping

ECCV 2024poster

"Vectorized high-definition (HD) maps contain detailed information about surrounding road elements, which are crucial for various downstream tasks in modern autonomous vehicles, such as motion planning and vehicle control. Recent works attempt to directly detect the vectorized HD map as a point set…

Cited by 7SourcePDFScholar
2024

PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis

CVPR 2024highlight

Recent advancements in large-scale pre-trained text-to-image models have led to remarkable progress in semantic image synthesis. Nevertheless synthesizing high-quality images with consistent semantics and layout remains a challenge. In this paper we propose the adaPtive LAyout-semantiC fusion modulE…

2023

HeadSculpt: Crafting 3D Head Avatars with Text

NeurIPS 2023poster

Recently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models. However, existing methods still struggle to create high-fidelity 3D head avatars in t…

Cited by 51SourcePDFScholar
2023

Learning Attention As Disentangler for Compositional Zero-Shot Learning

CVPR 2023poster

Compositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the disentanglement of the attribute-object composition. To this end, we propose to exp…

2023

RIGID: Recurrent GAN Inversion and Editing of Real Face Videos

ICCV 2023poster

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a unified recurrent framework, named Recurrent vIdeo GAN Inversi…

Cited by 8PDFcodeScholar
2023

SeSDF: Self-Evolved Signed Distance Field for Implicit 3D Clothed Human Reconstruction

CVPR 2023poster

We address the problem of clothed human reconstruction from a single image or uncalibrated multi-view images. Existing methods struggle with reconstructing detailed geometry of a clothed human and often require a calibrated setting for multi-view reconstruction. We propose a flexible framework which…

2023

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

NeurIPS 2023poster

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often struggle to adequately convey detailed controls, even when composed…

2022

Blind Image Super-Resolution With Elaborate Degradation Modeling on Noise and Kernel

CVPR 2022poster

While researches on model-based blind single image super-resolution (SISR) have achieved tremendous successes recently, most of them do not consider the image degradation sufficiently. Firstly, they always assume image noise obeys an independent and identically distributed (i.i.d.) Gaussian or Lapla…

Cited by 76PDFcodeScholar
2022

JIFF: Jointly-Aligned Implicit Face Function for High Quality Single View Clothed Human Reconstruction

CVPR 2022oral

This paper addresses the problem of single view 3D human reconstruction. Recent implicit function based methods have shown impressive results, but they fail to recover fine face details in their reconstructions. This largely degrades user experience in applications like 3D telepresence. In this pape…

Cited by 38PDFScholar
2022

PS-NeRF: Neural Inverse Rendering for Multi-View Photometric Stereo

ECCV 2022poster

"Traditional multi-view photometric stereo (MVPS) methods are often composed of multiple disjoint stages, resulting in noticeable accumulated errors. In this paper, we present a neural inverse rendering method for MVPS based on implicit representation. Given multi-view images of a non-Lambertian obj…

2022

S$^3$-NeRF: Neural Reflectance Field from Shading and Shadow under a Single Viewpoint

NeurIPS 2022accept

In this paper, we address the "dual problem" of multi-view scene reconstruction in which we utilize single-view images captured under different point lights to learn a neural scene representation. Different from existing single-view methods which can only recover a 2.5D scene representation (i.e., a…

2021

HDR Video Reconstruction: A Coarse-To-Fine Network and a Real-World Benchmark Dataset

ICCV 2021poster

High dynamic range (HDR) video reconstruction from sequences captured with alternating exposures is a very challenging problem. Existing methods often align low dynamic range (LDR) input sequence in the image space using optical flow, and then merge the aligned images to produce HDR output. However,…

Cited by 74PDFScholar
2021

Progressive Semantic-Aware Style Transformation for Blind Face Restoration

CVPR 2021poster

Face restoration is important in face image processing, and has been widely studied in recent years. However, previous works often fail to generate plausible high quality (HQ) results for real-world low quality (LQ) face images. In this paper, we propose a new progressive semantic-aware style transf…

Cited by 195PDFcodeScholar
2020

Cops-Ref: A New Dataset and Task on Compositional Referring Expression Comprehension

CVPR 2020poster

Referring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed f…

Cited by 75PDFScholar
2020

What is Learned in Deep Uncalibrated Photometric Stereo?

ECCV 2020poster

This paper targets at discovering what a deep uncalibrated photometric stereo network learns to resolve the problem’s inherent ambiguity, and designing an effective network architecture based on the new insight to improve the performance. The recently proposed deep uncalibrated photometric stereo me…

Cited by 63SourcePDFScholar
2019

Self-Calibrating Deep Photometric Stereo Networks

CVPR 2019oral

This paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a sc…

Cited by 182PDFcodeScholar
2018

TOM-Net: Learning Transparent Object Matting From a Single Image

CVPR 2018poster

This paper addresses the problem of transparent object matting. Existing image matting approaches for transparent objects often require tedious capturing procedures and long processing time, which limit their practical use. In this paper, we first formulate transparent object matting as a refractive…

2017

SCNet: Learning Semantic Correspondence

ICCV 2017poster

This paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearan…

Cited by 159PDFcodeScholar
2015

A Fixed Viewpoint Approach for Dense Reconstruction of Transparent Objects

CVPR 2015poster

This paper addresses the problem of reconstructing the surface shape of transparent objects. The difficulty of this problem originates from the viewpoint dependent appearance of a transparent object, which quickly makes reconstruction methods tailored for diffuse surfaces fail disgracefully. In this…

Cited by 47SourcePDFScholar