← Search

Rynson W. H. Lau

24 accepted papers

2026

GenSplat: Bridging the Generalization Gap in 3DGS Language Comprehension

CVPR 2026

In this paper, we propose GenSplat, a novel approach for language comprehension in 3D Gaussian Splatting (3DGS). Unlike previous methods that either achieve cross-scene generalization by being bounded to a predefined vocabulary or handle free-form language by overfitting to individual scenes, GenSpl

Cited by 0SourcecodeScholar
2026

OpenScan: A Benchmark for Generalized Open-Vocabulary 3D Scene Understanding

AAAI 2026technical

Open-vocabulary 3D scene understanding (OV-3D) aims to localize and classify novel objects beyond the closed set of object classes. However, existing approaches and benchmarks primarily focus on the open vocabulary problem within the context of object classes, which is insufficient in providing a ho

Cited by 0SourcePDFScholar
2026

RefSTAR: Blind Face Image Restoration with Reference Selection, Transfer, and Reconstruction

AAAI 2026technical

Introducing high-quality references can largely alleviate the uncertainty in blind face image restoration tasks, yet the equivocal utilization of reference priors makes it still a struggle to well preserve the human identity. We attribute the identity inconsistency to two deficiencies of existing re

Cited by 0SourcePDFScholar
2026

Video Mirror Detection with the Motion-in-Depth Cue

AAAI 2026technical

Detecting mirror regions in RGB videos is essential for scene understanding in applications such as scene reconstruction and robotic navigation. Existing video mirror detectors typically rely on cues like inside-outside mirror correspondences and 2D motion inconsistencies. However, these methods oft

Cited by 0SourcePDFScholar
2025

DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors

AAAI 2025technical

Dynamic 3D interaction has been attracting a lot of attention recently. However, creating such 4D content remains challenging. One solution is to animate 3D scenes with physics-based simulation, which requires manually assigning precise physical properties to the object or the simulated results woul…

2025

GenColor: Generative and Expressive Color Enhancement with Pixel-Perfect Texture Preservation

NeurIPS 2025spotlight

Color enhancement is a crucial yet challenging task in digital photography. It demands methods that are (i) expressive enough for fine-grained adjustments, (ii) adaptable to diverse inputs, and (iii) able to preserve texture. Existing approaches typically fall short in at least one of these aspects,…

Cited by 0SourceScholar
2025

Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection

AAAI 2025technical

Open-vocabulary 3D object detection (OV-3DOD) aims at localizing and classifying novel objects beyond closed sets. The recent success of vision-language models (VLMs) has demonstrated their remarkable capabilities to understand open vocabularies. Existing works that leverage VLMs for 3D object detec…

Cited by 0SourcePDFScholar
2025

Language-Guided Salient Object Ranking

CVPR 2025poster

Salient Object Ranking (SOR) aims to study human attention shifts across different objects in the scene. It is a challenging task, as it requires comprehension of the relations among the salient objects in the scene. However, existing works often overlook such relations or model them implicitly. In…

Cited by 0SourcePDFScholar
2025

Leveraging RGB-D Data with Cross-Modal Context Mining for Glass Surface Detection

AAAI 2025technical

Glass surfaces are becoming increasingly ubiquitous as modern buildings tend to use a lot of glass panels. This, however, poses substantial challenges to the operations of autonomous systems such as robots, self-driving cars, and drones, as the glass panels can become transparent obstacles to naviga…

Cited by 1SourcePDFScholar
2025

Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion

ICLR 2025poster

Generative 3D modeling has made significant advances recently, but it remains constrained by its inherently ill-posed nature, leading to challenges in quality and controllability. Inspired by the real-world workflow that designers typically refer to existing 3D models when creating new ones, we prop…

Cited by 5SourcePDFScholar
2025

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

NeurIPS 2025poster

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-shot solution in STVG. We reveal two key insights about MLLMs: (1) MLLMs tend to…

Cited by 0SourcecodeScholar
2024

Boosting Weakly Supervised Referring Image Segmentation via Progressive Comprehension

NeurIPS 2024poster

This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-text pairs. We note that the input text description typically already contains detailed information on how to localize t…

Cited by 2SourcePDFScholar
2024

LuSh-NeRF: Lighting up and Sharpening NeRFs for Low-light Scenes

NeurIPS 2024poster

Neural Radiance Fields (NeRFs) have shown remarkable performances in producing novel-view images from high-quality scene images. However, hand-held low-light photography challenges NeRFs as the captured images may simultaneously suffer from low visibility, noise, and camera shakes. While existing Ne…

2024

Revisiting the Integration of Convolution and Attention for Vision Backbone

NeurIPS 2024poster

Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate both, they apply the two operators simultaneously at the finest pixel granularity. With Convs responsible for per-pixel…

2024

TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields

ICLR 2024poster

Recent works learn 3D representation explicitly under text-3D guidance. However, limited text-3D data restricts the vocabulary scale and text control of generations. Generators may easily fall into a stereotype concept for certain text prompts, thus losing open-vocabulary generation ability. To tack…

Cited by 12SourcePDFScholar
2022

Geometry-aware Two-scale PIFu Representation for Human Reconstruction

NeurIPS 2022accept

Although PIFu-based 3D human reconstruction methods are popular, the quality of recovered details is still unsatisfactory. In a sparse (e.g., 3 RGBD sensors) capture setting, the depth noise is typically amplified in the PIFu representation, resulting in flat facial surfaces and geometry-fallible bo…

Cited by 17SourcePDFScholar
2021

What Makes Instance Discrimination Good for Transfer Learning?

ICLR 2021poster

Contrastive visual pretraining based on the instance discrimination pretext task has made significant progress. Notably, recent work on unsupervised pretraining has shown to surpass the supervised counterpart for finetuning downstream applications such as object detection and segmentation. It come…

Cited by 201SourcePDFScholar
2017

CREST: Convolutional Residual Learning for Visual Tracking

ICCV 2017poster

Discriminative correlation filters (DCFs) have \ryn been shown to perform superiorly in visual tracking. They \ryn only need a small set of training samples from the initial frame to generate an appearance model. However, existing DCFs learn the filters separately from feature extraction, and upda…

Cited by 652PDFScholar
2017

DeshadowNet: A Multi-Context Embedding Deep Network for Shadow Removal

CVPR 2017spotlight

Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a…

Cited by 367PDFcodeScholar
2017

Learning Fully Convolutional Networks for Iterative Non-Blind Deconvolution

CVPR 2017poster

In this paper, we propose a fully convolutional network for iterative non-blind deconvolution. We decompose the non-blind deconvolution problem into image denoising and image deconvolution. We train a FCNN to remove noise in the gradient domain and use the learned gradients to guide the image deconv…

Cited by 215PDFScholar