← Search

Rang Nguyen

10 accepted papers

2025

Semi-supervised 3D Semantic Scene Completion with 2D Vision Foundation Model Guidance

AAAI 2025technical

Accurate prediction of 3D semantic occupancy from 2D visual images is vital in enabling autonomous agents to comprehend their surroundings for planning and navigation. State-of-the-art methods typically employ fully supervised approaches, necessitating a huge labeled dataset acquired through expensi…

Cited by 0SourcePDFScholar
2025

SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation

CVPR 2025poster

We propose SharpDepth, a novel approach to monocular metric depth estimation that combines the metric accuracy of discriminative depth estimation methods (e.g., Metric3D, UniDepth) with the fine-grained boundary sharpness typically achieved by generative methods (e.g., Marigold, Lotus). Traditional…

Cited by 2SourcePDFScholar
2025

SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models

AAAI 2025technical

Given an input video of a person and a new garment, the objective of this paper is to synthesize a new video where the person is wearing the specified garment while maintaining spatiotemporal consistency. Although significant advances have been made in image-based virtual try-on, extending these suc…

Cited by 15SourcePDFScholar
2024

Blur2Blur: Blur Conversion for Unsupervised Image Deblurring on Unknown Domains

CVPR 2024poster

This paper presents an innovative framework designed to train an image deblurring algorithm tailored to a specific camera device. This algorithm works by transforming a blurry input image which is challenging to deblur into another blurry image that is more amenable to deblurring. The transformation…

2023

HyperCUT: Video Sequence From a Single Blurry Image Using Unsupervised Ordering

CVPR 2023poster

We consider the challenging task of training models for image-to-video deblurring, which aims to recover a sequence of sharp images corresponding to a given blurry image input. A critical issue disturbing the training of an image-to-video model is the ambiguity of the frame ordering since both the f…

2022

HyperInverter: Improving StyleGAN Inversion via Hypernetwork

CVPR 2022poster

Real-world image manipulation has achieved fantastic progress in recent years as a result of the exploration and utilization of GAN latent spaces. GAN inversion is the first step in this pipeline, which aims to map the real image to the latent code faithfully. Unfortunately, the majority of existing…

Cited by 152PDFcodeScholar
2022

Inductive and Transductive Few-Shot Video Classification via Appearance and Temporal Alignments

ECCV 2022poster

"We present a novel method for few-shot video classification, which performs appearance and temporal alignments. In particular, given a pair of query and support videos, we conduct appearance alignment via frame-level feature matching to achieve the appearance similarity score between the videos, wh…

2021

POODLE: Improving Few-shot Learning via Penalizing Out-of-Distribution Samples

NeurIPS 2021poster

In this work, we propose to use out-of-distribution samples, i.e., unlabeled samples coming from outside the target classes, to improve few-shot learning. Specifically, we exploit the easily available out-of-distribution samples to drive the classifier to avoid irrelevant features by maximizing the…

2019

Project AutoVision: Localization and 3D Scene Perception for an Autonomous Vehicle with a Multi-Camera System

ICRA 2019poster

Project AutoVision aims to develop localization and 3D scene perception capabilities for a self-driving vehicle. Such capabilities will enable autonomous navigation in urban and rural environments, in day and night, and with cameras as the only exteroceptive sensors. The sensor suite employs many ca…

Cited by 149SourceScholar