← Search

Alex Wong

40 accepted papers

2026

CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning

ICLR 2026poster

Unsupervised 3D representation learning reduces the burden of labeling multimodal 3D data for fusion perception tasks. Among different pre-training paradigms, differentiable-rendering-based methods have shown most promise. However, existing works separately conduct pre-training for each modalities d…

Cited by 0SourcecodeScholar
2026

Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression

ICLR 2026poster

We propose a method for audio-visual knowledge distillation. Existing methods typically distill from the latent embeddings or outputs. The former requires matching feature dimensions, if not the same architecture, between teacher and student models while the latter supports any teacher-student pairi…

Cited by 0SourceScholar
2026

Iris: Integrating Language into Diffusion-based Monocular Depth Estimation

CVPR 2026

Conventional monocular depth estimators suffer from visual ambiguities and nuisances. We demonstrate that language can improve the fidelity of estimates by providing additional information through text as a condition, thereby reducing the solution space for depth estimates. This conditional distribu

Cited by 0SourceScholar
2026

ODE-GS: Latent ODEs for Dynamic Scene Extrapolation with 3D Gaussian Splatting

ICLR 2026poster

We introduce ODE-GS, a novel approach that integrates 3D Gaussian Splatting with latent neural ordinary differential equations (ODEs) to enable future extrapolation of dynamic 3D scenes. Unlike existing dynamic scene reconstruction methods, which rely on time-conditioned deformation networks and are…

Cited by 0SourcecodeScholar
2026

ORCaS: Unsupervised Depth Completion via Occluded Region Completion as Supervision

ICLR 2026poster

We propose a method for inferring an egocentric dense depth map from an RGB image and a sparse point cloud. The crux of our method lies in modeling the 3D scene implicitly within the latent space and learning an inductive bias in an unsupervised manner through principles of Structure-from-Motion. T…

Cited by 0SourceScholar
2026

SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild

CVPR 2026

Accurate 3D understanding of human hands and objects during manipulation remains a significant challenge for egocentric computer vision. Existing hand-object interaction datasets are predominantly captured in controlled studio settings, which limits both environmental diversity and the ability of mo

Cited by 0SourcecodeScholar
2025

ETA: Energy-based Test-time Adaptation for Depth Completion

ICCV 2025poster

We propose a method of adapting pretrained depth completion models to test time data in an unsupervised manner. Depth completion models are (pre)trained to produce dense depth maps from pairs of RGB image and sparse depth maps in ideal capture conditions (source domain), e.g., well-illuminated, high…

Cited by 0SourcePDFScholar
2025

Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens

ICCV 2025accepted

We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite being trained on tens of millions of images, FMDEs are susceptible to the covariate shift introduced by changes in camera calibration (intrinsic, distortion) param…

2025

Progressive Test Time Energy Adaptation for Medical Image Segmentation

ICCV 2025poster

We propose a model-agnostic, progressive test-time energy adaptation approach for medical image segmentation. Maintaining model performance across diverse medical datasets is challenging, as distribution shifts arise from inconsistent imaging protocols and patient variations. Unlike domain adaptatio…

2025

ProtoDepth: Unsupervised Continual Depth Completion with Prototypes

CVPR 2025poster

We present ProtoDepth, a novel prototype-based approach for continual learning of unsupervised depth completion, the multimodal 3D reconstruction task of predicting dense depth maps from RGB images and sparse point clouds. The unsupervised learning paradigm is well-suited for continual learning, as…

Cited by 1SourcePDFScholar
2025

STree: Speculative Tree Decoding for Hybrid State Space Models

NeurIPS 2025poster

Speculative decoding is a technique to leverage hardware concurrency in order to enable multiple steps of token generation in a single forward pass, thus improving the efficiency of large-scale autoregressive (AR) Transformer models. State-space models (SSMs) are already more efficient than AR Trans…

Cited by 0SourcecodeScholar
2025

TREND: Unsupervised 3D Representation Learning via Temporal Forecasting for LiDAR Perception

NeurIPS 2025spotlight

Labeling LiDAR point clouds is notoriously time-and-energy-consuming, which spurs recent unsupervised 3D representation learning methods to alleviate the labeling burden in LiDAR perception via pretrained weights. Existing work focus on either masked auto encoding or contrastive learning on LiDAR po…

Cited by 0SourceScholar
2024

Adaptive Correspondence Scoring for Unsupervised Medical Image Registration

ECCV 2024oral

"We propose an adaptive training scheme for unsupervised medical image registration. Existing methods rely on image reconstruction as the primary supervision signal. However, nuisance variables (e.g. noise and covisibility), violation of the Lambertian assumption in physical waves (e.g. ultrasound),…

2024

All-day Depth Completion

IROS 2024poster

We propose a method for depth estimation under different illumination conditions, i.e., day and night time. As photometry is uninformative in regions under low-illumination, we tackle the problem through a multi-sensor fusion approach, where we take as input an additional synchronized sparse point c…

Cited by 3SourcecodeScholar
2024

AugUndo: Scaling Up Augmentations for Monocular Depth Completion and Estimation

ECCV 2024poster

"Unsupervised depth completion and estimation methods are trained by minimizing reconstruction error. Block artifacts from resampling, intensity saturation, and occlusions are amongst the many undesirable by-products of common data augmentation schemes that affect image reconstruction quality, and t…

2024

Binding Touch to Everything: Learning Unified Multimodal Tactile Representations

CVPR 2024poster

The ability to associate touch with other modalities has huge implications for humans and computational systems. However multimodal learning with touch remains challenging due to the expensive data collection process and non-standardized sensor outputs. We introduce UniTouch a unified tactile model…

Cited by 53SourcePDFScholar
2024

Diffeomorphic Template Registration for Atmospheric Turbulence Mitigation

CVPR 2024highlight

We describe a method for recovering the irradiance underlying a collection of images corrupted by atmospheric turbulence. Since supervised data is often technically impossible to obtain assumptions and biases have to be imposed to solve this inverse problem and we choose to model them explicitly. Ra…

Cited by 5SourcePDFScholar
2024

On the Viability of Monocular Depth Pre-training for Semantic Segmentation

ECCV 2024poster

"The question of whether pre-training on geometric tasks is viable for downstream transfer to semantic tasks is important for two reasons, one practical and the other scientific. If the answer is positive, we may be able to reduce pre-training costs and bias from human annotators significantly. If t…

2024

RSA: Resolving Scale Ambiguities in Monocular Depth Estimators through Language Descriptions

NeurIPS 2024poster

We propose a method for metric-scale monocular depth estimation. Inferring depth from a single image is an ill-posed problem due to the loss of scale from perspective projection during the image formation process. Any scale chosen is a bias, typically stemming from training on a dataset; hence, exis…

2024

Sub-token ViT Embedding via Stochastic Resonance Transformers

ICML 2024poster

Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This representation trades spatial granularity for embedding dimensionality, and results in semantically rich but spatially coarsel…

2024

WorDepth: Variational Language Prior for Monocular Depth Estimation

CVPR 2024poster

Three-dimensional (3D) reconstruction from a single image is an ill-posed problem with inherent ambiguities i.e. scale. Predicting a 3D scene from text description(s) is similarly ill-posed i.e. spatial arrangements of objects described. We investigate the question of whether two inherently ambiguou…

2023

Depth Estimation From Camera Image and mmWave Radar Point Cloud

CVPR 2023poster

We present a method for inferring dense depth from a camera image and a sparse noisy radar point cloud. We first describe the mechanics behind mmWave radar point cloud formation and the challenges that it poses, i.e. ambiguous elevation and noisy depth and azimuth components that yields incorrect po…

Cited by 53SourcePDFScholar
2023

WeatherStream: Light Transport Automation of Single Image Deweathering

CVPR 2023poster

Today single image deweathering is arguably more sensitive to the dataset type, rather than the model. We introduce WeatherStream, an automatic pipeline capturing all real-world weather effects (rain, snow, and rain fog degradations), along with their clean image pairs. Previous state-of-the-art met…

Cited by 23SourcePDFScholar
2022

Monitored Distillation for Positive Congruent Depth Completion

ECCV 2022poster

"We propose a method to infer a dense depth map from a single image, its calibration, and the associated sparse point cloud. In order to leverage existing models (teachers) that produce putative depth maps, we propose an adaptive knowledge distillation approach that yields a positive congruent train…

2022

Not Just Streaks: Towards Ground Truth for Single Image Deraining

ECCV 2022poster

"We propose a large-scale dataset of real-world rainy and clean image pairs and a method to remove degradations, induced by rain streaks and rain accumulation, from the image. As there exists no real-world dataset for deraining, current state-of-the-art methods rely on synthetic data and thus are li…

2022

Stereoscopic Universal Perturbations Across Different Architectures and Datasets

CVPR 2022poster

We study the effect of adversarial perturbations of images on deep stereo matching networks for the disparity estimation task. We present a method to craft a single set of perturbations that, when added to any stereo image pair in a dataset, can fool a stereo network to significantly alter the perce…

Cited by 19PDFcodeScholar
2021

Stereopagnosia: Fooling Stereo Networks with Adversarial Perturbations

AAAI 2021technical

We study the effect of adversarial perturbations of images on the estimates of disparity by deep learning models trained for stereo. We show that imperceptible additive perturbations can significantly alter the disparity map, and correspondingly the perceived geometry of the scene. These perturbatio…

2020

Targeted Adversarial Perturbations for Monocular Depth Prediction

NeurIPS 2020poster

We study the effect of adversarial perturbations on the task of monocular depth prediction. Specifically, we explore the ability of small, imperceptible additive perturbations to selectively alter the perceived geometry of the scene. We show that such perturbations can not only globally re-scale the…

2019

Bilateral Cyclic Constraint and Adaptive Regularization for Unsupervised Monocular Depth Prediction

CVPR 2019poster

Supervised learning methods to infer (hypothesize) depth of a scene from a single image require costly per-pixel ground-truth. We follow a geometric approach that exploits abundant stereo imagery to learn a model to hypothesize scene structure without direct supervision. Although we train a network…

Cited by 113PDFcodeScholar