← Search

Sai-Kit Yeung

42 accepted papers

2026

MaskGuide: Efficient Distillation for Deployable Lightweight Segmentation in Marine Environments

RA-L 2026

The growing demand for efficient image segmentation in marine ecological studies is currently constrained by two key factors: the high computational requirements of models such as the Segment Anything Model (SAM) and the degraded accuracy of lightweight models in underwater environments. To overcome

Cited by 0SourceScholar
2025

Align3R: Aligned Monocular Depth Estimation for Dynamic Videos

CVPR 2025highlight

Recent developments in monocular depth estimation methods enable high-quality depth estimation of single-view images but fail to estimate consistent video depth across different frames. Recent works address this problem by applying a video diffusion model to generate video depth conditioned on the i…

Cited by 14SourcePDFScholar
2025

CoralSRT: Revisiting Coral Reef Semantic Segmentation by Feature Rectification via Self-supervised Guidance

ICCV 2025poster

We investigate coral reef semantic segmentation, in which coral reefs are governed by multifaceted factors, like genes, environmental changes, and internal interactions. Unlike segmenting structural units/instances, which are predictable and follow a set pattern, also referred to as commonsense or p…

Cited by 0SourcePDFScholar
2025

SC-OmniGS: Self-Calibrating Omnidirectional Gaussian Splatting

ICLR 2025poster

360-degree cameras streamline data collection for radiance field 3D reconstruction by capturing comprehensive scene data. However, traditional radiance field methods do not address the specific challenges inherent to 360-degree images. We present SC-OmniGS, a novel self-calibrating omnidirectional G…

Cited by 0SourcePDFScholar
2025

TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels

NeurIPS 2025poster

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall short in separating the camera motion from foreground dynamic…

Cited by 0SourceScholar
2024

360Loc: A Dataset and Benchmark for Omnidirectional Visual Localization with Cross-device Queries

CVPR 2024poster

Portable 360^\circ cameras are becoming a cheap and efficient tool to establish large visual databases. By capturing omnidirectional views of a scene these cameras could expedite building environment models that are essential for visual localization. However such an advantage is often overlooked due…

2024

CoralSCOP: Segment any COral Image on this Planet

CVPR 2024highlight

Underwater visual understanding has recently gained increasing attention within the computer vision community for studying and monitoring underwater ecosystems. Among these coral reefs play an important and intricate role often referred to as the rainforests of the sea due to their rich biodiversity…

Cited by 8SourcePDFScholar
2024

Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset Updates

CVPR 2024poster

Neural radiance field (NeRF) is an emerging technique for 3D scene reconstruction and modeling. However current NeRF-based methods are limited in the capabilities of adding or removing objects. This paper fills the aforementioned gap by proposing a new language-driven method for object manipulation…

2024

MarineInst: A Foundation Model for Marine Image Analysis with Instance Visual Description

ECCV 2024oral

"Recent foundation models trained on a tremendous scale of data have shown great promise in a wide range of computer vision tasks and application domains. However, less attention has been paid to the marine realms, which in contrast cover the majority of our blue planet. The scarcity of labeled data…

Cited by 8SourcePDFScholar
2024

Photo-SLAM: Real-time Simultaneous Localization and Photorealistic Mapping for Monocular Stereo and RGB-D Cameras

CVPR 2024poster

The integration of neural rendering and the SLAM system recently showed promising results in joint localization and photorealistic view reconstruction. However existing methods fully relying on implicit representations are so resource-hungry that they cannot run on portable devices which deviates fr…

2024

StyleCity: Large-Scale 3D Urban Scenes Stylization

ECCV 2024poster

"Creating large-scale virtual urban scenes with variant styles is inherently challenging. To facilitate prototypes of virtual production and bypass the need for complex materials and lighting setups, we introduce the first vision-and-text-driven texture stylization system for large-scale urban scene…

2023

360VOT: A New Benchmark Dataset for Omnidirectional Visual Object Tracking

ICCV 2023poster

360deg images can provide an omnidirectional field of view which is important for stable and long-term scene perception. In this paper, we explore 360deg images for visual object tracking and perceive new challenges caused by large distortion, stitching artifacts, and other unique attributes of 360d…

Cited by 10PDFcodeScholar
2023

A Novel Platform to Control Biofouling in Pearl Oysters Cultivation

ICRA 2023poster

This paper presents a simple yet effective design of a platform to automate the task of shellfish aquaculture, specifically pearl oysters. Compared to traditional methods, our platform can eliminate the tedious task of cleaning the pearl oysters due to fouling. Inspired by the low and high tide char…

Cited by 1SourceScholar
2023

CompUDA: Compositional Unsupervised Domain Adaptation for Semantic Segmentation Under Adverse Conditions

IROS 2023poster

In autonomous driving, performing robust semantic segmentation under adverse weather conditions is a long-standing challenge. Imperfect camera observations under adverse conditions result in images with reduced visibility, which hinders label annotation and semantic scene understanding based on thes…

Cited by 5SourcecodeScholar
2023

Conditional 360-degree Image Synthesis for Immersive Indoor Scene Decoration

ICCV 2023poster

In this paper, we address the problem of conditional scene decoration for 360deg images. Our method takes a 360deg background photograph of an indoor scene and generates decorated images of the same scene in the panorama view. To do this, we develop a 360-aware object layout generator that learns la…

Cited by 8PDFcodeScholar
2023

Cross-Domain Autonomous Driving Perception Using Contrastive Appearance Adaptation

IROS 2023poster

Addressing domain shifts for complex perception tasks in autonomous driving has long been a challenging problem. In this paper, we show that existing domain adaptation methods pay little attention to the content mismatch issue between source and target domains, thus weakening the domain adaptation p…

Cited by 1SourceScholar
2022

360ST-Mapping: An Online Semantics-Guided Topological Mapping Module for Omnidirectional Visual SLAM

IROS 2022poster

As an abstract representation of the environment structure, a topological map has advantageous properties for path-planning and navigation. Here we proposed an online topological mapping method, 360ST-Mapping, using omnidirectional vision. The 360° field-of-view allows the agent to obtain consistent…

Cited by 7SourceScholar
2022

Neural Scene Decoration from a Single Photograph

ECCV 2022poster

"Furnishing and rendering indoor scenes has been a long-standing task for interior design, where artists create a conceptual design for the space, build a 3D model of the space, decorate, and then perform rendering. Although the task is important, it is tedious and requires tremendous effort. In thi…

2022

RFNet-4D: Joint Object Reconstruction and Flow Estimation from 4D Point Clouds

ECCV 2022poster

"Object reconstruction from 3D point clouds has achieved impressive progress in the computer vision and computer graphics research field. However, reconstruction from time-varying point clouds (a.k.a. 4D point clouds) is generally overlooked. In this paper, we propose a new network architecture, nam…

2022

SiamX: An Efficient Long-term Tracker Using Cross-level Feature Correlation and Adaptive Tracking Scheme

ICRA 2022poster

Siamese network based trackers have achieved significant progress in visual object tracking. For the sake of speed, they mainly rely on offline training to learn a mono-level feature correlation between a target template and a search region. During the tracking period, they use a fixed strategy to i…

Cited by 3SourceScholar
2020

Dual-SLAM: A framework for robust single camera navigation

IROS 2020poster

SLAM (Simultaneous Localization And Mapping) seeks to provide a moving agent with real-time self-localization. To achieve real-time speed, SLAM incrementally propagates position estimates. This makes SLAM fast but also makes it vulnerable to local pose estimation failures. As local pose estimation i…

Cited by 17SourceScholar
2020

SideInfNet: A Deep Neural Network for Semi-Automatic Semantic Segmentation with Side Information

ECCV 2020poster

Fully-automatic execution is the ultimate goal for many Computer Vision applications. However, this objective is not always realistic in tasks associated with high failure costs, such as medical applications. For these tasks, semi-automatic methods allowing minimal effort from users to guide compute…

Cited by 6SourcePDFScholar
2019

Force-based Heterogeneous Traffic Simulation for Autonomous Vehicle Testing

ICRA 2019poster

Recent failures in real-world self-driving tests have suggested a paradigm shift from directly learning in real-world roads to building a high-fidelity driving simulator as an alternative, effective, and safe tool to handle intricate traffic environments in urban areas. To date, traffic simulation c…

Cited by 44SourceScholar
2019

JSIS3D: Joint Semantic-Instance Segmentation of 3D Point Clouds With Multi-Task Pointwise Networks and Multi-Value Conditional Random Fields

CVPR 2019oral

Deep learning techniques have become the to-go models for most vision-related tasks on 2D images. However, their power has not been fully realised on several tasks in 3D space, e.g., 3D scene understanding. In this work, we jointly address the problems of semantic and instance segmentation of 3D poi…

Cited by 262PDFcodeScholar
2019

Revisiting Point Cloud Classification: A New Benchmark Dataset and Classification Model on Real-World Data

ICCV 2019oral

Deep learning techniques for point cloud data have demonstrated great potentials in solving classical problems in 3D computer vision such as 3D object classification and segmentation. Several recent 3D object classification methods have reported state-of-the-art performance on CAD model datasets suc…

Cited by 1063PDFcodeScholar
2019

ShellNet: Efficient Point Cloud Convolutional Neural Networks Using Concentric Shells Statistics

ICCV 2019oral

Deep learning with 3D data has progressed significantly since the introduction of convolutional neural networks that can handle point order ambiguity in point cloud data. While being able to achieve good accuracies in various scene understanding tasks, previous methods often have low training speed…

Cited by 481PDFcodeScholar
2018

Uncalibrated Photometric Stereo Under Natural Illumination

CVPR 2018poster

This paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve ea…

Cited by 43SourcePDFScholar
2018

Urban Zoning Using Higher-Order Markov Random Fields on Multi-View Imagery Data

ECCV 2018poster

Urban zoning enables various applications in land use analysis and urban planning. As cities evolve, it is important to constantly update the zoning maps of cities to reflect urban pattern changes. This paper proposes a method for automatic urban zoning using higher-order Markov random fields (HO-MR…

Cited by 21SourcePDFScholar
2017

GMS: Grid-based Motion Statistics for Fast, Ultra-Robust Feature Correspondence

CVPR 2017poster

Incorporating smoothness constraints into feature matching is known to enable ultra-robust matching. However, such formulations are both complex and slow, making them unsuitable for video applications. This paper proposes GMS (Grid-based Motion Statistics), a simple means of encapsulating motion smo…

Cited by 862PDFcodeScholar
2016

A Benchmark Dataset and Evaluation for Non-Lambertian and Uncalibrated Photometric Stereo

CVPR 2016poster

Recent progress on photometric stereo extends the technique to deal with general materials and unknown illumination conditions. However, due to the lack of suitable benchmark data with ground truth shapes (normals), quantitative comparison and evaluation is difficult to achieve. In this paper, we fi…

Cited by 358PDFScholar