← Search

Ronald Clark

25 accepted papers

2026

ForestVO: Enhancing Visual Odometry in Forest Environments through ForestGlue

ICRA 2026poster

Recent advancements in visual odometry systems have improved autonomous navigation, yet challenges persist in complex environments like forests, where dense foliage, variable lighting, and repetitive textures compromise the accuracy of feature correspondences. To address these challenges, we introdu…

2025

Cloud4D: Estimating Cloud Properties at a High Spatial and Temporal Resolution

NeurIPS 2025spotlight

There has been great progress in improving numerical weather prediction and climate models using machine learning. However, most global models act at a kilometer-scale, making it challenging to model individual clouds and factors such as extreme precipitation, wind gusts, turbulence, and surface irr…

Cited by 0SourceScholar
2025

ForestVO: Enhancing Visual Odometry in Forest Environments Through ForestGlue

RA-L 2025

Recent advancements in visual odometry systems have improved autonomous navigation, yet challenges persist in complex environments like forests, where dense foliage, variable lighting, and repetitive textures compromise the accuracy of feature correspondences. To address these challenges, we introdu

Cited by 7SourceScholar
2025

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

NeurIPS 2025poster

Although diffusion-based models can generate high-quality and high-resolution video sequences from textual or image inputs, they lack explicit integration of geometric cues when controlling scene lighting and visual appearance across frames. To address this limitation, we propose IllumiCraft, an end…

Cited by 0SourceScholar
2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

NeurIPS 2025poster

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as `safety' and `robustness' requires strong construct validity, that is, having measures t…

Cited by 0SourceScholar
2025

Olympus: A Universal Task Router for Computer Vision Tasks

CVPR 2025highlight

We introduce Olympus, a new approach that transforms Multimodal Large Language Models (MLLMs) into a unified framework capable of handling a wide array of computer vision tasks. Utilizing a controller MLLM, Olympus delegates over 20 specialized tasks across images, videos, and 3D objects to dedicate…

2024

Instant Uncertainty Calibration of NeRFs Using a Meta-Calibrator

ECCV 2024poster

"Neural Radiance Fields (NeRFs) have markedly improved novel view synthesis, but accurate uncertainty quantification in their image predictions remains an open problem. The prevailing methods for estimating uncertainty, including the state-of-the-art Density-aware NeRF Ensembles (DANE) [?], quantify…

2023

Learning Tethered Perching for Aerial Robots

ICRA 2023poster

Aerial robots have a wide range of applications, such as collecting data in hard-to-reach areas. This requires the longest possible operation time. However, because currently available commercial batteries have limited specific energy of roughly 300 W h kg-1, a drone's flight time is a bottleneck fo…

Cited by 16SourceScholar
2020

Anomaly Detection for Time Series Using VAE-LSTM Hybrid Model

ICASSP 2020accepted

In this work, we propose a VAE-LSTM hybrid model as an unsupervised approach for anomaly detection in time series. Our model utilizes both a VAE module for forming robust local features over short windows and a LSTM module for estimating the long term correlation in the series on top of the features…

Cited by 0SourceScholar
2020

DeepFactors: Real-Time Probabilistic Dense Monocular SLAM

RA-L 2020

The ability to estimate rich geometry and camera motion from monocular imagery is fundamental to future interactive robotics and augmented reality applications. Different approaches have been proposed that vary in scene geometry representation (sparse landmarks, dense maps), the consistency metric u

Cited by 225SourcecodeScholar
2020

Scalable Uncertainty for Computer Vision With Functional Variational Inference

CVPR 2020poster

As Deep Learning continues to yield successful applications in Computer Vision, the ability to quantify all forms of uncertainty is a paramount requirement for its safe and reliable deployment in the real-world. In this work, we leverage the formulation of variational inference in function space, wh…

Cited by 26PDFScholar
2020

Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction

ICRA 2020poster

The best way to combine the results of deep learning with standard 3D reconstruction pipelines remains an open problem. While systems that pass the output of traditional multi-view stereo approaches to a network for regularisation or refinement currently seem to get the best results, it may be prefe…

Cited by 3SourceScholar
2019

Learning Meshes for Dense Visual SLAM

ICCV 2019poster

Estimating motion and surrounding geometry of a moving camera remains a challenging inference problem. From an information theoretic point of view, estimates should get better as more information is included, such as is done in dense SLAM, but this is strongly dependent on the validity of the underl…

Cited by 28PDFScholar
2019

Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds

NeurIPS 2019spotlight

We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cl…

2018

CodeSLAM — Learning a Compact, Optimisable Representation for Dense Visual SLAM

CVPR 2018poster

The representation of geometry in real-time 3D perception systems continues to be a critical research issue. Dense maps capture complete surface shape and can be augmented with semantic labels, but their high dimensionality makes them computationally costly to store and process, and unsuitable for r…

Cited by 463SourcePDFScholar
2018

Learning to Solve Nonlinear Least Squares for Monocular Stereo

ECCV 2018poster

Sum-of-squares objective functions are very popular in computer vision algorithms. However, these objective functions are not always easy to optimize. The underlying assumptions made by solvers are often not satisfied and many problems are inherently ill-posed. In this paper, we propose a neural non…

Cited by 100SourcePDFScholar
2017

DeepVO: Towards end-to-end visual odometry with deep Recurrent Convolutional Neural Networks

ICRA 2017poster

This paper studies monocular visual odometry (VO) problem. Most of existing VO algorithms are developed under a standard pipeline including feature extraction, feature matching, motion estimation, local optimisation, etc. Although some of them have demonstrated superior performance, they usually nee…

Cited by 1123SourceScholar
2017

VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization

CVPR 2017poster

Machine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of th…

Cited by 325PDFcodeScholar
2016

Keyframe based large-scale indoor localisation using geomagnetic field and motion pattern

IROS 2016poster

This paper studies indoor localisation problem by using low-cost and pervasive sensors. Most of existing indoor localisation algorithms rely on camera, laser scanner, floor plan or other pre-installed infrastructure to achieve sub-meter or sub-centimetre localisation accuracy. However, in some circu…

Cited by 69SourceScholar