← Search

Andrew Markham

55 accepted papers

2026

Aurelius: Relation Aware Text-to-Audio Generation At Scale

ICLR 2026poster

We present Aurelius, a new framework that enables relation aware text-to-audio (TTA) generation research at scale. Given the lack of essential audio event and relation corpora, \emph{Aurelius} contributes a large-scale audio event corpus \emph{AudioEventSet} and another large-scale relation corpus \…

Cited by 0SourcecodeScholar
2026

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

ICRA 2026poster

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while …

2025

COOPERA: Continual Open-Ended Human-Robot Assistance

NeurIPS 2025spotlight

To understand and collaborate with humans, robots must account for individual human traits, habits, and activities over time. However, most robotic assistants lack these abilities, as they primarily focus on predefined tasks in structured environments and lack a human model to learn from. This work…

Cited by 0SourceScholar
2025

DiffRefine: Diffusion-based Proposal Specific Point Cloud Densification for Cross-Domain Object Detection

ICCV 2025poster

The robustness of 3D object detection in large-scale outdoor point clouds degrades significantly when deployed in an unseen environment due to domain shifts. To minimize the domain gap, existing works on domain adaptive detection focuses on several factors, including point density, object shape and…

Cited by 0SourcePDFScholar
2025

Manydepth2: Motion-Aware Self-Supervised Monocular Depth Estimation in Dynamic Scenes

RA-L 2025

Despite advancements in self-supervised monocular depth estimation, challenges persist in dynamic scenarios due to the dependence on assumptions about a static world. In this paper, we present Manydepth2, to achieve precise depth estimation for both dynamic objects and static backgrounds, all while

Cited by 18SourceScholar
2025

RiTTA: Modeling Event Relations in Text-to-Audio Generation

EMNLP 2025

Existing text-to-audio (TTA) generation methods have neither systematically explored audio event relation modeling, nor proposed any new framework to enhance this capability. In this work, we systematically study audio event relation modeling in TTA generation models. We first establish a benchmark

2025

Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments

NeurIPS 2025poster

Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as conditioning inputs. However, such clean audio examples are not a…

Cited by 0SourcecodeScholar
2024

Dusk Till Dawn: Self-supervised Nighttime Stereo Depth Estimation using Visual Foundation Models

ICRA 2024poster

Self-supervised depth estimation algorithms rely heavily on frame-warping relationships, exhibiting substantial performance degradation when applied in challenging circumstances, such as low-visibility and nighttime scenarios with varying illumination conditions. Addressing this challenge, we introd…

Cited by 4SourcecodeScholar
2024

Learning Continuous 3D Words for Text-to-Image Generation

CVPR 2024poster

Current controls over diffusion models (e.g. through text or ControlNet) for image generation fall short in recognizing abstract continuous attributes like illumination direction or non-rigid shape change. In this paper we present an approach for allowing users of text-to-image models to have fine-g…

2024

Learning Generalizable Manipulation Policy with Adapter-Based Parameter Fine-Tuning

IROS 2024

This study investigates the use of adapters in reinforcement learning for robotic skill generalization across multiple robots and tasks. Traditional methods are typically reliant on robot-specific retraining and face challenges such as efficiency and adaptability, particularly when scaling to robots

Cited by 5SourcecodeScholar
2024

Learning to Catch Reactive Objects with a Behavior Predictor

ICRA 2024poster

Tracking and catching moving objects is an important ability for robots in a dynamic world. Whilst some objects have highly predictable state evolution e.g., the ballistic trajectory of a tennis ball, reactive targets alter their behavior in response to motion of the manipulator. Reactive applicatio…

Cited by 2SourcecodeScholar
2024

SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound Classification

ICASSP 2024accepted

Efficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods typically demand extensively labeled audio datasets and have highl…

Cited by 0SourceScholar
2024

SoundCount: Sound Counting from Raw Audio with Dyadic Decomposition Neural Network

AAAI 2024technical

In this paper, we study an underexplored, yet important and challenging problem: counting the number of distinct sounds in raw audio characterized by a high degree of polyphonicity. We do so by systematically proposing a novel end-to-end trainable neural network~(which we call DyDecNet, consisting o…

Cited by 2SourcePDFScholar
2024

SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors

NeurIPS 2024poster

Current state-of-the-art spatial reasoning-enhanced VLMs are trained to excel at spatial visual question answering (VQA). However, we believe that higher-level 3D-aware tasks, such as articulating dynamic scene changes and motion planning, require a fundamental and explicit 3D understanding beyond c…

Cited by 5SourcePDFScholar
2024

Spherical Mask: Coarse-to-Fine 3D Point Cloud Instance Segmentation with Spherical Representation

CVPR 2024poster

Coarse-to-fine 3D instance segmentation methods show weak performances compared to recent Grouping-based Kernel-based and Transformer-based methods. We argue that this is due to two limitations: 1) Instance size overestimation by axis-aligned bounding box(AABB) 2) False negative error accumulation f…

2024

Towards Learning Group-Equivariant Features for Domain Adaptive 3D Detection

NeurIPS 2024poster

The performance of 3D object detection in large outdoor point clouds deteriorates significantly in an unseen environment due to the inter-domain gap. To address these challenges, most existing methods for domain adaptation harness self-training schemes and attempt to bridge the gap by focusing on a…

Cited by 0SourcePDFScholar
2024

WSCLoc: Weakly-Supervised Sparse-View Camera Relocalization via Radiance Field

IROS 2024poster

Despite the advancements in deep learning for camera relocalization tasks, obtaining ground truth pose labels required for the training process remains a costly endeavor. While current weakly supervised methods excel in lightweight label generation, their performance notably declines in scenarios wi…

Cited by 0SourceScholar
2024

ZeST: Zero-Shot Material Transfer from a Single Image

ECCV 2024poster

"We propose , a method for zero-shot material transfer to an object in the input image given a material exemplar image. leverages existing diffusion adapters to extract implicit material representation from the exemplar image. This representation is used to transfer the material using pre-trained in…

2023

3DMiner: Discovering Shapes from Large-Scale Unannotated Image Datasets

ICCV 2023poster

We present 3DMiner -- a pipeline for mining 3D shapes from challenging large-scale unannotated image datasets. Unlike other unsupervised 3D reconstruction methods, we assume that, within a large-enough dataset, there must exist images of objects with similar shapes but varying backgrounds, textures,…

Cited by 0PDFcodeScholar
2023

Decoupling Skill Learning from Robotic Control for Generalizable Object Manipulation

ICRA 2023poster

Recent works in robotic manipulation through reinforcement learning (RL) or imitation learning (IL) have shown potential for tackling a range of tasks e.g., opening a drawer or a cupboard. However, these techniques generalize poorly to unseen objects. We conjecture that this is due to the high-dimen…

Cited by 5SourcecodeScholar
2023

DynPoint: Dynamic Neural Point For View Synthesis

NeurIPS 2023poster

The introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario. To tackle t…

Cited by 18SourcePDFScholar
2023

Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion Estimation

NeurIPS 2023poster

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion estimates, we present an SE(3) equivariant architecture and a traini…

2023

RADA: Robust Adversarial Data Augmentation for Camera Localization in Challenging Conditions

IROS 2023poster

Camera localization is a fundamental problem for many applications in computer vision, robotics, and autonomy. Despite recent deep learning-based approaches, the lack of robustness in challenging conditions persists due to changes in appearance caused by texture-less planes, repeating structures, re…

Cited by 3SourcecodeScholar
2023

Sample, Crop, Track: Self-Supervised Mobile 3D Object Detection for Urban Driving LiDAR

ICRA 2023poster

Deep learning has led to great progress in the detection of mobile (i.e. movement-capable) objects in urban driving scenes in recent years. Supervised approaches typically require the annotation of large training sets; there has thus been great interest in leveraging weakly, semi- or self- supervise…

Cited by 2SourceScholar
2023

SoundSynp: Sound Source Detection from Raw Waveforms with Multi-Scale Synperiodic Filterbanks

AISTATS 2023poster

We propose synperiodic filter banks, a novel multi-scale learnable filter bank construction strategy that all filters are synchronized by their rotating periodicity. By synchronizing in a certain periodicity, we naturally get filters whose temporal length are reduced if they carry higher frequency r…

Cited by 8SourcePDFScholar
2022

DeepCIR: Insights into CIR-based Data-driven UWB Error Mitigation

IROS 2022poster

Ultra-Wide-Band (UWB) ranging sensors have been widely adopted for robotic navigation thanks to their extremely high bandwidth and hence high resolution. However, off-the-shelf devices may output ranges with significant errors in cluttered, severe non-line-of-sight (NLOS) environments. Recently, neu…

Cited by 9SourceScholar
2022

Meta-Sampler: Almost-Universal yet Task-Oriented Sampling for Point Clouds

ECCV 2022poster

"Sampling is a key operation in point-cloud task and acts to increase computational efficiency and tractability by discarding redundant points. Universal sampling algorithms (e.g., Farthest Point Sampling) work without modification across different tasks, models, and datasets, but by their very natu…

2022

No Pain, Big Gain: Classify Dynamic Point Cloud Sequences With Static Models by Fitting Feature-Level Space-Time Surfaces

CVPR 2022poster

Scene flow is a powerful tool for capturing the motion field of 3D point clouds. However, it is difficult to directly apply flow-based models to dynamic point cloud classification since the unstructured points make it hard or even impossible to efficiently and effectively trace point-wise correspond…

Cited by 29PDFcodeScholar
2022

Real-Time Hybrid Mapping of Populated Indoor Scenes using a Low-Cost Monocular UAV

IROS 2022poster

Unmanned aerial vehicles (UAVs) have been used for many applications in recent years, from urban search and rescue, to agricultural surveying, to autonomous underground mine exploration. However, deploying UAVs in tight, indoor spaces, especially close to humans, remains a challenge. One solution, w…

Cited by 4SourceScholar
2022

SQN: Weakly-Supervised Semantic Segmentation of Large-Scale 3D Point Clouds

ECCV 2022poster

"Labelling point clouds fully is highly time-consuming and costly. As larger point cloud datasets containing billions of points become more common, we ask whether the full annotation is even necessary, demonstrating that existing baselines designed under a fully annotated assumption only degrade sli…

2022

When the Sun Goes Down: Repairing Photometric Losses for All-Day Depth Estimation

CoRL 2022poster

Self-supervised deep learning methods for joint depth and ego-motion estimation can yield accurate trajectories without needing ground-truth training data. However, as they typically use photometric losses, their performance can degrade significantly when the assumptions these losses make (e.g. temp…

Cited by 27SourceScholar
2021

3D Motion Capture of an Unmodified Drone with Single-chip Millimeter Wave Radar

ICRA 2021poster

Accurate motion capture of aerial robots in 3D is a key enabler for autonomous operation in indoor environments such as warehouses or factories, as well as driving forward research in these areas. The most commonly used solutions at present are optical motion capture (e.g. VICON) and Ultrawide-band…

Cited by 32SourceScholar
2021

P2-Net: Joint Description and Detection of Local Features for Pixel and Point Matching

ICCV 2021poster

Accurately describing and detecting 2D and 3D keypoints is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint dete…

Cited by 62PDFcodeScholar
2021

SoundDet: Polyphonic Moving Sound Event Detection and Localization from Raw Waveform

ICML 2021spotlight

We present a new framework SoundDet, which is an end-to-end trainable and light-weight framework, for polyphonic moving sound event detection and localization. Prior methods typically approach this problem by preprocessing raw waveform into time-frequency representations, which is more amenable to p…

Cited by 36SourcePDFScholar
2021

SpinNet: Learning a General Surface Descriptor for 3D Point Cloud Registration

CVPR 2021poster

Extracting robust and general 3D local features is key to downstream tasks such as point cloud registration and reconstruction. Existing learning-based local descriptors are either sensitive to rotation transformations, or rely on classical handcrafted features which are neither general nor represen…

Cited by 380PDFcodeScholar
2021

Towards Semantic Segmentation of Urban-Scale 3D Point Clouds: A Dataset, Benchmarks and Challenges

CVPR 2021poster

An essential prerequisite for unleashing the potential of supervised deep learning algorithms in the area of 3D scene understanding is the availability of large-scale and richly annotated datasets. However, publicly available datasets are either in relatively small spatial scales or have limited sem…

Cited by 236PDFcodeScholar
2021

VMLoc: Variational Fusion For Learning-Based Multimodal Camera Localization

AAAI 2021technical

Recent learning-based approaches have achieved impressive results in the field of single-shot camera localization. However, how best to fuse multiple modalities (e.g., image and depth) and to deal with degraded or missing input are less well studied. In particular, we note that previous approaches t…

2020

DeepTIO: A Deep Thermal-Inertial Odometry With Visual Hallucination

RA-L 2020

Visual odometry shows excellent performance in a wide range of environments. However, in visually-denied scenarios (e.g. heavy smoke or darkness), pose estimates degrade or even fail. Thermal cameras are commonly used for perception and inspection when the environment has low visibility. However, th

Cited by 72SourceScholar
2020

Heart Rate Sensing with a Robot Mounted mmWave Radar

ICRA 2020poster

Heart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitiv…

Cited by 90SourceScholar
2020

RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds

CVPR 2020oral

We study the problem of efficient semantic segmentation for large-scale 3D point clouds. By relying on expensive sampling techniques or computationally heavy pre/post-processing steps, most existing approaches are only able to be trained and operate over small-scale point clouds. In this paper, we i…

Cited by 2143PDFcodeScholar
2020

SnapNav: Learning Mapless Visual Navigation with Sparse Directional Guidance and Visual Reference

ICRA 2020poster

Learning-based visual navigation still remains a challenging problem in robotics, with two overarching issues: how to transfer the learnt policy to unseen scenarios, and how to deploy the system on real robots. In this paper, we propose a deep neural network based visual navigation system, SnapNav.…

Cited by 17SourceScholar
2019

DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural Network

IROS 2019poster

Odometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches hav…

Cited by 61SourceScholar
2019

Distilling Knowledge From a Deep Pose Regressor Network

ICCV 2019poster

This paper presents a novel method to distill knowledge from a deep pose regressor network for efficient Visual Odometry (VO). Standard distillation relies on "dark knowledge" for successful knowledge transfer. As this knowledge is not available in pose regression and the teacher prediction is not a…

Cited by 133PDFScholar
2019

GANVO: Unsupervised Deep Monocular Visual Odometry and Depth Estimation with Generative Adversarial Networks

ICRA 2019poster

In the last decade, supervised deep learning approaches have been extensively employed in visual odometry (VO) applications, which is not feasible in environments where labelled data is not abundant. On the other hand, unsupervised deep learning approaches for localization and mapping in unknown env…

Cited by 199SourceScholar
2019

Learning Monocular Visual Odometry through Geometry-Aware Curriculum Learning

ICRA 2019poster

Inspired by the cognitive process of humans and animals, Curriculum Learning (CL) trains a model by gradually increasing the difficulty of the training data. In this paper, we study whether CL can be applied to complex geometry problems like estimating monocular Visual Odometry (VO). Unlike existing…

Cited by 58SourceScholar
2019

Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds

NeurIPS 2019spotlight

We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cl…

2019

Selective Sensor Fusion for Neural Visual-Inertial Odometry

CVPR 2019poster

Deep learning approaches for Visual-Inertial Odometry (VIO) have proven successful, but they rarely focus on incorporating robust fusion strategies for dealing with imperfect input sensory data. We propose a novel end-to-end selective sensor fusion framework for monocular VIO, which fuses monocular…

Cited by 192PDFcodeScholar
2018

DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks

ICRA 2018poster

Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditiona…

Cited by 10SourceScholar
2018

Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning

ICRA 2018poster

Deep Reinforcement Learning (DRL) has been applied successfully to many robotic applications. However, the large number of trials needed for training is a key issue. Most of existing techniques developed to improve training efficiency (e.g. imitation) target on general tasks rather than being tailor…

Cited by 105SourcecodeScholar
2017

GraphTinker: Outlier rejection and inlier injection for pose graph SLAM

IROS 2017poster

In pose graph Simultaneous Localization and Mapping (SLAM) systems, incorrect loop closures can seriously hinder optimizers from converging to correct solutions, significantly degrading both localization accuracy and map consistency. Therefore, it is crucial to enhance their robustness in the presen…

Cited by 14SourceScholar
2017

VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization

CVPR 2017poster

Machine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of th…

Cited by 325PDFcodeScholar