← Search

Ian Reid

93 accepted papers

2026

Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models

AAAI 2026technical

Recent advancements in Large Video Language Models (LVLMs) have highlighted their potential for multi-modal understanding, yet evaluating their factual grounding in videos remains a critical unsolved challenge. To address this gap, we introduce Video SimpleQA, the first comprehensive benchmark tailo

Cited by 0SourcePDFScholar
2025

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

CVPR 2025poster

Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we…

2025

FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion Generation

NeurIPS 2025poster

Diffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net…

Cited by 0SourcecodeScholar
2025

Hier-SLAM: Scaling-Up Semantics in SLAM with a Hierarchically Categorical Gaussian Splatting

ICRA 2025

We propose Hier-SLAM, a semantic 3D Gaussian Splatting SLAM method featuring a novel hierarchical categorical representation, which enables accurate global 3D semantic mapping, scaling-up capability, and explicit semantic label prediction in the 3D world. The parameter usage in semantic SLAM systems

Cited by 19SourcecodeScholar
2025

Learning from 10 Demos: Generalisable and Sample-Efficient Policy Learning with Oriented Affordance Frames

CoRL 2025poster

Imitation learning has unlocked the potential for robots to exhibit highly dexterous behaviours. However, it still struggles with long-horizon, multi-object tasks due to poor sample efficiency and limited generalisation. Existing methods require a substantial number of demonstrations to cover possib…

Cited by 0SourcecodeScholar
2025

ObjectReact: Learning Object-Relative Control for Visual Navigation

CoRL 2025poster

Visual navigation using only a single camera and a topological map has recently become an appealing alternative to methods that require additional sensors and 3D maps. This is typically achieved through an "image-relative" approach to estimating control from a given pair of current observation and s…

Cited by 0SourceScholar
2025

TANGO: Traversability-Aware Navigation with Local Metric Control for Topological Goals

ICRA 2025

Visual navigation in robotics traditionally relies on globally-consistent 3D maps or learned controllers, which can be computationally expensive and difficult to generalize across diverse environments. In this work, we present a novel RGB-only, object-level topometric navigation pipeline that enable

Cited by 4SourcecodeScholar
2024

GaussCtrl: Multi-View Consistent Text-Driven 3D Gaussian Splatting Editing

ECCV 2024poster

"We propose , a text-driven method to edit a 3D scene reconstructed by the 3D Gaussian Splatting (3DGS). Our method first renders a collection of images by using the 3DGS and edits them by using a pre-trained 2D diffusion model (ControlNet) based on the input prompt, which is then used to optimise t…

2024

ItTakesTwo: Leveraging Peer Representations for Semi-supervised LiDAR Semantic Segmentation

ECCV 2024poster

"The costly and time-consuming annotation process to produce large training sets for modelling semantic LiDAR segmentation methods has motivated the development of semi-supervised learning (SSL) methods. However, such SSL approaches often concentrate on employing consistency learning only for indivi…

2024

JRDB-PanoTrack: An Open-world Panoptic Segmentation and Tracking Robotic Dataset in Crowded Human Environments

CVPR 2024poster

Autonomous robot systems have attracted increasing research attention in recent years where environment understanding is a crucial step for robot navigation human-robot interaction and decision. Real-world robot systems usually collect visual data from multiple sensors and are required to recognize…

Cited by 2SourcePDFScholar
2024

Motion Mamba: Efficient and Long Sequence Motion Generation

ECCV 2024poster

"Human motion generation stands as a significant pursuit in generative computer vision, while achieving long-sequence and efficient motion generation remains challenging. Recent advancements in state space models (SSMs), notably Mamba, have showcased considerable promise in long sequence modeling wi…

2024

RoboHop: Segment-based Topological Map Representation for Open-World Visual Navigation

ICRA 2024poster

Mapping is crucial for spatial reasoning, planning and robot navigation. Existing approaches range from metric, which require precise geometry-based optimization, to purely topological, where image-as-node based graphs lack explicit object-level reasoning and interconnectivity. In this paper, we pro…

Cited by 19SourcecodeScholar
2023

Residual Pattern Learning for Pixel-Wise Out-of-Distribution Detection in Semantic Segmentation

ICCV 2023poster

Semantic segmentation models classify pixels into a set of known ("in-distribution") visual classes. When deployed in an open world, the reliability of these models depends on their ability to not only classify in-distribution pixels but also to detect out-of-distribution (OoD) pixels. Historicall…

Cited by 45PDFcodeScholar
2023

SayPlan: Grounding Large Language Models using 3D Scene Graphs for Scalable Robot Task Planning

CoRL 2023oral

Large language models (LLMs) have demonstrated impressive results in developing generalist planning agents for diverse tasks. However, grounding these plans in expansive, multi-floor, and multi-room environments presents a significant challenge for robotics. We introduce SayPlan, a scalable approach…

Cited by 315SourcecodeScholar
2022

Asynchronous Optimisation for Event-based Visual Odometry

ICRA 2022poster

Event cameras open up new possibilities for robotic perception due to their low latency and high dynamic range. On the other hand, developing effective event-based vision algorithms that fully exploit the beneficial properties of event cameras remains work in progress. In this paper, we focus on eve…

Cited by 16SourceScholar
2022

Autonomy and Perception for Space Mining

ICRA 2022poster

Future Moon bases will likely be constructed using resources mined from the surface of the Moon. The difficulty of maintaining a human workforce on the Moon and communications lag with Earth means that mining will need to be conducted using collaborative robots with a high degree of autonomy. In thi…

Cited by 8SourceScholar
2022

JRDB-Act: A Large-Scale Dataset for Spatio-Temporal Action, Social Group and Activity Detection

CVPR 2022poster

The availability of large-scale video action understanding datasets has facilitated advances in the interpretation of visual scenes containing people. However, learning to recognise human actions and their social interactions in an unconstrained real-world environment comprising numerous people, wit…

Cited by 46PDFScholar
2022

You Only Cut Once: Boosting Data Augmentation with a Single Cut

ICML 2022spotlight

We present You Only Cut Once (YOCO) for performing data augmentations. YOCO cuts one image into two pieces and performs data augmentations individually within each piece. Applying YOCO improves the diversity of the augmentation per sample and encourages neural networks to recognize objects from part…

2021

MOLTR: Multiple Object Localization, Tracking and Reconstruction From Monocular RGB Videos

RA-L 2021

Semantic aware reconstruction is more advantageous than geometric-only reconstruction for future robotic and AR/VR applications because it represents not only where things are, but also what things are. Object-centric mapping is a task to build an object-level reconstruction where objects are separa

Cited by 25SourceScholar
2021

ODAM: Object Detection, Association, and Mapping Using Posed RGB Video

ICCV 2021poster

Localizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system…

Cited by 34PDFcodeScholar
2021

Rotation Coordinate Descent for Fast Globally Optimal Rotation Averaging

CVPR 2021poster

Under mild conditions on the noise level of the measurements, rotation averaging satisfies strong duality, which enables global solutions to be obtained via semidefinite programming (SDP) relaxation. However, generic solvers for SDP are rather slow in practice, even on rotation averaging instances o…

Cited by 19PDFcodeScholar
2021

TRiPOD: Human Trajectory and Pose Dynamics Forecasting in the Wild

ICCV 2021poster

Joint forecasting of human trajectory and pose dynamics is a fundamental building block of various applications ranging from robotics and autonomous driving to surveillance systems. Predicting body dynamics requires capturing subtle information embedded in the humans' interactions with each other an…

Cited by 64PDFScholar
2020

Depth Based Semantic Scene Completion With Position Importance Aware Loss

RA-L 2020

Semantic scene completion (SSC) refers to the task of inferring the 3D semantic segmentation of a scene while simultaneously completing the 3D shapes. We propose PALNet, a novel hybrid network for SSC based on single depth. PALNet utilizes a two-stream network to extract both 2D and 3D features from

Cited by 71SourcecodeScholar
2020

Joint Learning of Social Groups, Individuals Action and Sub-group Activities in Videos

ECCV 2020poster

Individuals Action and Sub-group Activities in Videos","The state-of-the art solutions for human activity understanding from a video stream formulate the task as a spatio-temporal problem which requires joint localization of all individuals in the scene and classification of their actions or group a…

2020

Meta Learning with Differentiable Closed-form Solver for Fast Video Object Segmentation

IROS 2020poster

Video object segmentation plays a vital role to many robotic tasks, beyond the satisfied accuracy, quickly adapt to the new scenario with very limited annotations and conduct a quick inference are also important. In this paper, we are specifically concerned with the task of fast segmenting all pixel…

Cited by 14SourceScholar
2020

SG-VAE: Scene Grammar Variational Autoencoder to generate new indoor scenes

ECCV 2020poster

Deep generative models have been used in recent years to learn coherent latent representations in order to synthesize high-quality images. In this work, we propose a neural network to learn a generative model for sampling consistent indoor scene layouts. Our method learns the co-occurrences, and app…

Cited by 56SourcePDFScholar
2020

Socially and Contextually Aware Human Motion and Pose Forecasting

RA-L 2020

Smooth and seamless robot navigation while interacting with humans depends on predicting human movements. Forecasting such human dynamics often involves modeling human trajectories (global motion) or detailed body joint movements (local motion). Prior work typically tackled local and global human mo

Cited by 94SourceScholar
2020

Training Quantized Neural Networks With a Full-Precision Auxiliary Module

CVPR 2020oral

In this paper, we seek to tackle a challenge in training low-precision networks: the notorious difficulty in propagating gradient through a low-precision network due to the non-differentiable quantization function. We propose a solution by training the low-precision network with a full-precision aux…

Cited by 95PDFScholar
2020

Visual Odometry Revisited: What Should Be Learnt?

ICRA 2020poster

In this work we present a monocular visual odometry (VO) algorithm which leverages geometry-based methods and deep learning. Most existing VO/SLAM systems with superior performance are based on geometry and have to be carefully designed for different application scenarios. Moreover, most monocular s…

Cited by 230SourcecodeScholar
2019

A Theoretically Sound Upper Bound on the Triplet Loss for Improving the Efficiency of Deep Distance Metric Learning

CVPR 2019poster

We propose a method that substantially improves the efficiency of deep distance metric learning based on the optimization of the triplet loss function. One epoch of such training process based on a na"ive optimization of the triplet loss function has a run-time complexity O(N^3), where N is the numb…

Cited by 77PDFScholar
2019

Attention-Guided Network for Ghost-Free High Dynamic Range Imaging

CVPR 2019poster

Ghosting artifacts caused by moving objects or misalignments is a key challenge in high dynamic range (HDR) imaging for dynamic scenes. Previous methods first register the input low dynamic range (LDR) images using optical flow before merging them, which are error-prone and cause ghosts in results.…

Cited by 342PDFScholar
2019

Fast Neural Architecture Search of Compact Semantic Segmentation Models via Auxiliary Cells

CVPR 2019poster

Automated design of neural network architectures tailored for a specific task is an extremely promising, albeit inherently difficult, avenue to explore. While most results in this domain have been achieved on image classification and language modelling problems, here we concentrate on dense per-pixe…

Cited by 196PDFcodeScholar
2019

Generalized Intersection Over Union: A Metric and a Loss for Bounding Box Regression

CVPR 2019poster

Intersection over Union (IoU) is the most popular evaluation metric used in the object detection benchmarks. However, there is a gap between optimizing the commonly used distance losses for regressing the parameters of a bounding box and maximizing this metric value. The optimal objective for a metr…

Cited by 6789PDFcodeScholar
2019

RGBD Based Dimensional Decomposition Residual Network for 3D Semantic Scene Completion

CVPR 2019poster

RGB images differentiate from depth as they carry more details about the color and texture information, which can be utilized as a vital complement to depth for boosting the performance of 3D semantic scene completion (SSC). SSC is composed of 3D shape completion (SC) and semantic scene labeling whi…

Cited by 97PDFScholar
2019

Real-Time Joint Semantic Segmentation and Depth Estimation Using Asymmetric Annotations

ICRA 2019poster

Deployment of deep learning models in robotics as sensory information extractors can be a daunting task to handle, even using generic GPU cards. Here, we address three of its most prominent hurdles, namely, i) the adaptation of a single model to perform multiple tasks at once (in this work, we consi…

Cited by 169SourcecodeScholar
2019

Scalable Place Recognition Under Appearance Change for Autonomous Driving

ICCV 2019oral

A major challenge in place recognition for autonomous driving is to be robust against appearance changes due to short-term (e.g., weather, lighting) and long-term (seasons, vegetation growth, etc.) environmental variations. A promising solution is to continuously accumulate images to maintain an ade…

Cited by 82PDFScholar
2019

Self-supervised Learning for Single View Depth and Surface Normal Estimation

ICRA 2019poster

In this work we present a self-supervised learning framework to simultaneously train two Convolutional Neural Networks (CNNs) to predict depth and surface normals from a single image. In contrast to most existing frameworks which represent outdoor scenes as fronto-parallel planes at piece-wise smoot…

Cited by 38SourceScholar
2019

Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks

NeurIPS 2019poster

Predicting the future trajectories of multiple interacting pedestrians in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to security and surveillance. This problem is compounded by the presence of soc…

Cited by 838SourcePDFScholar
2019

Structured Binary Neural Networks for Accurate Image Classification and Semantic Segmentation

CVPR 2019poster

In this paper, we propose to train convolutional neural networks (CNNs) with both binarized weights and activations, leading to quantized models specifically for mobile devices with limited power capacity and computation resources. By assuming the same architecture to full-precision networks, previo…

Cited by 192PDFScholar
2019

Unsupervised Scale-consistent Depth and Ego-motion Learning from Monocular Video

NeurIPS 2019poster

Recent work has shown that CNN-based depth and ego-motion estimators can be learned using unlabelled monocular videos. However, the performance is limited by unidentified moving objects that violate the underlying static scene assumption in geometric image reconstruction. More significantly, due to…

2018

Addressing Challenging Place Recognition Tasks Using Generative Adversarial Networks

ICRA 2018poster

Place recognition is an essential component of Simultaneous Localization And Mapping (SLAM). Under severe appearance change, reliable place recognition is a difficult perception task since the same place is perceptually very different in the morning, at night, or over different seasons. This work ad…

Cited by 45SourceScholar
2018

AffordanceNet: An End-to-End Deep Learning Approach for Object Affordance Detection

ICRA 2018poster

We propose AffordanceNet, a new deep learning approach to simultaneously detect multiple objects and their affordances from RGB images. Our AffordanceNet has two branches: an object detection branch to localize and classify the object, and an affordance detection branch to assign each pixel in the o…

Cited by 366SourcecodeScholar
2018

Are You Talking to Me? Reasoned Visual Dialog Generation Through Adversarial Learning

CVPR 2018poster

The Visual Dialogue task requires an agent to engage in a conversation about an image with a human. It represents an extension of the Visual Question Answering task in that the agent needs to answer a question about an image, but it needs to do so in light of the previous dialogue that has taken pl…

Cited by 148SourcePDFScholar
2018

Bayesian Semantic Instance Segmentation in Open Set World

ECCV 2018poster

This paper addresses the semantic instance segmentation task in the open-set conditions, where input images can contain known and unknown object classes. The training process of existing semantic instance segmentation methods requires annotation masks for all object instances, which is expensive to…

2018

Bootstrapping the Performance of Webly Supervised Semantic Segmentation

CVPR 2018poster

Fully supervised methods for semantic segmentation require pixel-level class masks to train, the creation of which are expensive in terms of manual labour and time. In this work, we focus on weak supervision, developing a method for training a high-quality pixel-level classifier for semantic segment…

2018

Deep Regression Tracking with Shrinkage Loss

ECCV 2018poster

Regression trackers directly learn a mapping from regularly dense samples of target objects to soft labels, which are usually generated by a Gaussian function, to estimate target positions. Due to the potential for fast-tracking and easy implementation, regression trackers have received increasing a…

2018

Efficient Dense Point Cloud Object Reconstruction using Deformation Vector Fields

ECCV 2018poster

Most existing CNN-based methods for single-view 3D object reconstruction represent a 3D object as either a 3D voxel occupancy grid or multiple depth-mask image pairs. However, these representations are inefficient since empty voxels or background pixels are wasteful. We propose a novel approach that…

Cited by 49SourcePDFScholar
2018

Just-in-Time Reconstruction: Inpainting Sparse Maps Using Single View Depth Predictors as Priors

ICRA 2018poster

We present “just-in-time reconstruction” as realtime image-guided inpainting of a map with arbitrary scale and sparsity to generate a fully dense depth map for the image. In particular, our goal is to inpaint a sparse map - obtained from either a monocular visual SLAM system or a sparse sensor - usi…

Cited by 37SourceScholar
2018

Multi-modal Cycle-consistent Generalized Zero-Shot Learning

ECCV 2018poster

In generalized zero shot learning (GZSL), the set of classes are split into seen and unseen classes, where training relies on the semantic features of the seen and unseen classes and the visual representations of only the seen classes, while testing uses the visual representations of the seen and un…

2018

Parallel Attention: A Unified Framework for Visual Object Discovery Through Dialogs and Queries

CVPR 2018poster

Recognising objects according to a pre-defined fixed set of class labels has been well studied in the Computer Vision. There are a great many practical applications where the subjects that may be of interest are not known beforehand, or so easily delineated, however. In many of these cases natural l…

Cited by 158SourcePDFScholar
2018

SceneCut: Joint Geometric and Object Segmentation for Indoor Scenes

ICRA 2018poster

This paper presents SceneCut, a novel approach to jointly discover previously unseen objects and non-object surfaces using a single RGB-D image. SceneCut's joint reasoning over scene semantics and geometry allows a robot to detect and segment object instances in complex scenes where modern deep lear…

Cited by 50SourceScholar
2018

Towards Effective Low-Bitwidth Convolutional Neural Networks

CVPR 2018poster

This paper tackles the problem of training a deep convolutional neural network with both low-precision weights and low-bitwidth activations. Optimizing a low-precision network is very challenging since the training process can easily get trapped in a poor local minima, which results in substantial a…

2018

Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction

CVPR 2018poster

Despite learning based methods showing promising results in single view depth estimation and visual odometry, most existing approaches treat the tasks in a supervised manner. Recent approaches to single view depth estimation explore the possibility of learning without full supervision via minimizing…

2018

Vision-and-Language Navigation: Interpreting Visually-Grounded Navigation Instructions in Real Environments

CVPR 2018poster

A robot that can carry out a natural-language instruction has been a dream since before the Jetsons cartoon series imagined a life of leisure mediated by a fleet of attentive robot helpers. It is a dream that remains stubbornly distant. However, recent advances in vision and language methods have m…

2018

Visual Question Answering With Memory-Augmented Networks

CVPR 2018poster

In this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We…

Cited by 134SourcePDFScholar
2017

"Maximizing Rigidity" Revisited: A Convex Programming Approach for Generic 3D Shape Reconstruction From Multiple Perspective Views

ICCV 2017poster

Rigid structure-from-motion (RSfM) and non-rigid structure-from-motion (NRSfM) have long been treated in the literature as separate (different) problems. Inspired by a previous work which solved directly for 3D scene structure by factoring the relative camera poses out, we revisit the principle of "…

Cited by 20PDFScholar
2017

A Bayesian Data Augmentation Approach for Learning Deep Models

NeurIPS 2017poster

Data augmentation is an essential part of the training process applied to deep learning models. The motivation is that a robust training process for deep learning models depends on large annotated datasets, which are expensive to be acquired, stored and processed. Therefore a reasonable alternativ…

2017

A branch-and-bound algorithm for checkerboard extraction in camera-laser calibration

ICRA 2017poster

We address the problem of camera-to-laserscanner calibration using a checkerboard and multiple imagelaser scan pairs. Distinguishing which laser points measure the checkerboard and which lie on the background is essential to any such system. We formulate the checkerboard extraction as a combinatoria…

Cited by 8SourceScholar
2017

A discrete-time attitude observer on SO(3) for vision and GPS fusion

ICRA 2017poster

This paper proposes a discrete-time geometric attitude observer for fusing monocular vision with GPS velocity measurements. The observer takes the relative transformations obtained from processing monocular images with any visual odometry algorithm and fuses them with GPS velocity measurements. The…

Cited by 12SourceScholar
2017

Attend in Groups: A Weakly-Supervised Deep Learning Framework for Learning From Web Data

CVPR 2017poster

Large-scale datasets have driven the rapid development of deep neural networks for visual recognition. However, annotating a massive dataset is expensive and time-consuming. Web images and their labels are, in comparison, much easier to obtain, but direct training on such automatially harvested imag…

Cited by 103PDFScholar
2017

Deep learning features at scale for visual place recognition

ICRA 2017poster

The success of deep learning techniques in the computer vision domain has triggered a range of initial investigations into their utility for visual place recognition, all using generic features from networks that were trained for other types of recognition tasks. In this paper, we train, at large sc…

Cited by 437SourceScholar
2017

DeepSetNet: Predicting Sets With Deep Neural Networks

ICCV 2017spotlight

This paper addresses the task of set prediction using deep learning. This is important because the output of many computer vision tasks, including image tagging and object detection, are naturally expressed as sets of entities rather than vectors. As opposed to a vector, the size of a set is not fix…

Cited by 56PDFScholar
2017

From Motion Blur to Motion Flow: A Deep Learning Solution for Removing Heterogeneous Motion Blur

CVPR 2017poster

Removing pixel-wise heterogeneous motion blur is challenging due to the ill-posed nature of the problem. The predominant solution is to estimate the blur kernel by adding a prior, but extensive literature on the subject indicates the difficulty in identifying a prior which is suitably informative, a…

Cited by 504PDFScholar
2017

Meaningful maps with object-oriented semantic mapping

IROS 2017poster

For intelligent robots to interact in meaningful ways with their environment, they must understand both the geometric and semantic properties of the scene surrounding them. The majority of research to date has addressed these mapping challenges separately, focusing on either geometric or semantic ma…

Cited by 292SourceScholar
2017

RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation

CVPR 2017poster

Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense classification problems such as semantic segmentation. However, repeated subsampling operations like pooling or convolution striding in deep…

Cited by 4039PDFcodeScholar
2017

Towards Context-Aware Interaction Recognition for Visual Relationship Detection

ICCV 2017poster

Recognizing how objects interact with each other is a crucial task in visual recognition. If we define the context of the interaction to be the objects involved, then most current methods can be categorized as either: (i) training a single classifier on the combination of the interaction and its con…

Cited by 199PDFcodeScholar
2016

Efficient Piecewise Training of Deep Structured Models for Semantic Segmentation

CVPR 2016spotlight

Recent advances in semantic image segmentation have mostly been achieved by training deep convolutional neural networks(CNNs). We show how to improve semantic segmentation through the use of contextual information; specifically, we explore 'patch-patch' context between image regions, and 'patch-back…

Cited by 1221PDFScholar
2016

Efficient Point Process Inference for Large-Scale Object Detection

CVPR 2016poster

We tackle the problem of large-scale object detection in images, where the number of objects can be arbitrarily large, and can exhibit significant overlap/occlusion. A successful approach to modelling the large-scale nature of this problem has been via point process density functions which jointly…

Cited by 38PDFScholar
2016

Fast Training of Triplet-Based Deep Binary Embedding Networks

CVPR 2016accepted

In this paper, we aim to learn a mapping (or embedding) from images to a compact binary space in which Hamming distances correspond to a ranking measure for the image retrieval task. We make use of a triplet loss because this has been shown to be most effective for ranking problems. How- ever, train…

Cited by 146SourcePDFScholar
2016

Geometrically consistent plane extraction for dense indoor 3D maps segmentation

IROS 2016poster

Modern SLAM systems with a depth sensor are able to reliably reconstruct dense 3D geometric maps of indoor scenes. Representing these maps in terms of meaningful entities is a step towards building semantic maps for autonomous robots. One approach is to segment the 3D maps into semantic objects usin…

Cited by 81SourceScholar
2016

Joint Probabilistic Matching Using m-Best Solutions

CVPR 2016oral

Matching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems c…

Cited by 41PDFScholar
2016

Learning Local Image Descriptors With Deep Siamese and Triplet Convolutional Networks by Minimising Global Loss Functions

CVPR 2016spotlight

Recent innovations in training deep convolutional neural network (ConvNet) models have motivated the design of new methods to automatically learn local image descriptors. The latest deep ConvNets proposed for this task consist of a siamese network that is trained by penalising misclassification of p…

Cited by 394PDFcodeScholar
2015

Deeply Learning the Messages in Message Passing Inference

NeurIPS 2015poster

Deep structured output learning shows great promise in tasks like semantic image segmentation. We proffer a new, efficient deep structured model learning scheme, in which we show how deep Convolutional Neural Networks (CNNs) can be used to directly estimate the messages in message passing inference…

Cited by 81SourcePDFScholar
2015

Hierarchical Higher-Order Regression Forest Fields: An Application to 3D Indoor Scene Labelling

ICCV 2015poster

This paper addresses the problem of semantic segmentation of 3D indoor scenes reconstructed from RGB-D images.Traditionally label prediction for 3D points is tackled by employing graphical models that capture scene features and complex relations between different class labels. However, the existing…

Cited by 31PDFScholar
2015

Joint Probabilistic Data Association Revisited

ICCV 2015poster

In this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applicatio…

Cited by 454PDFcodeScholar
2015

The k-Support Norm and Convex Envelopes of Cardinality and Rank

CVPR 2015poster

Sparsity, or cardinality, as a tool for feature selection is extremely common in a vast number of current computer vision applications. The $k$-support norm is a recently proposed norm with the proven property of providing the tightest convex bound on cardinality over the Euclidean norm unit ball. I…

Cited by 29SourcePDFScholar