← Search

Sen Wang

80 accepted papers

2026

AQUA-SLAM: Tightly-Coupled Underwater Acoustic-Visual-Inertial SLAM with Sensor Calibration

ICRA 2026poster

Underwater environments pose significant challenges for visual Simultaneous Localization and Mapping (SLAM) systems due to limited visibility, inadequate illumination, and sporadic loss of structural features in images. Addressing these challenges, this paper introduces a novel, tightly-coupled Acou…

2026

Déjà Vu: Unlocking Transparent Action Reasoning for Object-Goal Navigation Via Large Language Models

ICRA 2026poster

The remarkable interaction and reasoning capabilities of Large Language Models (LLMs) make them promising in collaborative Embodied AI tasks, particularly for Object-goal Navigation (ObjNav) tasks that require both decision-making and transparent explanation. However, existing work mainly uses LLMs …

Cited by 0Scholar
2026

EKF-Based Radar-Inertial Odometry with Online Temporal Calibration

ICRA 2026poster

Accurate time synchronization between heterogeneous sensors is crucial for ensuring robust state estimation in multi-sensor fusion systems. Sensor delays often cause discrepancies between the actual time when the event was captured and the time of sensor measurement, leading to temporal misalignment…

2026

Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration

CVPR 2026

An ideal embodied agent should possess lifelong learning capabilities to handle long-horizon and complex tasks, enabling continuous operation in general environments. This not only requires the agent to accurately accomplish given tasks but also to leverage long-term episodic memory to optimize deci

Cited by 0SourcecodeScholar
2026

InclusiveVidPose: Bridging the Pose Estimation Gap for Individuals with Limb Deficiencies in Video-Based Motion

ICLR 2026poster

Approximately 445.2 million individuals worldwide are living with traumatic amputations, and an estimated 31.64 million children aged 0–14 have congenital limb differences, yet they remain largely underrepresented in human pose estimation (HPE) research. Accurate HPE could significantly benefit this…

Cited by 0SourcecodeScholar
2026

Near-Field Driven Origami-Based Bio-Inspired Jellyfish Robot

ICRA 2026poster

The development of bio-inspired jellyfish robots holds significant benefits for autonomous aquatic systems due to jellyfish’s efficient water jet propulsion. However, the current design of jellyfish robots still faces challenges in balancing high biological fidelity with the demands of lightweight, …

Cited by 0Scholar
2026

PhysiXDeform: Real-Time Vision-Guided Soft-Tissue Deformation Prediction With Physical Priors

RA-L 2026

Accurate prediction of soft tissue deformation from endoscopic video is critical for robot-assisted, image-guided interventions. However, it remains challenging due to occlusions, complex dynamics, and the absence of direct physical measurements. Existing vision-based approaches often impose rigid o

Cited by 0SourceScholar
2026

Robust Test-time Video-Text Retrieval: Benchmarking and Adapting for Query Shifts

ICLR 2026poster

Modern video-text retrieval (VTR) models excel on in-distribution benchmarks but are highly vulnerable to real-world *query shifts*, where the distribution of query data deviates from the training domain, leading to a sharp performance drop. Existing image-focused robustness solutions are inadequate…

Cited by 0SourceScholar
2025

An Inflatable Deployable Origami Grasper for Adaptive and High-Load Grasping

IROS 2025

Robotic graspers are essential for enhancing the efficiency and versatility of robots in grasping tasks. In this paper, we propose a novel inflatable deployable origami grasper with a rigid-flexible coupling structure. The proposed grasper can achieve multiple deployment configurations under a singl

Cited by 0SourceScholar
2025

DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation

NeurIPS 2025poster

Learning generalizable robotic manipulation policies remains a key challenge due to the scarcity of diverse real-world training data. While recent approaches have attempted to mitigate this through self-supervised representation learning, most either rely on 2D vision pretraining paradigms such as m…

Cited by 0SourceScholar
2025

EKF-Based Radar-Inertial Odometry With Online Temporal Calibration

RA-L 2025

Accurate time synchronization between heterogeneous sensors is crucial for ensuring robust state estimation in multi-sensor fusion systems. Sensor delays often cause discrepancies between the actual time when the event was captured and the time of sensor measurement, leading to temporal misalignment

Cited by 9SourcecodeScholar
2025

FlowRAM: Grounding Flow Matching Policy with Region-Aware Mamba Framework for Robotic Manipulation

CVPR 2025poster

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational efficiency due to the iterative denoising process during in…

Cited by 0SourcePDFScholar
2025

From Enhancement to Understanding: Build a Generalized Bridge for Low-light Vision via Semantically Consistent Unsupervised Fine-tuning

ICCV 2025poster

Low-level enhancement and high-level visual understanding in low-light vision have traditionally been treated separately. Low-light enhancement improves image quality for downstream tasks, but existing methods rely on physical or geometric priors, limiting generalization. Evaluation mainly focuses o…

Cited by 0SourcePDFScholar
2025

General Scene Adaptation for Vision-and-Language Navigation

ICLR 2025poster

Vision-and-Language Navigation (VLN) tasks mainly evaluate agents based on one-time execution of individual instructions across multiple environments, aiming to develop agents capable of functioning in any environment in a zero-shot manner. However, real-world navigation robots often operate in pers…

2025

Is Less More? Exploring Token Condensation as Training-free Test-time Adaptation

ICCV 2025poster

Contrastive Language-Image Pretraining (CLIP) excels at learning generalizable image representations but often falls short in zero-shot inference on certain downstream datasets. Test-time adaptation (TTA) mitigates this issue by adjusting components like normalization layers or context prompts, yet…

2025

M3GYM: A Large-Scale Multimodal Multi-view Multi-person Pose Dataset for Fitness Activity Understanding in Real-world Settings

CVPR 2025poster

Human pose estimation is a critical task in computer vision for applications in sports analysis, healthcare monitoring, and human-computer interaction. However, existing human pose datasets are collected either from custom-configured laboratories with complex devices or they only include data on sin…

Cited by 0SourcePDFScholar
2025

Medium-Difficulty Samples Constitute Smoothed Decision Boundary for Knowledge Distillation on Pruned Datasets

ICLR 2025poster

This paper tackles a new problem of dataset pruning for Knowledge Distillation (KD), from a fresh perspective of Decision Boundary (DB) preservation and drifts. Existing dataset pruning methods generally assume that the post-pruning DB formed by the selected samples can be well-captured by future ne…

2025

Multimodal Retina Image Analysis Survey: Datasets, Tasks and Methods

IJCAI 2025

Retina images provide a noninvasive view of the central nervous system and microvasculature, making it essential for clinical applications. Changes in the retina often indicate both ophthalmic and systemic diseases, aiding in diagnosis and early intervention. While deep learning algorithms have adva

Cited by 0SourcePDFScholar
2025

PDFactor: Learning Tri-Perspective View Policy Diffusion Field for Multi-Task Robotic Manipulation

CVPR 2025poster

Robotic manipulation based on visual observations and natural language instructions is a long-standing challenge in robotics. Yet prevailing approaches model action distribution by adopting explicit or implicit representations, which often struggle to achieve a trade-off between accuracy and efficie…

Cited by 0SourcePDFScholar
2025

Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization

ICCV 2025poster

Despite recent advances, Text-to-video retrieval (TVR) is still hindered by multiple inherent uncertainties, such as ambiguous textual queries, indistinct text-video mappings, and low-quality video frames. Although interactive systems have emerged to address these challenges by refining user intent…

2025

RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers

ICML 2025poster

We reveal that feedforward network (FFN) layers, rather than attention layers, are the primary contributors to Vision Transformer (ViT) inference latency, with their impact signifying as model size increases. This finding highlights a critical opportunity for optimizing the efficiency of large-scale…

2025

SAMPO: Scale-wise Autoregression with Motion Prompt for Generative World Models

NeurIPS 2025poster

World models allow agents to simulate the consequences of actions in imagined environments for planning, control, and long-horizon decision-making. However, existing autoregressive world models struggle with visually coherent predictions due to disrupted spatial structure, inefficient decoding, and…

Cited by 0SourceScholar
2025

VoxNeRF: Bridging Voxel Representation and Neural Radiance Fields for Enhanced Indoor View Synthesis

RA-L 2025

The generation of high-fidelity view synthesis is essential for robotic navigation and interaction but remains challenging, particularly in indoor environments and real-time scenarios. Existing techniques often require significant computational resources for both training and rendering, and they fre

Cited by 2SourceScholar
2025

When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions

NeurIPS 2025poster

Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video temporal grounding. By revisiting the gap between current MR…

Cited by 0SourcecodeScholar
2024

CURL-MAP: Continuous Mapping and Positioning with CURL Representation†

ICRA 2024poster

Maps of LiDAR Simultaneous Localisation and Mapping (SLAM) are often represented as point clouds. They usually take up a huge amount of storage space for large-scale environments, otherwise much structural detail may not be kept. In this paper, a novel paradigm of LiDAR mapping and odometry is desig…

Cited by 1SourceScholar
2024

Event-Content-Oriented Dialogue Generation in Short Video

NAACL 2024long

Understanding complex events from different modalities, associating to external knowledge and generating response in a clear point of view are still unexplored in today’s multi-modal dialogue research. The great challenges include 1) lack of event-based multi-modal dialogue dataset; 2) understanding…

2024

HI-SLAM: Monocular Real-Time Dense Mapping With Hybrid Implicit Fields

RA-L 2024

In this letter, we present a neural field-based real-time monocular mapping framework for accurate and dense Simultaneous Localization and Mapping (SLAM). Recent neural mapping frameworks show promising results, but rely on RGB-D or pose inputs, or cannot run in real-time. To address these limitatio

Cited by 46SourceScholar
2024

MoMask: Generative Masked Modeling of 3D Human Motions

CVPR 2024poster

We introduce MoMask a novel masked modeling framework for text-driven 3D human motion generation. In MoMask a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer with a sequence of motion…

2024

Reachability Verification Based Reliability Assessment for Deep Reinforcement Learning Controlled Robotics and Autonomous Systems

RA-L 2024

Deep Reinforcement Learning (DRL) has achieved impressive performance in robotics and autonomous systems (RAS). A key challenge to its deployment in real-life operations is the presence of spuriously unsafe DRL policies. Unexplored states may lead the agent to make wrong decisions that could result

Cited by 8SourceScholar
2024

Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts

IJCAI 2024poster

Current Vision-and-Language Navigation (VLN) tasks mainly employ textual instructions to guide agents. However, being inherently abstract, the same textual instruction can be associated with different visual signals, causing severe ambiguity and limiting the transfer of prior knowledge in the vision…

2024

rWiFiSLAM: Effective WiFi Ranging Based SLAM System in Ambient Environments

RA-L 2024

In this paper, we propose rWiFiSLAM, an indoor localisation system based on WiFi ranging measurements. Indoor localisation techniques play an important role in mobile robots when they cannot access good quality GPS signals in indoor environments. Indoor localisation also has many other applications,

Cited by 6SourceScholar
2023

A Safety-Performance Metric Enabling Computational Awareness in Autonomous Robots

RA-L 2023

This letter takes a first step towards the analysis of safety <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">and</i> performance critical computational tasks for autonomous robots. Our contribution is a safety-performance (SP) metric that ensures sa

Cited by 5SourceScholar
2023

BAMF-SLAM: Bundle Adjusted Multi-Fisheye Visual-Inertial SLAM Using Recurrent Field Transforms

ICRA 2023poster

In this paper, we present BAMF-SLAM, a novel multi-fisheye visual-inertial SLAM system that utilizes Bundle Adjustment (BA) and recurrent field transforms (RFT) to achieve accurate and robust state estimation in challenging scenarios. First, our system directly operates on raw fisheye images, enabli…

Cited by 18SourceScholar
2023

DSVT: Dynamic Sparse Voxel Transformer With Rotated Sets

CVPR 2023poster

Designing an efficient yet deployment-friendly 3D backbone to handle sparse point clouds is a fundamental problem in 3D perception. Compared with the customized sparse convolution, the attention mechanism in Transformers is more appropriate for flexibly modeling long-range relationships and is easie…

2023

Observability-Aware Active Extrinsic Calibration of Multiple Sensors

ICRA 2023poster

The extrinsic parameters play a crucial role in multi-sensor fusion, such as visual-inertial Simultaneous Localization and Mapping(SLAM), as they enable the accurate alignment and integration of measurements from different sensors. However, extrinsic calibration is challenging in scenarios, such as…

Cited by 8SourceScholar
2023

RVD: A Handheld Device-Based Fundus Video Dataset for Retinal Vessel Segmentation

NeurIPS 2023poster

Retinal vessel segmentation is generally grounded in image-based datasets collected with bench-top devices. The static images naturally lose the dynamic characteristics of retina fluctuation, resulting in diminished dataset richness, and the usage of bench-top devices further restricts dataset scal…

Cited by 10SourcePDFScholar
2022

Generating Diverse and Natural 3D Human Motions From Text

CVPR 2022poster

Automated generation of 3D human motions from text is a challenging problem. The generated motions are expected to be sufficiently diverse to explore the text-grounded motion space, and more importantly, accurately depicting the content in prescribed text descriptions. Here we tackle this problem wi…

Cited by 615PDFcodeScholar
2022

Improved Feature Distillation via Projector Ensemble

NeurIPS 2022accept

In knowledge distillation, previous feature distillation methods mainly focus on the design of loss functions and the selection of the distilled layers, while the effect of the feature projector between the student and the teacher remains under-explored. In this paper, we first discuss a plausible m…

2022

Object Wake-Up: 3D Object Rigging from a Single Image

ECCV 2022poster

"Given a single chair image, could we wake it up by reconstructing its 3D shape and skeleton, as well as animating its plausible articulations and motions, similar to that of human modeling? It is a new problem that not only goes beyond image-based object reconstruction but also involves articulated…

Cited by 7SourcePDFScholar
2022

TM2T: Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts

ECCV 2022poster

"Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task, shorthanded for text2motion and motion2text, respectively. To tack…

2021

A Lightweight Soft Gripper Driven by Self-Sensing Super-Coiled Polymer Actuator

RA-L 2021

This study proposes a soft gripper driven by a simple, lightweight, low-cost, self-sensing super-coiled polymer (SCP) actuator with a high power-to-weight ratio. SCP generates an untwisting motion when heated, and therefore, is suitable for designing the actuator. The actuator is fabricated with a m

Cited by 22SourceScholar
2021

EventHPE: Event-Based 3D Human Pose and Shape Estimation

ICCV 2021poster

Event camera is an emerging imaging sensor for capturing dynamics of moving objects as events, which motivates our work in estimating 3D human pose and shape from the event signals. Events, on the other hand, have their unique challenges: rather than capturing static body postures, the event signals…

Cited by 58PDFcodeScholar
2021

Improving Embedding-based Large-scale Retrieval via Label Enhancement

EMNLP 2021finding

Current embedding-based large-scale retrieval models are trained with 0-1 hard label that indicates whether a query is relevant to a document, ignoring rich information of the relevance degree. This paper proposes to improve embedding-based retrieval from the perspective of better characterizing the…

Cited by 6SourcePDFScholar
2021

LiDAR-Aug: A General Rendering-Based Augmentation Framework for 3D Object Detection

CVPR 2021poster

Annotating the LiDAR point cloud is crucial for deep learning-based 3D object detection tasks. Due to expensive labeling costs, data augmentation has been taken as a necessary module and plays an important role in training the neural network. "Copy" and "paste" (i.e., GT-Aug) is the most commonly us…

Cited by 75PDFScholar
2021

QuadrupletBERT: An Efficient Model For Embedding-Based Large-Scale Retrieval

NAACL 2021long

The embedding-based large-scale query-document retrieval problem is a hot topic in the information retrieval (IR) field. Considering that pre-trained language models like BERT have achieved great success in a wide variety of NLP tasks, we present a QuadrupletBERT model for effective and efficient re…

Cited by 10SourcePDFScholar
2021

RADIATE: A Radar Dataset for Automotive Perception in Bad Weather

ICRA 2021poster

Datasets for autonomous cars are essential for the development and benchmarking of perception systems. However, most existing datasets are captured with camera and LiDAR sensors in good weather conditions. In this paper, we present the RAdar Dataset In Adverse weaThEr (RADIATE), aiming to facilitate…

Cited by 302SourcecodeScholar
2021

Robust Underwater Visual SLAM Fusing Acoustic Sensing

ICRA 2021poster

In this paper, we propose an approach for robust visual Simultaneous Localisation and Mapping (SLAM) in underwater environments leveraging acoustic, inertial and altimeter/depth sensors. Underwater visual SLAM is challenging due to factors including poor visibility caused by suspended particles in w…

Cited by 61SourceScholar
2021

Self-Supervised Adversarial Distribution Regularization for Medication Recommendation

IJCAI 2021poster

Medication recommendation is a significant healthcare application due to its promise in effectively prescribing medications. Avoiding fatal side effects related to Drug-Drug Interaction (DDI) is among the critical challenges. Most existing methods try to mitigate the problem by providing models with…

2021

Semantics Disentangling for Generalized Zero-Shot Learning

ICCV 2021poster

Generalized zero-shot learning (GZSL) aims to classify samples under the assumption that some classes are not observable during training. To bridge the gap between the seen and unseen classes, most GZSL methods attempt to associate the visual features of seen classes with attributes or to generate u…

Cited by 149PDFcodeScholar
2021

Underwater Visual Acoustic SLAM with Extrinsic Calibration

IROS 2021poster

Underwater scenarios are challenging for visual Simultaneous Localization and Mapping (SLAM) due to limited visibility and intermittently losing structures in image views. In this paper, we propose a visual acoustic bundle adjustment system which fuses a camera and a Doppler Velocity Log (DVL) in a…

Cited by 32SourceScholar
2020

3D Human Shape Reconstruction from a Polarization Image

ECCV 2020poster

This paper tackles the problem of estimating 3D body shape of clothed humans from single polarized 2D images, i.e. polarization images. Polarization images are known to be able to capture polarized reflected lights that preserve rich geometric cues of an object, which has motivated its recent applic…

Cited by 56SourcePDFScholar
2020

Interacting Vehicle Trajectory Prediction with Convolutional Recurrent Neural Networks

ICRA 2020poster

Anticipating the future trajectories of surrounding vehicles is a crucial and challenging task in path planning for autonomy. We propose a novel Convolutional Long Short Term Memory (Conv-LSTM) based neural network architecture to predict the future positions of cars using several seconds of histori…

Cited by 18SourceScholar
2020

Quadratic Sparse Gaussian Graphical Model Estimation Method for Massive Variables

IJCAI 2020poster

We consider the problem of estimating a sparse Gaussian Graphical Model with a special graph topological structure and more than a million variables. Most previous scalable estimators still contain expensive calculation steps (e.g., matrix inversion or Hessian matrix calculation) and become infeasib…

Cited by 0SourcePDFScholar
2020

Robot Calligraphy using Pseudospectral Optimal Control in Conjunction with a Novel Dynamic Brush Model

IROS 2020poster

Chinese calligraphy is a unique art form with great artistic value but difficult to master. In this paper, we formulate the calligraphy writing problem as a trajectory optimization problem, and propose an improved virtual brush model for simulating the real writing process. Our approach is inspired…

Cited by 31SourceScholar
2020

SolarSLAM: Battery-free Loop Closure for Indoor Localisation

IROS 2020poster

In this paper, we propose SolarSLAM, a batteryfree loop closure method for indoor localisation. Inertial Measurement Unit (IMU) based indoor localisation method has been widely used due to its ubiquity in mobile devices, such as mobile phones, smartwatches and wearable bands. However, it suffers fro…

Cited by 6SourceScholar
2019

Detailed Human Shape Estimation From a Single Image by Hierarchical Mesh Deformation

CVPR 2019oral

This paper presents a novel framework to recover detailed human body shapes from a single image. It is a challenging task due to factors such as variations in human shapes, body poses, and viewpoints. Prior methods typically attempt to recover the human body shape using a parametric based template…

Cited by 173PDFcodeScholar
2019

Learning Monocular Visual Odometry through Geometry-Aware Curriculum Learning

ICRA 2019poster

Inspired by the cognitive process of humans and animals, Curriculum Learning (CL) trains a model by gradually increasing the difficulty of the training data. In this paper, we study whether CL can be applied to complex geometry problems like estimating monocular Visual Odometry (VO). Unlike existing…

Cited by 58SourceScholar
2019

Learning Object Bounding Boxes for 3D Instance Segmentation on Point Clouds

NeurIPS 2019spotlight

We propose a novel, conceptually simple and general framework for instance segmentation on 3D point clouds. Our method, called 3D-BoNet, follows the simple design philosophy of per-point multilayer perceptrons (MLPs). The framework directly regresses 3D bounding boxes for all instances in a point cl…

2019

TextPlace: Visual Place Recognition and Topological Localization Through Reading Scene Texts

ICCV 2019poster

Visual place recognition is a fundamental problem for many vision based applications. Sparse feature and deep learning based methods have been successful and dominant over the decade. However, most of them do not explicitly leverage high-level semantic information to deal with challenging scenarios…

Cited by 67PDFcodeScholar
2018

DEFO-NET: Learning Body Deformation Using Generative Adversarial Networks

ICRA 2018poster

Modelling the physical properties of everyday objects is a fundamental prerequisite for autonomous robots. We present a novel generative adversarial network (DEFO-NET), able to predict body deformations under external forces from a single RGB-D image. The network is based on an invertible conditiona…

Cited by 10SourceScholar
2018

Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning

ICRA 2018poster

Deep Reinforcement Learning (DRL) has been applied successfully to many robotic applications. However, the large number of trials needed for training is a key issue. Most of existing techniques developed to improve training efficiency (e.g. imitation) target on general tasks rather than being tailor…

Cited by 105SourcecodeScholar
2018

UnDeepVO: Monocular Visual Odometry Through Unsupervised Deep Learning

ICRA 2018poster

We propose a novel monocular visual odometry (VO) system called UnDeepVO in this paper. UnDeepVO is able to estimate the 6-DoF pose of a monocular camera and the depth of its view by using deep neural networks. There are two salient features of the proposed UnDeepVo:one is the unsupervised deep lear…

Cited by 696SourceScholar
2017

A generative human-robot motion retargeting approach using a single depth sensor

ICRA 2017poster

The goal of human-robot motion retargeting is to let a robot follow the movements performed by a human subject. This is traditionally achieved by applying the estimated poses from a human pose tracking system to a robot via explicit joint mapping strategies. In this paper, we present a novel approac…

Cited by 24SourceScholar
2017

DeepVO: Towards end-to-end visual odometry with deep Recurrent Convolutional Neural Networks

ICRA 2017poster

This paper studies monocular visual odometry (VO) problem. Most of existing VO algorithms are developed under a standard pipeline including feature extraction, feature matching, motion estimation, local optimisation, etc. Although some of them have demonstrated superior performance, they usually nee…

Cited by 1123SourceScholar
2017

Detailed Surface Geometry and Albedo Recovery From RGB-D Video Under Natural Illumination

ICCV 2017poster

In this paper we present a novel approach for depth map enhancement from an RGB-D video sequence. The basic idea is to exploit the photometric information in the color sequence. Instead of making any assumption about surface albedo or controlled object motion and lighting, we use the lighting variat…

Cited by 15PDFScholar
2017

GraphTinker: Outlier rejection and inlier injection for pose graph SLAM

IROS 2017poster

In pose graph Simultaneous Localization and Mapping (SLAM) systems, incorrect loop closures can seriously hinder optimizers from converging to correct solutions, significantly degrading both localization accuracy and map consistency. Therefore, it is crucial to enhance their robustness in the presen…

Cited by 14SourceScholar
2017

VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip Relocalization

CVPR 2017poster

Machine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of th…

Cited by 325PDFcodeScholar
2016

Keyframe based large-scale indoor localisation using geomagnetic field and motion pattern

IROS 2016poster

This paper studies indoor localisation problem by using low-cost and pervasive sensors. Most of existing indoor localisation algorithms rely on camera, laser scanner, floor plan or other pre-installed infrastructure to achieve sub-meter or sub-centimetre localisation accuracy. However, in some circu…

Cited by 69SourceScholar
2015

Interactive Visual Hull Refinement for Specular and Transparent Object Surface Reconstruction

ICCV 2015poster

In this paper we present a method of using standard multi-view images for 3D surface reconstruction of non-Lambertian objects. We extend the original visual hull concept to incorporate 3D cues presented by internal occluding contours, i.e., occluding contours that are inside the object's silhouettes…

Cited by 23PDFScholar