← Search

Hong Zhang

113 accepted papers

2026

Adap-RPF: Adaptive Trajectory Sampling for Robot Person Following in Dynamic Crowded Environments

ICRA 2026poster

Robot person following (RPF) is a core capability in human–robot interaction, enabling robots to assist users in daily activities, collaborative work, and other service scenarios. However, achieving practical RPF remains challenging due to frequent occlusions, particularly in dynamic and crowded env…

2026

CogStereo: Neural Stereo Matching with Implicit Spatial Cognition Embedding

ICRA 2026poster

Deep stereo matching has advanced significantly on benchmark datasets through fine-tuning but falls short of the zero-shot generalization seen in foundation models in other vision tasks. We introduce CogStereo, a novel framework that addresses challenging regions, such as occlusions or weak textures…

2026

Don't Force the Fit: Bounded Log-Likelihood Loss for Enhanced Reasoning in Large Language Models

ICML 2026oral

Supervised fine-tuning (SFT) is central to aligning large language models (LLMs) with instruction following and task-specific reasoning. Despite its success, SFT optimizes token-level likelihoods under the implicit assumption that strictly fitting all tokens in expert demonstrations induces the desi…

Cited by 0SourceScholar
2026

EZREAL: Enhancing Zero-Shot Outdoor Robot Navigation Toward Distant Targets under Varying Visibility

ICRA 2026poster

Zero-shot object navigation (ZSON) in large-scale outdoor environments faces many challenges; we specifically address a coupled one: long-range targets that reduce to tiny projections and intermittent visibility due to partial or complete occlusion. We present a unified, lightweight closed-loop syst…

2026

EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence

CVPR 2026

Video-based spatial reasoning -- such as estimating distances, judging directions, or understanding layouts from multiple views -- requires selecting informative frames and, when needed, actively seeking additional viewpoints during inference. Existing multimodal large language models (MLLMs) consum

Cited by 0SourceScholar
2026

FullPart: Generating each 3D Part at Full Resolution

ICLR 2026poster

Part-based 3D generation holds great potential for various applications. Previous part generators that represent parts using implicit vector-set tokens often suffer from insufficient geometric details. Another line of work adopts an explicit voxel representation but shares a global voxel grid among…

Cited by 0SourcecodeScholar
2026

GraphCoT-VLA: A 3D Spatial-Aware Reasoning Vision-Language-Action Model for Robotic Manipulation with Ambiguous Instructions

AAAI 2026technical

Vision-language-action models have emerged as a crucial paradigm in robotic manipulation. However, existing VLA models exhibit notable limitations in handling ambiguous language instructions and unknown environmental states. Furthermore, their perception is largely constrained to static two-dimensio

Cited by 0SourcePDFScholar
2026

Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting

CVPR 2026

Visual relocalization is a fundamental task in the field of 3D computer vision, estimating a camera's pose when it revisits a previously known scene. While point-based hierarchical relocalization methods have shown strong scalability and efficiency, they are often limited by sparse image observation

Cited by 0SourceScholar
2026

Online Navigation Refinement: Achieving Lane-Level Guidance by Associating Standard-Definition and Online Perception Maps

ICLR 2026poster

Lane-level navigation is critical for geographic information systems and navigation-based tasks, offering finer-grained guidance than road-level navigation by standard definition (SD) maps. However, it currently relies on expansive global HD maps that cannot adapt to dynamic road conditions. Recentl…

Cited by 0SourcecodeScholar
2026

Towards Explainable Video Camouflaged Object Detection: SAM2 with Eventstream-Inspired Data

AAAI 2026technical

Video Camouflaged Object Detection (VCOD) poses significant challenges due to the subtle appearance of camouflaged objects, especially under dynamic motion and occlusion. Existing methods predominantly rely on optical flow or black-box features for motion modeling, which often entail substantial com

Cited by 0SourcePDFScholar
2026

TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking

ICRA 2026poster

Embodied Visual Tracking (EVT) is a fundamental ability that underpins practical applications, such as companion robots, guidance robots and service assistants, where continuously following moving targets is essential. Recent advances have enabled language-guided tracking in complex and unstructured…

2026

Unveiling Uncertainty-Aware Autonomous Cooperative Learning Based Planning Strategy

ICRA 2026poster

In future intelligent transportation systems, autonomous cooperative planning (ACP), becomes a promising technique to increase the effectiveness and security of multi-vehicle interactions. However, multiple uncertainties cannot be fully addressed for existing ACP strategies, e.g. perception, plannin…

2025

3D Lane Detection Based on Projection-Consistent Reference Points and Intra- & Inter-lane Context

ICRA 2025

3D lane detection aims to identify lane categories and trends in 3D space, which is a vital and challenging task in autonomous driving. Existing methods introduce various priors to guide 3D lane prediction, which generally consist of a series of reference points for context aggregation. However, due

Cited by 1SourceScholar
2025

C2MIL: Synchronizing Semantic and Topological Causalities in Multiple Instance Learning for Robust and Interpretable Survival Analysis

ICCV 2025poster

Graph-based Multiple Instance Learning (MIL) is widely used in survival analysis with Hematoxylin and Eosin (H&E)-stained whole slide images (WSIs) due to its ability to capture topological information. However, variations in staining and scanning can introduce semantic bias, while topological subgr…

2025

Correcting on Graph: Faithful Semantic Parsing over Knowledge Graphs with Large Language Models

ACL 2025finding

Complex multi-hop questions often require comprehensive retrieval and reasoning. As a result, effectively parsing such questions and establishing an efficient interaction channel between large language models (LLMs) and knowledge graphs (KGs) is essential for ensuring reliable reasoning. In this pap…

2025

DEPTHOR: Depth Enhancement from a Practical Light-Weight dToF Sensor and RGB Image

ICCV 2025poster

Depth enhancement, which uses RGB images as guidance to convert raw signals from dToF into high-precision, dense depth maps, is a critical task in computer vision. Although existing super-resolution-based methods show promising results on public datasets, they often rely on idealized assumptions lik…

2025

Density Adaptive Registration of Large-Scale Point Clouds in Diverse Outdoor Environments

RA-L 2025

Point cloud registration is the foundation of collaborative multi-robot mapping tasks in outdoor environments. Due to the dynamic changes in communication bandwidth, the density of point clouds transmitted from the robot to the server will also change simultaneously, which will significantly affect

Cited by 1SourceScholar
2025

Dynamic Perception-Enhanced Motion Planning and Control for UAVs Flights in Challenging Dynamic Environments

ICRA 2025

The autonomous flights of unmanned aerial vehicles (UAVs) in unknown environments have garnered significant attention. However, most existing methods only achieve safe navigation in static environments or spacious scenes with few moving obstacles. Motivated by this open problem, this paper presents

Cited by 0SourceScholar
2025

ExVideo: Extending Video Diffusion Models via Parameter-Efficient Post-Tuning

IJCAI 2025

Recently, advancements in video synthesis have attracted significant attention. Video synthesis models have demonstrated the practical applicability of diffusion models in creating dynamic visual content. Despite these advancements, the extension of video lengths remains constrained by computational

2025

FLAF: Focal Line and Feature-Constrained Active View Planning for Visual Teach and Repeat

ICRA 2025

This paper presents FLAF, a focal line and feature-constrained active view planning method for autonomous orientation adjustment of a rotatable active camera during mobile robot navigation. FLAF is built on a visual teach-and-repeat (VT&R) system, which enables robots to cruise various paths that fu

Cited by 1SourceScholar
2025

FlowPlan: Zero-Shot Task Planning with LLM Flow Engineering for Robotic Instruction Following

IROS 2025

Robotic instruction following tasks require seamless integration of visual perception, task planning, target localization, and motion execution. However, existing task planning methods for instruction following are either data-driven or underperform in zero-shot scenarios due to difficulties in grou

Cited by 1SourcecodeScholar
2025

GMMCL: Adaptive Concept Drift in Data Streams with Gaussian Mixture Models based on Contrastive Learning

ICASSP 2025accepted

Classical classification methods often fail in dynamic environments where data distributions shift over time, known as concept drift. Applications like flight delay prediction and weather forecasting require handling such dynamic data streams. Concept drift can be either virtual, affecting unconditi…

Cited by 0SourceScholar
2025

HGDiffuser: Efficient Task-Oriented Grasp Generation via Human-Guided Grasp Diffusion Models

IROS 2025

Task-oriented grasping (TOG) is essential for robots to perform manipulation tasks, requiring grasps that are both stable and compliant with task-specific constraints. Humans naturally grasp objects in a task-oriented manner to facilitate subsequent manipulation tasks. By leveraging human grasp demo

Cited by 5SourceScholar
2025

Heterogeneous Graph Network-Based UWB Localization for Complex Indoor Environments

IROS 2025

Accurate indoor location-based services are important for mobile robots, especially in complex indoor environments. In this paper, we propose a heterogeneous graph network-based ultra-wide band (UWB) localization method to provide accurate and robust localization results for mobile robots in complex

Cited by 0SourceScholar
2025

Improving Flow Matching by Aligning Flow Divergence

ICML 2025poster

Conditional flow matching (CFM) stands out as an efficient, simulation-free approach for training flow-based generative models, achieving remarkable performance for data generation. However, CFM is insufficient to ensure accuracy in learning probability paths. In this paper, we introduce a new parti…

Cited by 0SourcePDFScholar
2025

Inductive Reasoning on Few-Shot Knowledge Graphs with Task-Aware Language Models

EMNLP 2025

Knowledge graphs are dynamic structures that continuously evolve as new entities emerge, often accompanied by only a handful of associated triples. Current knowledge graph reasoning methods struggle in these few-shot scenarios due to their reliance on extensive structural information.To address this

Cited by 0SourcePDFScholar
2025

JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction

IROS 2025

Predicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability modes in multi-agent joint prediction. To tackle this issue, we propose a two-stage multi-agent interactive prediction frame

Cited by 0SourcecodeScholar
2025

MimicFunc: Imitating Tool Manipulation from a Single Human Video via Functional Correspondence

CoRL 2025poster

Imitating tool manipulation from human videos offers an intuitive approach to teaching robots, while also providing a promising and scalable alternative to labor-intensive teleoperation data collection for visuomotor policy learning. While humans can mimic tool manipulation behavior by observing oth…

Cited by 0SourceScholar
2025

MuRAR: A Simple and Effective Multimodal Retrieval and Answer Refinement Framework for Multimodal Question Answering

COLING 2025system demonstrations

Recent advancements in retrieval-augmented generation have demonstrated impressive performance on the question-answering task. However, most previous work predominantly focuses on text-based answers. Although some studies have explored multimodal data, they still fall short in generating comprehensi…

2025

P2d-DO: Degeneracy Optimization for LiDAR SLAM With Point-to-Distribution Detection Factors

RA-L 2025

Although the LiDAR SLAM technique has been already widely deployed on various robots, it may still suffers from degeneracy caused by inadequate constraints in scenes with sparse geometric features. If the degeneracy is not detected and properly processed, the accuracy of localization and mapping wil

Cited by 11SourceScholar
2025

Priority on High-Quality: Selecting Instruction Data via Consistency Verification of Noise Injection

EMNLP 2025

Large Language Models (LLMs) have demonstrated a remarkable understanding of language nuances through instruction tuning, enabling them to effectively tackle various natural language processing tasks. Recent research has focused on the quality of instruction data rather than the quantity of instruct

2025

RTAGrasp: Learning Task-Oriented Grasping from Human Videos via Retrieval, Transfer, and Alignment

ICRA 2025

Task-oriented grasping (TOG) is crucial for robots to accomplish manipulation tasks, requiring the determination of TOG positions and directions. Existing methods either rely on costly manual TOG annotations or only extract coarse grasping positions or regions from human demonstrations, limiting the

Cited by 12SourceScholar
2025

Reducing Redundancy in VSLAM: VLMs-driven Keyframe Selection using Multi-dimensional Semantic Information

IROS 2025

Keyframe selection plays a crucial role in balancing computational efficiency and localization accuracy in Visual Simultaneous Localization and Mapping (VSLAM) systems. Existing keyframe selection methods often struggle to capture high-level semantic information in environments where multiple semant

Cited by 2SourceScholar
2025

SP2T: Sparse Proxy Attention for Dual-stream Point Transformer

ICCV 2025poster

Point transformers have demonstrated remarkable progress in 3D understanding through expanded receptive fields (RF), but further expanding the RF leads to dilution in group attention and decreases detailed feature extraction capability. Proxy, which serves as abstract representations for simplifying…

2025

SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective Fusion

CVPR 2025poster

Lightweight direct Time-of-Flight (dToF) sensors are ideal for 3D sensing on mobile devices. However, due to the manufacturing constraints of compact devices and the inherent physical principles of imaging, dToF depth maps are sparse and noisy. In this paper, we propose a novel video depth completio…

2025

SuperPromptSeg: A Novel Fine-Tuning-Free Segmentation Method Leveraging Superpixel-Based Point Prompts

ICASSP 2025accepted

The efficient segmentation of histopathological tissues plays a crucial role in aiding diagnostics and prognosis. However, the existing segmentation methods not only typically demand extensive human annotation and/or time-consuming model fine-tuning, but also exhibit limited generalization capabilit…

Cited by 0SourceScholar
2025

TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification

IROS 2025

Visual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information. However, existing methods remain limited in indoor settings due to the highly repetitive structures inherent in such environments.

Cited by 1SourcecodeScholar
2025

Unveiling Uncertainty-Aware Autonomous Cooperative Learning Based Planning Strategy

RA-L 2025

In future intelligent transportation systems, autonomous cooperative planning (ACP), becomes a promising technique to increase the effectiveness and security of multi-vehicle interactions. However, multiple uncertainties cannot be fully addressed for existing ACP strategies, e.g. perception, plannin

Cited by 1SourceScholar
2024

A Point-to-distribution Degeneracy Detection Factor for LiDAR SLAM using Local Geometric Models

ICRA 2024poster

Limited by the working principles, LiDAR-SLAM systems suffer from the degeneration phenomenon in environments such as long corridors and tunnels, due to the lack of sufficient geometric features for frame-to-frame matching. The accuracy and sensitivity of existing degeneracy detection methods need t…

Cited by 5SourcecodeScholar
2024

DA-Net: A Disentangled and Adaptive Network for Multi-Source Cross-Lingual Transfer Learning

AAAI 2024technical

Multi-Source cross-lingual transfer learning deals with the transfer of task knowledge from multiple labelled source languages to an unlabeled target language under the language shift. Existing methods typically focus on weighting the predictions produced by language-specific classifiers of differen…

Cited by 1SourcePDFScholar
2024

Discrepancy and Uncertainty Aware Denoising Knowledge Distillation for Zero-Shot Cross-Lingual Named Entity Recognition

AAAI 2024technical

The knowledge distillation-based approaches have recently yielded state-of-the-art (SOTA) results for cross-lingual NER tasks in zero-shot scenarios. These approaches typically employ a teacher network trained with the labelled source (rich-resource) language to infer pseudo-soft labels for the unl…

2024

EFP: Efficient Frontier-Based Autonomous UAV Exploration Strategy for Unknown Environments

RA-L 2024

The optimization of quadrotors for the efficient and autonomous exploration of complex, unknown environments and the construction of corresponding maps with integrity is of high priority in unmanned aerial vehicle(UAV) research. To overcome the challenges of inefficient and incomplete map constructi

Cited by 41SourceScholar
2024

GV-Bench: Benchmarking Local Feature Matching for Geometric Verification of Long-term Loop Closure Detection

IROS 2024poster

Visual loop closure detection is an important module in visual simultaneous localization and mapping (SLAM), which associates current camera observation with previously visited places. Loop closures correct drifts in trajectory estimation to build a globally consistent map. However, a false loop clo…

Cited by 3SourcecodeScholar
2024

PISR: Polarimetric Neural Implicit Surface Reconstruction for Textureless and Specular Objects

ECCV 2024poster

"Neural implicit surface reconstruction has achieved remarkable progress recently. Despite resorting to complex radiance modeling, state-of-the-art methods still struggle with textureless and specular surfaces. Different from RGB images, polarization images can provide direct constraints on the azim…

2024

Parsing All Adverse Scenes: Severity-Aware Semantic Segmentation with Mask-Enhanced Cross-Domain Consistency

AAAI 2024technical

Although recent methods in Unsupervised Domain Adaptation (UDA) have achieved success in segmenting rainy or snowy scenes by improving consistency, they face limitations when dealing with more challenging scenarios like foggy and night scenes. We argue that these prior methods excessively focus on w…

Cited by 8SourcePDFScholar
2024

Person Re-Identification for Robot Person Following With Online Continual Learning

RA-L 2024

Robot person following (RPF) is a crucial capability in human-robot interaction (HRI) applications, allowing a robot to persistently follow a designated person. In practical RPF scenarios, the person can often be occluded by other objects or people. Consequently, it is necessary to re-identify the p

Cited by 18SourceScholar
2024

QI-IRA: Quantum-Inspired Interactive Ranking Aggregation for Person Re-identification

AAAI 2024technical

Ranking aggregation (RA), the process of aggregating multiple rankings derived from multiple search strategies, has been proved effective in person re-identification (re-ID) because of a single re-ID method can not always achieve consistent superiority for different scenarios. Existing RA research m…

2024

SWCF-Net: Similarity-weighted Convolution and Local-global Fusion for Efficient Large-scale Point Cloud Semantic Segmentation

IROS 2024poster

Large-scale point cloud consists of a multitude of individual objects, thereby encompassing rich structural and underlying semantic contextual information, resulting in a challenging problem in efficiently segmenting a point cloud. Most existing researches mainly focus on capturing intricate local f…

Cited by 2SourcecodeScholar
2024

Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and Visual Analysis Strategy

ECCV 2024poster

"In the domain of Camouflaged Object Segmentation (COS), despite continuous improvements in segmentation performance, the underlying mechanisms of effective camouflage remain poorly understood, akin to a black box. To address this gap, we present the first comprehensive study to examine the impact o…

2023

An Open-Source Robotic Chinese Chess Player

IROS 2023poster

Consumer robots can accompany children growing up, improving their abilities while playing and entertaining. This paper presents an open-source, practical, low-cost robotic Chinese chess player. The proposed system includes an elaborate mechanical structure, a simple kinematic solution, a novel robo…

Cited by 1SourcecodeScholar
2023

Combining Scene Coordinate Regression and Absolute Pose Regression for Visual Relocalization

ICRA 2023poster

Visual relocalization is a fundamental problem in computer vision and robotics. Recently, regression-based methods become popular and they can be categorized into two classes: absolute pose regression and scene coordinate regression. In this work, we present a combined regression network that jointl…

Cited by 4SourceScholar
2023

Continuous-Time LiDAR-Inertial-Vehicle Odometry Method with Lateral Acceleration Constraint

ICRA 2023poster

In this paper, we propose a continuous-time-based LiDAR-inertial-vehicle odometry method, which can tightly fuse the data from Light Detection And Ranging (LiDAR), inertial measurement units (IMU), and vehicle measurements. The lateral acceleration constraint is further added to trajectory estimatio…

Cited by 3SourceScholar
2023

Curiosity-based Robot Navigation under Uncertainty in Crowded Environments

RA-L 2023

Mobile robots have become more and more popular in large-scale and crowded environments, such as airports, shopping malls, etc. However, due to sparse landmarks and crowd noise, localization in this environment is a great challenge. Furthermore, it is unreliable for the robot to navigate safely in c

Cited by 10SourceScholar
2023

GraspGPT: Leveraging Semantic Knowledge From a Large Language Model for Task-Oriented Grasping

RA-L 2023

Task-oriented grasping (TOG) refers to the problem of predicting grasps on an object that enable subsequent manipulation tasks. To model the complex relationships between objects, tasks, and grasps, existing methods incorporate semantic knowledge as priors into TOG pipelines. However, the existing s

Cited by 122SourceScholar
2023

Multi-Scale Point Octree Encoding Network for Point Cloud Based Place Recognition

IROS 2023poster

Over the past decades, point cloud-based place recognition has garnered significant attention. This research paper presents a pioneering approach, denoted as the Multi-scale Point Octree Encoding Network (MPOE-Net), designed to acquire a discriminative global descriptor for efficient retrieval of pl…

Cited by 1SourcecodeScholar
2023

Multi-View Robust Graph Representation Learning for Graph Classification

IJCAI 2023poster

The robustness of graph classification models plays an essential role in providing highly reliable applications. Previous studies along this line primarily focus on seeking the stability of the model in terms of overall data metrics (e.g., accuracy) when facing data perturbations, such as removing…

Cited by 10SourcePDFScholar
2023

ProKD: An Unsupervised Prototypical Knowledge Distillation Network for Zero-Resource Cross-Lingual Named Entity Recognition

AAAI 2023technical

For named entity recognition (NER) in zero-resource languages, utilizing knowledge distillation methods to transfer language-independent knowledge from the rich-resource source languages to zero-resource languages is an effective means. Typically, these approaches adopt a teacher-student architectur…

Cited by 10SourcePDFScholar
2023

Recovering a Molecule's 3D Dynamics from Liquid-phase Electron Microscopy Movies

ICCV 2023poster

The dynamics of biomolecules are crucial for our understanding of their functioning in living systems. However, current 3D imaging techniques, such as cryogenic electron microscopy (cryo-EM), require freezing the sample, which limits the observation of their conformational changes in real time. The…

Cited by 5PDFScholar
2023

Robot Person Following Under Partial Occlusion

ICRA 2023poster

Robot person following (RPF) is a capability that supports many useful human-robot-interaction (HRI) applications. However, existing solutions to person following often as-sume full observation of the tracked person. As a consequence, they cannot track the person reliably under partial occlusion whe…

Cited by 20SourcecodeScholar
2023

Task-Oriented Grasp Prediction with Visual-Language Inputs

IROS 2023poster

To perform household tasks, assistive robots receive commands in the form of user language instructions for tool manipulation. The initial stage involves selecting the intended tool (i.e., object grounding) and grasping it in a task-oriented manner (i.e., task grounding). Nevertheless, prior researc…

Cited by 40SourceScholar
2022

Are We Ready for Robust and Resilient SLAM? A Framework For Quantitative Characterization of SLAM Datasets

IROS 2022poster

Reliability of SLAM systems is considered one of the critical requirements in modern autonomous systems. This directed the efforts to developing many state-of-the-art systems, creating challenging datasets, and introducing rigorous metrics to measure SLAM performance. However, the link between datas…

Cited by 8SourceScholar
2022

Deep Tri-Training for Semi-Supervised Image Segmentation

RA-L 2022

Semantic segmentation is of great value to autonomous driving and many robotic applications, while it highly depends on costly and time-consuming pixel-level annotation. To make full use of unlabeled data, this work proposes a deep tri-training framework (dubbed DTT) to utilize labeled along with un

Cited by 11SourceScholar
2022

E-VarM: Enhanced Variational Word Masks to Improve the Interpretability of Text Classification Models

COLING 2022main

Enhancing the interpretability of text classification models can help increase the reliability of these models in real-world applications. Currently, most researchers focus on extracting task-specific words from inputs to improve the interpretability of the model. The competitive approaches exploit…

2022

FedDUAP: Federated Learning with Dynamic Update and Adaptive Pruning Using Shared Data on the Server

IJCAI 2022poster

Despite achieving remarkable performance, Federated Learning (FL) suffers from two critical challenges, i.e., limited computational resources and low training efficiency. In this paper, we propose a novel FL framework, i.e., FedDUAP, with two original contributions, to exploit the insensitive data o…

Cited by 52SourcePDFScholar
2022

Fire Together Wire Together: A Dynamic Pruning Approach With Self-Supervised Mask Prediction

CVPR 2022poster

Dynamic model pruning is a recent direction that allows for the inference of a different sub-network for each input sample during deployment. However, current dynamic methods rely on learning a continuous channel gating through regularization by inducing sparsity loss. This formulation introduces co…

Cited by 45PDFcodeScholar
2022

Keyframe Selection with Information Occupancy Grid Model for Long-term Data Association

IROS 2022poster

As the basics of Visual Simultaneous Localization And Mapping (VSLAM), keyframes play an essential role. In previous works, keyframes are selected according to a series of view change-based strategies for short-term data association (STDA). However, the texture enrichment of frames is always ignored…

Cited by 2SourceScholar
2022

MonoDistill: Learning Spatial Features for Monocular 3D Object Detection

ICLR 2022poster

3D object detection is a fundamental and challenging task for 3D scene understanding, and the monocular-based methods can serve as an economical alternative to the stereo-based or LiDAR-based methods. However, accurately locating objects in the 3D space from a single image is extremely difficult due…

2022

NDD: A 3D Point Cloud Descriptor Based on Normal Distribution for Loop Closure Detection

IROS 2022poster

Loop closure detection is a key technology for long-term robot navigation in complex environments. In this paper, we present a global descriptor, named Normal Distribution Descriptor (NDD), for 3D point cloud loop closure detection. The descriptor encodes both the probability density score and entro…

Cited by 14SourcecodeScholar
2022

Open-Topic False Information Detection on Social Networks with Contrastive Adversarial Learning

EMNLP 2022main

Current works about false information detection based on conversation graphs on social networks focus primarily on two research streams from the standpoint of topic distribution: in-topic and cross-topic techniques, which assume that the data topic distribution is identical or cross, respectively. T…

2022

Perspective Phase Angle Model for Polarimetric 3D Reconstruction

ECCV 2022poster

"Current polarimetric 3D reconstruction methods, including those in the well-established shape from polarization literature, are all developed under the orthographic projection assumption. In the case of a large field of view, however, this assumption does not hold and may result in significant reco…

2022

Provable Benefit of Multitask Representation Learning in Reinforcement Learning

NeurIPS 2022accept

As representation learning becomes a powerful technique to reduce sample complexity in reinforcement learning (RL) in practice, theoretical understanding of its advantage is still limited. In this paper, we theoretically characterize the benefit of representation learning under the low-rank Markov d…

Cited by 28SourcePDFScholar
2022

Relationship Oriented Semantic Scene Understanding for Daily Manipulation Tasks

IROS 2022poster

Assistive robot systems have been developed to help people accomplish daily manipulation tasks especially for those with disabilities, where scene understanding plays a crucial role in enabling robots to interpret the surroundings and behave accordingly. Most of the current systems approach scene un…

Cited by 5SourceScholar
2021

Cost Affinity Learning Network for Stereo Matching

ICASSP 2021accepted

Existing stereo matching methods mainly tend to directly aggregate features output from Convolutional Neural Network to obtain more discriminative cost features, but ignore the affinity of each element in the cost feature which also plays a key role in enhancing the cost feature. In this work, we pr…

Cited by 0SourceScholar
2021

DeepRelativeFusion: Dense Monocular SLAM using Single-Image Relative Depth Prediction

IROS 2021poster

Traditional monocular visual simultaneous localization and mapping (SLAM) algorithms have been extensively studied and proven to reliably recover a sparse structure and camera motion. Nevertheless, the sparse structure is still insufficient for scene interaction, e.g., visual navigation and augmente…

Cited by 13SourceScholar
2021

Frequency Domain Image Translation: More Photo-Realistic, Better Identity-Preserving

ICCV 2021poster

Image-to-image translation has been revolutionized with GAN-based methods. However, existing methods lack the ability to preserve the identity of the source domain. As a result, synthesized images can often over-adapt to the reference domain, losing important structural characteristics and suffering…

Cited by 106PDFcodeScholar
2021

Robust Improvement in 3D Object Landmark Inference for Semantic Mapping

ICRA 2021poster

Recent works on semantic Simultaneous Localization and Mapping (SLAM) utilizing object landmarks have shown superiority in terms of robustness and accuracy in tracking and localization. 3D object landmarks represented by a cubic or quadric surface are inferred from 2D object bounding boxes which are…

Cited by 4SourceScholar
2020

Keypoint Description by Descriptor Fusion Using Autoencoders

ICRA 2020poster

Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo conditional…

Cited by 5SourceScholar
2019

A Comparison of CNN-Based and Hand-Crafted Keypoint Descriptors

ICRA 2019poster

Keypoint matching is an important operation in computer vision and its applications such as visual simultaneous localization and mapping (SLAM) in robotics. This matching operation heavily depends on the descriptors of the keypoints, and it must be performed reliably when images undergo condition ch…

Cited by 34SourceScholar
2019

CNN-SVO: Improving the Mapping in Semi-Direct Visual Odometry Using Single-Image Depth Prediction

ICRA 2019poster

Reliable feature correspondence between frames is a critical step in visual odometry (VO) and visual simultaneous localization and mapping (V-SLAM) algorithms. In comparison with existing VO and V-SLAM algorithms, semi-direct visual odometry (SVO) has two main advantages that lead to state-of-the-ar…

Cited by 137SourcecodeScholar
2019

Implicit Semantic Data Augmentation for Deep Networks

NeurIPS 2019poster

In this paper, we propose a novel implicit semantic data augmentation (ISDA) approach to complement traditional augmentation techniques like flipping, translation or rotation. Our work is motivated by the intriguing property that deep networks are surprisingly good at linearizing features, such that…

2019

Improving Keypoint Matching Using a Landmark-Based Image Representation

ICRA 2019poster

Motivated by the need to improve the performance of visual loop closure verification via multi-view geometry (MVG) under significant illumination and viewpoint changes, we propose a keypoint matching method that uses landmarks as an intermediate image representation in order to leverage the power of…

Cited by 7SourceScholar
2019

Moving Object Detection Under Discontinuous Change in Illumination Using Tensor Low-Rank and Invariant Sparse Decomposition

CVPR 2019poster

Although low-rank and sparse decomposition based methods have been successfully applied to the problem of moving object detection using structured sparsity-inducing norms, they are still vulnerable to significant illumination changes that arise in certain applications. We are interested in moving ob…

Cited by 27PDFScholar
2018

Fast Shadow Detection from a Single Image Using a Patched Convolutional Neural Network

IROS 2018poster

In recent years, various shadow detection methods from a single image have been proposed and used in vision systems; however, most of them are not appropriate for the robotic applications due to the expensive time complexity. This paper introduces a fast shadow detection method using a deep learning…

Cited by 70SourceScholar
2018

Real-Time Segmentation with Appearance, Motion and Geometry

IROS 2018poster

Real-time Segmentation is of crucial importance to robotics related applications such as autonomous driving, driving assisted systems, and traffic monitoring from unmanned aerial vehicles imagery. We propose a novel two-stream convolutional network for motion segmentation, which exploits flow and ge…

Cited by 19SourceScholar
2018

Submap-Based Pose-Graph Visual SLAM: A Robust Visual Exploration and Localization System

IROS 2018poster

For VSLAM (Visual Simultaneous Localization and Mapping), localization is a challenging task, especially for some challenging situations: textureless frames, motion blur, etc. To build a robust exploration and localization system in a given space, a submap-based VSLAM system is proposed in this pape…

Cited by 13SourceScholar
2018

Submap-Based Pose-Graph Visual SLAM: A Robust Visual Exploration and Localization System* The work in this paper is supported by the National Natural Science Foundation of China (61603103, 61673125), the Natural Science Foundation of Guangdong of China (2016A030310293), and the Major Scientific and Technological Special Project of Guangdong of China (2016B090910003)

IROS 2018

For VSLAM (Visual Simultaneous Localization and Mapping), localization is a challenging task, especially for some challenging situations: textureless frames, motion blur, etc. To build a robust exploration and localization system in a given space, a submap-based VSLAM system is proposed in this pape

Cited by 15SourceScholar
2017

Moving Object Detection in Time-Lapse or Motion Trigger Image Sequences Using Low-Rank and Invariant Sparse Decomposition

ICCV 2017poster

Low-rank and sparse representation based methods have attracted wide attention in background subtraction and moving object detection, where moving objects in the scene are modeled as pixel-wise sparse outliers. Since in real scenarios moving objects are also structurally sparse, recently researchers…

Cited by 28PDFScholar
2016

Multi-Scale Patch Aggregation (MPA) for Simultaneous Detection and Segmentation

CVPR 2016oral

Aiming at simultaneous detection and segmentation (SDS), we propose a proposal-free framework, which detect and segment object instances via mid-level patches. We design a unified trainable network on patches, which is followed by a fast and effective patch aggregation algorithm to infer object inst…

Cited by 110PDFScholar
2015

Morphologies and swimming characteristics of rotating magnetic swimmers with soft tails at low Reynolds numbers

IROS 2015poster

Helical microswimmers capable of propulsion at low Reynolds numbers have been proposed for numerous applications. Several different kinds of helical swimmers inspired by E. coli bacteria have been proposed by researchers, and most of them have rigid helical tails. However, high softness could make s…

Cited by 8SourceScholar