← Search

Chen Feng

89 accepted papers

2026

CRAG: Can 3D Generative Models Help 3D Assembly?

ICML 2026poster

Most existing 3D assembly methods treat the problem as pure pose estimation, rearranging observed parts via rigid transformations. In contrast, human assembly naturally couples structural reasoning with holistic shape inference. Inspired by this intuition, we reformulate 3D assembly as a joint probl…

Cited by 0SourceScholar
2026

Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis

CVPR 2026

Statistically consistent methods based on the noise transition matrix (T) offer a theoretically grounded solution to Learning with Noisy Labels (LNL), with guarantees of convergence to the optimal clean-data classifier. In practice, however, these methods are often outperformed by empirical approach

Cited by 0SourceScholar
2026

Emergent Outlier View Rejection in Visual Geometry Grounded Transformers

CVPR 2026

Reliable 3D reconstruction from in-the-wild image collections is often hindered by noisy images--irrelevant inputs with little or no view overlap with others. While traditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed-forward 3D rec

Cited by 0SourcecodeScholar
2026

Flexible and Efficient Spatio-Temporal Transformer for Sequential Visual Place Recognition

ICRA 2026poster

Sequential Visual Place Recognition (Seq-VPR) leverages transformers to capture spatio-temporal features effectively. In practice, a transformer-based Seq-VPR model should be flexible to the number of frames per sequence (sequence length), deliver fast inference, and use little memory to meet real-t…

2026

Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy

CVPR 2026

The development of affective multimodal language models (MLMs) has long been constrained by a gap between low-level perception and high-level interaction, leading to fragmented affective capabilities and limited generalization. To bridge this gap, we propose a cognitively inspired three-level hierar

Cited by 0SourceScholar
2026

Noisy but Valid: Robust Statistical Evaluation of LLMs with Imperfect Judges

ICLR 2026poster

Reliable certification of Large Language Models (LLMs)—verifying that failure rates are below a safety threshold—is critical yet challenging. While "LLM-as-a-Judge" offers scalability, judge imperfections, noise, and bias can invalidate statistical guarantees. We introduce a "Noisy but Valid" hypoth…

Cited by 0SourceScholar
2026

SAVMap: Structure-Aided Visual Mapping of Large-Scale 2.5D Manhattan Wireframes from Panoramic Video

ICRA 2026poster

Precise 3D representations of industrial environments enable tasks such as robot localization and digital twin generation. We propose SAVMap, a method for generating a wireframe map of warehouse shelf and light structures using only a panoramic video camera as the sensor input. Sequences of rectifie…

2026

Thinking in 360deg: Humanoid Visual Search in the Wild

CVPR 2026

Humans rely on the synergistic control of head (cephalomotor) and eye (oculomotor) to efficiently search for visual information in 360deg. However, prior approaches to visual search are limited to a static image, neglecting the physical embodiment and its interaction with the 3D world. How can we de

Cited by 0SourcecodeScholar
2026

Visual-Auditory Proprioception of Soft Finger Shape and Contact

ICRA 2026poster

Soft robotic fingers require precise proprioception of both global deformation and local contact to enable safe and dexterous manipulation. Vision-based methods can reconstruct overall shape but struggle under severe occlusion, while audio-only approaches provide complementary cues but lack spatial …

Cited by 0codeScholar
2026

Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI

CVPR 2026

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded interaction in complex, open-world urban environments. Although rece

Cited by 0SourcecodeScholar
2025

Adversarial Exploitation of Data Diversity Improves Visual Localization

ICCV 2025poster

Visual localization, which estimates a camera's pose within a known scene, is a fundamental capability for autonomous systems. While absolute pose regression (APR) methods have shown promise for efficient inference, they often struggle with generalization. Recent approaches attempt to address this t…

Cited by 0SourcePDFScholar
2025

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

CVPR 2025poster

Navigating dynamic urban environments presents significant challenges for embodied agents, requiring advanced spatial reasoning and adherence to common-sense norms. Despite progress, existing visual navigation methods struggle in map-free or off-street settings, limiting the deployment of autonomous…

2025

Extrapolated Urban View Synthesis Benchmark

ICCV 2025poster

Photorealistic simulators are essential for the training and evaluation of vision-centric autonomous vehicles (AVs). At their core is Novel View Synthesis (NVS), a crucial capability that generates diverse unseen viewpoints to accommodate the broad and continuous pose distribution of AVs. Recent adv…

2025

Fusionsense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction

ICRA 2025

Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observ

Cited by 6SourceScholar
2025

GARF: Learning Generalizable 3D Reassembly for Real-World Fractures

ICCV 2025poster

3D reassembly is a challenging spatial intelligence task with broad applications across scientific domains. While large-scale synthetic datasets have fueled promising learning-based approaches, their generalizability to different domains is limited. Critically, it remains uncertain whether models tr…

Cited by 0SourcePDFScholar
2025

NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments

ICRA 2025

Visual place recognition (VPR) enables autonomous robots to identify previously visited locations, which contributes to tasks like simultaneous localization and mapping (SLAM). VPR faces challenges such as accurate image neighbor retrieval and appearance change in scenery. Event cameras, also known

Cited by 8SourcecodeScholar
2025

OmniDraft: A cross-vocabulary, online adaptive drafter for on-device speculative decoding

NeurIPS 2025poster

Speculative decoding generally dictates having a small, efficient draft model that is either pretrained or distilled offline to a particular target model series, for instance, Llama or Qwen models. However, within online deployment settings, there are two major challenges: 1) usage of a target model…

Cited by 0SourceScholar
2025

PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks

AAAI 2025technical

It is widely known that state-of-the-art machine learning models, including vision and language models, can be seriously compromised by adversarial perturbations. It is therefore increasingly relevant to develop capabilities to certify their performance in the presence of the most effective adversar…

Cited by 0SourcePDFScholar
2025

Robot-Based Automatic Charging for Electric Vehicles Using Incremental Learning and Biomimetic Control

ICRA 2025

With the growing popularity of electric vehicles, the demand for robot-based unmanned automatic charging has become both urgent and challenging. Two key challenges need to be addressed: how to efficiently locate the charging port, and how to compliantly insert the connector into the port. In this pa

Cited by 0SourceScholar
2025

Self-Supervised Place Recognition by Refining Temporal and Featural Pseudo Labels From Panoramic Data

RA-L 2025

Visual place recognition (VPR) using deep networks has achieved state-of-the-art performance. However, most of them require a training set with ground truth sensor poses to obtain positive and negative samples of each observation's spatial neighborhood for supervised learning. When such information

Cited by 6SourcecodeScholar
2025

Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings

ICLR 2025poster

Recent advancements in brain-computer interfaces (BCIs) and deep learning have made decoding lexical tones from intracranial recordings possible, providing the potential to restore the communication ability of speech-impaired tonal language speakers. However, data heterogeneity induced by both physi…

Cited by 0SourcePDFScholar
2025

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

IROS 2025

Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interp

Cited by 45SourcecodeScholar
2024

Biomedical Knowledge Graph Embedding with Householder Projection (Student Abstract)

AAAI 2024technical

Researchers have applied knowledge graph embedding (KGE) techniques with advanced neural network techniques, such as capsule networks, for predicting drug-drug interactions (DDIs) and achieved remarkable results. However, most ignore molecular structure and position features between drug pairs. They…

Cited by 1SourcePDFScholar
2024

Collaborative Multi-Object Tracking With Conformal Uncertainty Propagation

RA-L 2024

Object detection and multiple object tracking (MOT) are essential components of self-driving systems. Accurate detection and uncertainty quantification are both critical for onboard modules, such as perception, prediction, and planning, to improve the safety and robustness of autonomous vehicles. Co

Cited by 44SourceScholar
2024

EgoPAT3Dv2: Predicting 3D Action Target from 2D Egocentric Vision for Human-Robot Interaction

ICRA 2024poster

A robot’s ability to anticipate the 3D action target location of a hand’s movement from egocentric videos can greatly improve safety and efficiency in human-robot interaction (HRI). While previous research predominantly focused on semantic action classification or 2D target region prediction, we arg…

Cited by 2SourceScholar
2024

FC-Planner: A Skeleton-guided Planning Framework for Fast Aerial Coverage of Complex 3D Scenes

ICRA 2024poster

3D coverage path planning for UAVs is a crucial problem in diverse practical applications. However, existing methods have shown unsatisfactory system simplicity, computation efficiency, and path quality in large and complex scenes. To address these challenges, we propose FC-Planner, a skeleton-guide…

Cited by 12SourcecodeScholar
2024

LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition

CVPR 2024poster

In this work we focus on learning facial representations that can be adapted to train effective face recognition models particularly in the absence of labels. Firstly compared with existing labelled face datasets a vastly larger magnitude of unlabeled faces exists in the real world. We explore the l…

2024

LUWA Dataset: Learning Lithic Use-Wear Analysis on Microscopic Images

CVPR 2024highlight

Lithic Use-Wear Analysis (LUWA) using microscopic images is an underexplored vision-for-science research area. It seeks to distinguish the worked material which is critical for understanding archaeological artifacts material interactions tool functionalities and dental records. However this challeng…

Cited by 4SourcePDFScholar
2024

LiDAR-based 4D Occupancy Completion and Forecasting

IROS 2024poster

Scene completion and forecasting are two popular perception problems in research for mobile agents like autonomous vehicles. Existing approaches treat the two problems in isolation, resulting in a separate perception of the two aspects. In this paper, we introduce a novel LiDAR perception task of Oc…

Cited by 18SourcecodeScholar
2024

MTKD: Multi-Teacher Knowledge Distillation for Image Super-Resolution

ECCV 2024poster

"Knowledge distillation (KD) has emerged as a promising technique in deep learning, typically employed to enhance a compact student network through learning from their high-performance but more complex teacher variant. When applied in the context of image super-resolution, most KD approaches are mod…

2024

Memorize What Matters: Emergent Scene Decomposition from Multitraverse

NeurIPS 2024spotlight

Humans naturally retain memories of permanent elements, while ephemeral moments often slip through the cracks of memory. This selective retention is crucial for robotic perception, localization, and mapping. To endow robots with this capability, we introduce 3D Gaussian Mapping (3DGM), a self-superv…

2024

Multiagent Multitraversal Multimodal Self-Driving: Open MARS Dataset

CVPR 2024poster

Large-scale datasets have fueled recent advancements in AI-based autonomous vehicle research. However these datasets are usually collected from a single vehicle's one-time pass of a certain location lacking multiagent interactions or repeated traversals of the same place. Such information could lead…

2024

NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic Annotation

ICRA 2024poster

Visual Place Recognition (VPR) in indoor environments is beneficial to humans and robots for better localization and navigation. It is challenging due to appearance changes at various frequencies, and difficulties of obtaining ground truth metric trajectories for training and evaluation. This paper…

Cited by 2SourceScholar
2024

OmniNxt: A Fully Open-source and Compact Aerial Robot with Omnidirectional Visual Perception

IROS 2024poster

Adopting omnidirectional Field of View (FoV) cameras in aerial robots vastly improves perception ability, significantly advancing aerial robotics’s capabilities in inspection, reconstruction, and rescue tasks. However, such sensors also elevate system complexity, e.g., hardware design, and correspon…

Cited by 9SourcecodeScholar
2024

Robust Collaborative Perception without External Localization and Clock Devices

ICRA 2024poster

A consistent spatial-temporal coordination across multiple agents is fundamental for collaborative perception, which seeks to improve perception abilities through information exchange among agents. To achieve this spatial-temporal alignment, traditional methods depend on external devices to provide…

Cited by 4SourceScholar
2024

Roofus: Learning-based Robotic Moisture Mapping on Flat Rooftops with Ground Penetrating Radar

IROS 2024poster

Robust moisture detection is crucial for building maintenance and cost reduction. Current methods are often limited by the type of roofing material or are cumbersome and expensive. Ground Penetrating Radar (GPR) has shown promise in recent works in moisture detection due to its effectiveness across…

Cited by 0SourceScholar
2024

SOAR: Simultaneous Exploration and Photographing with Heterogeneous UAVs for Fast Autonomous Reconstruction

IROS 2024poster

Unmanned Aerial Vehicles (UAVs) have gained significant popularity in scene reconstruction. This paper presents SOAR, a LiDAR-Visual heterogeneous multi-UAV system specifically designed for fast autonomous reconstruction of complex environments. Our system comprises a LiDAR-equipped explorer with a…

Cited by 3SourcecodeScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

Star-Searcher: A Complete and Efficient Aerial System for Autonomous Target Search in Complex Unknown Environments

RA-L 2024

This paper tackles the challenge of autonomous target search using unmanned aerial vehicles (UAVs) in complex unknown environments. To fill the gap in systematic approaches for this task, we introduce Star-Searcher, an aerial system featuring specialized sensor suites, mapping, and planning modules

Cited by 36SourcecodeScholar
2024

Stepping Forward on the Last Mile

NeurIPS 2024poster

Continuously adapting pre-trained models to local data on resource constrained edge devices is the \emph{last mile} for model deployment. However, as models increase in size and depth, backpropagation requires a large amount of memory, which becomes prohibitive for edge devices. In addition, most ex…

Cited by 1SourcePDFScholar
2024

Temporal Knowledge Graph Embedding using Householder Transformations

ICASSP 2024accepted

The rapid development of Knowledge Graph (KG) technology has led to the emergence of Temporal Knowledge Graphs (TKGs), which hold significant research importance and value. Temporal Knowledge Graph Embedding (TKGE) techniques complement TKGs and predict links within them. The efficacy of TKGE hinges…

Cited by 0SourceScholar
2024

VIRL: Self-Supervised Visual Graph Inverse Reinforcement Learning

CoRL 2024poster

Learning dense reward functions from unlabeled videos for reinforcement learning exhibits scalability due to the vast diversity and quantity of video resources. Recent works use visual features or graph abstractions in videos to measure task progress as rewards, which either deteriorate in unseen do…

Cited by 0SourceScholar
2023

Among Us: Adversarially Robust Collaborative Perception by Consensus

ICCV 2023poster

Multiple robots could perceive a scene (e.g., detect objects) collaboratively better than individuals, although easily suffer from adversarial attacks when using deep learning. This could be addressed by the adversarial defense, but its training requires the often-unknown attacking mechanism. Differ…

Cited by 33PDFcodeScholar
2023

AutoTrans: A Complete Planning and Control Framework for Autonomous UAV Payload Transportation

RA-L 2023

The robotics community is increasingly interested in autonomous aerial transportation. Unmanned aerial vehicles with suspended payloads have advantages over other systems, including mechanical simplicity and agility, but pose great challenges in planning and control. To realize fully autonomous aeri

Cited by 47SourceScholar
2023

Boosting UAV Tracking With Voxel-Based Trajectory-Aware Pre-Training

RA-L 2023

Siamese network-based object tracking has remarkably promoted the automatic capability for highly-maneuvered unmanned aerial vehicles (UAVs). However, the leading-edge tracking framework often depends on template matching, making it trapped when facing multiple views of object in consecutive frames.

Cited by 9SourceScholar
2023

Concavity-Induced Distance for Unoriented Point Cloud Decomposition

RA-L 2023

We propose Concavity-induced Distance (CID) as a novel way to measure the dissimilarity between a pair of points in an unoriented point cloud. CID indicates the likelihood of two points or two sets of points belonging to different convex parts of an underlying shape represented as a point cloud. Aft

Cited by 0SourcecodeScholar
2023

DeepMapping2: Self-Supervised Large-Scale LiDAR Map Optimization

CVPR 2023poster

LiDAR mapping is important yet challenging in self-driving and mobile robotics. To tackle such a global point cloud registration problem, DeepMapping converts the complex map estimation into a self-supervised training of simple deep networks. Despite its broad convergence range on small datasets, De…

2023

Learning Simultaneous Navigation and Construction in Grid Worlds

ICLR 2023poster

We propose to study a new learning task, mobile construction, to enable an agent to build designed structures in 1/2/3D grid worlds while navigating in the same evolving environments. Unlike existing robot learning tasks such as visual navigation and object manipulation, this task is challenging bec…

2023

MacFormer: Map-Agent Coupled Transformer for Real-Time and Robust Trajectory Prediction

RA-L 2023

Predicting the future behavior of agents is a fundamental task in autonomous vehicle domains. Accurate prediction relies on comprehending the surrounding map, which significantly regularizes agent behaviors. However, existing methods have limitations in exploiting the map and exhibit a strong depend

Cited by 79SourceScholar
2023

Metric-Free Exploration for Topological Mapping by Task and Motion Imitation in Feature Space

RSS 2023poster

We propose DeepExplorer, a simple and lightweight metric-free exploration method for topological mapping of unknown environments. It performs task and motion planning (TAMP) entirely in image feature space. The task planner is a recurrent network using the latest image observation sequence to halluc…

2023

PredRecon: A Prediction-boosted Planning Framework for Fast and High-quality Autonomous Aerial Reconstruction

ICRA 2023poster

Autonomous UAV path planning for 3D reconstruction has been actively studied in various applications for high-quality 3D models. However, most existing works have adopted explore-then-exploit, prior-based or exploration-based strategies, demonstrating inefficiency with repeated flight and low autono…

Cited by 25SourcecodeScholar
2023

Robust Collaborative 3D Object Detection in Presence of Pose Errors

ICRA 2023poster

Collaborative 3D object detection exploits information exchange among multiple agents to enhance accuracy of object detection in presence of sensor impairments such as occlusion. However, in practice, pose estimation errors due to imperfect localization would cause spatial message misalignment and s…

Cited by 116SourcecodeScholar
2023

Toward Zero-Shot Sim-to-Real Transfer Learning for Pneumatic Soft Robot 3D Proprioceptive Sensing

ICRA 2023poster

Pneumatic soft robots present many advantages in manipulation tasks. Notably, their inherent compliance makes them safe and reliable in unstructured and fragile environments. However, full-body shape sensing for pneumatic soft robots is challenging because of their high degrees of freedom and comple…

Cited by 18SourcecodeScholar
2023

Uncertainty Quantification of Collaborative Detection for Self-Driving

ICRA 2023poster

Sharing information between connected and autonomous vehicles (CAVs) fundamentally improves the performance of collaborative object detection for self-driving. However, CAVs still have uncertainties on object detection due to practical challenges, which will affect the later modules in self-driving…

Cited by 69SourcecodeScholar
2023

VoxFormer: Sparse Voxel Transformer for Camera-Based 3D Semantic Scene Completion

CVPR 2023highlight

Humans can easily imagine the complete 3D geometry of occluded objects and scenes. This appealing ability is vital for recognition and understanding. To enable such capability in AI systems, we propose VoxFormer, a Transformer-based semantic scene completion framework that can output complete 3D vol…

2022

A Deep Reinforcement Learning Environment for Particle Robot Navigation and Object Manipulation

ICRA 2022poster

Particle robots are novel biologically-inspired robotic systems where locomotion can be achieved collectively and robustly, but not independently. While its control is currently limited to a hand-crafted policy for basic locomotion tasks, such a multi-robot system could be potentially controlled via…

Cited by 7SourceScholar
2022

Egocentric Prediction of Action Target in 3D

CVPR 2022poster

We are interested in anticipating as early as possible the target location of a person's object manipulation action in a 3D workspace from egocentric vision. It is important in fields like human-robot collaboration, but has not yet received enough attention from vision and learning communities. To s…

Cited by 17PDFScholar
2022

Multi-Robot Scene Completion: Towards Task-Agnostic Collaborative Perception

CoRL 2022poster

Collaborative perception learns how to share information among multiple robots to perceive the environment better than individually done. Past research on this has been task-specific, such as detection or segmentation. Yet this leads to different information sharing for different tasks, hindering th…

Cited by 55SourcecodeScholar
2022

Represent, Compare, and Learn: A Similarity-Aware Framework for Class-Agnostic Counting

CVPR 2022poster

Class-agnostic counting (CAC) aims to count all instances in a query image given few exemplars. A standard pipeline is to extract visual features from exemplars and match them with query images to infer object counts. Two essential components in this pipeline are feature representation and similarit…

Cited by 109PDFcodeScholar
2022

Self-Supervised Spatial Reasoning on Multi-View Line Drawings

CVPR 2022poster

Spatial reasoning on multi-view line drawings by state-of-the-art supervised deep networks is recently shown with puzzling low performances on the SPARE3D dataset. Based on the fact that self-supervised learning is helpful when a large number of data are available, we propose two self-supervised lea…

Cited by 4PDFcodeScholar
2022

V2X-Sim: Multi-Agent Collaborative Perception Dataset and Benchmark for Autonomous Driving

RA-L 2022

Vehicle-to-everything (V2X) communication techniques enable the collaboration between vehicles and many other entities in the neighboring environment, which could fundamentally improve the perception system for autonomous driving. However, the lack of a public dataset significantly restricts the res

Cited by 346SourceScholar
2021

Fooling LiDAR Perception via Adversarial Trajectory Perturbation

ICCV 2021poster

LiDAR point clouds collected from a moving vehicle are functions of its trajectories, because the sensor motion needs to be compensated to avoid distortions. When autonomous vehicles are sending LiDAR point clouds to deep networks for perception and planning, could the motion compensation consequent…

Cited by 66PDFcodeScholar
2021

Learning Distilled Collaboration Graph for Multi-Agent Perception

NeurIPS 2021poster

To promote better performance-bandwidth trade-off for multi-agent perception, we propose a novel distilled collaboration graph (DiscoGraph) to model trainable, pose-aware, and adaptive collaboration among agents. Our key novelties lie in two aspects. First, we propose a teacher-student framework to…

2021

Mobile 3D Printing Robot Simulation with Viscoelastic Fluids

IROS 2021poster

The system design and algorithm development of mobile 3D printing robots need a realistic simulation. They require a mobile robot simulation platform to interoperate with a physics-based material simulation for handling interactions between the time-variant deformable 3D printing materials and other…

Cited by 7SourceScholar
2021

NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization Influences

IROS 2021poster

Visual place recognition (VPR) is critical in not only localization and mapping for autonomous driving vehicles, but also assistive navigation for the visually impaired population. To enable a long-term VPR system on a large scale, several challenges need to be addressed. First, different applicatio…

Cited by 14SourceScholar
2021

Siamese Anchor Proposal Network for High-Speed Aerial Tracking

ICRA 2021poster

In the domain of visual tracking, most deep learning-based trackers highlight the accuracy but casting aside efficiency. Therefore, their real-world deployment on mobile platforms like the unmanned aerial vehicle (UAV) is impeded. In this work, a novel two-stage Siamese network-based method is propo…

Cited by 94SourcecodeScholar
2020

A Robust Speaker Clustering Method Based on Discrete Tied Variational Autoencoder

ICASSP 2020accepted

Recently, the speaker clustering model based on aggregation hierarchy cluster (AHC) is a common method to solve two main problems: no preset category number clustering and fix category number clustering. In general, model takes features like i-vectors as input of probability and linear discriminant…

Cited by 0SourceScholar
2020

Automatic Failure Recovery and Re-Initialization for Online UAV Tracking with Joint Scale and Aspect Ratio Optimization

IROS 2020poster

Current unmanned aerial vehicle (UAV) visual tracking algorithms are primarily limited with respect to: (i) the kind of size variation they can deal with, (ii) the implementation speed which hardly meets the real-time requirement. In this work, a real-time UAV tracking algorithm with powerful size e…

Cited by 12SourcecodeScholar
2020

DR2Track: Towards Real-Time Visual Tracking for UAV via Distractor Repressed Dynamic Regression

IROS 2020poster

Visual tracking has yielded promising applications with unmanned aerial vehicle (UAV). In literature, the advanced discriminative correlation filter (DCF) type trackers generally distinguish the foreground from the background with a learned regressor which regresses the implicit circulated samples i…

Cited by 13SourceScholar
2020

LUVLi Face Alignment: Estimating Landmarks' Location, Uncertainty, and Visibility Likelihood

CVPR 2020poster

Modern face alignment methods have become quite accurate at predicting the locations of facial landmarks, but they do not typically estimate the uncertainty of their predicted locations nor predict whether landmarks are visible. In this paper, we present a novel framework for jointly predicting land…

Cited by 196PDFcodeScholar
2020

P2B: Point-to-Box Network for 3D Object Tracking in Point Clouds

CVPR 2020oral

Towards 3D object tracking in point clouds, a novel point-to-box network termed P2B is proposed in an end-to-end learning manner. Our main idea is to first localize potential target centers in 3D search area embedded with target information. Then point-driven 3D target proposal and verification are…

Cited by 199PDFcodeScholar
2020

Real-Time Soft Body 3D Proprioception via Deep Vision-Based Sensing

RA-L 2020

Soft bodies made from flexible and deformable materials are popular in many robotics applications, but their proprioceptive sensing has been a long-standing challenge. In other words, there has hardly been a method to measure and model the high-dimensional 3D shapes of soft bodies with internal sens

Cited by 48SourcecodeScholar
2020

Regularizing Neural Networks via Minimizing Hyperspherical Energy

CVPR 2020poster

Inspired by the Thomson problem in physics where the distribution of multiple propelling electrons on a unit sphere can be modeled via minimizing some potential energy, hyperspherical energy minimization has demonstrated its potential in regularizing neural networks and improving their generalizatio…

Cited by 34PDFScholar
2020

SPARE3D: A Dataset for SPAtial REasoning on Three-View Line Drawings

CVPR 2020poster

Spatial reasoning is an important component of human intelligence. We can imagine the shapes of 3D objects and reason about their spatial relations by merely looking at their three-view line drawings in 2D, with different levels of competence. Can deep networks be trained to perform spatial reasonin…

Cited by 20PDFcodeScholar
2019

An assistive low-vision platform that augments spatial cognition through proprioceptive guidance: Point-to-Tell-and-Touch

IROS 2019poster

Spatial cognition, as gained through the sense of vision, is one of the most important capabilities of human beings. However, for the visually impaired (VI), lack of this perceptual capability poses great challenges in their life. Therefore, we have designed Point-to-Tell-and-Touch, a wearable syste…

Cited by 8SourceScholar
2018

Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling

CVPR 2018poster

Unlike on images, semantic learning on 3D point clouds using a deep network is challenging due to the naturally unordered data structure. Among existing works, PointNet has achieved promising results by directly learning on point sets. However, it does not take full advantage of a point's local neig…

Cited by 635SourcePDFScholar
2018

Simultaneous Edge Alignment and Learning

ECCV 2018poster

Edge detection is among the most fundamental vision problems for its role in perceptual grouping and its wide applications. Recent advances in representation learning have led to considerable improvements in this area. Many state of the art edge detection models are learned with fully convolutional…

Cited by 109SourcePDFScholar
2018

VLASE: Vehicle Localization by Aggregating Semantic Edges

IROS 2018poster

We propose VLASE, a framework to use semantic edge features from images to achieve on-road localization. Semantic edge features denote edge contours that separate pairs of distinct objects such as building-sky, road-sidewalk, and building-ground. While prior work has shown promising results by utili…

Cited by 53SourceScholar
2017

Contour-enhanced resampling of 3D point clouds via graphs

ICASSP 2017accepted

To reduce storage and computational cost for processing and visualizing large-scale 3D point clouds, an efficient resampling strategy is needed to select a representative subset of 3D points that can preserve contours in the original 3D point cloud. We tackle this problem by using graph-based techni…

Cited by 0SourceScholar
2016

A fully automated robotic system for three-dimensional cell rotation

ICRA 2016

Injection and extraction of materials (e.g. protein, sperms, DNA, and blastomeres) into and from cells are essential operation for In Vitro Fertilisation (IVF), intracytoplasmic sperm injection (ICSI) and preimplantation genetic diagnosis (PGD). In order to perform the injection and extraction witho

Cited by 9SourceScholar