← Search

Xu Liu

90 accepted papers

2026

A Closer Look at Knowledge Distillation in Spiking Neural Network Training

AAAI 2026technical

Spiking Neural Networks (SNNs) become popular due to excellent energy efficiency, yet facing challenges for effective model training. Recent works improve this by introducing knowledge distillation (KD) techniques, with the pre-trained artificial neural networks (ANNs) used as teachers and the targe

Cited by 2SourcePDFScholar
2026

DSGCR: Decomposed Spectral Geometry-Aware Cross-Modal Semantic Representation for 3D Visual Grounding

ICML 2026poster

3D visual grounding encompassing 3D referring expression comprehension (3DREC) and segmentation (3DRES) requires robust cross-modal representation to achieve fine-grained semantic alignment and precise geometric reasoning. However, most methods employ unimodal pre-trained encoders that transfer visu…

Cited by 0SourceScholar
2026

Delving Aleatoric Uncertainty in Medical Image Segmentation via Vision Foundation Models

CVPR 2026

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing pervasive data uncertainty that substantially undermines model rob

Cited by 0SourceScholar
2026

Demystifying Robot Diffusion Policies: Action Memorization and a Simple Lookup Table Alternative

ICLR 2026poster

Diffusion policies for visuomotor robot manipulation tasks achieve remarkable dexterity and robustness while only training on a small number of task demonstrations. However, the reason for this performance remains a mystery. In this paper, we offer a surprising hypothesis: diffusion policies essent…

Cited by 0SourcecodeScholar
2026

Denoising as Path Planning: Training-Free Acceleration of Diffusion Models with DPCache

CVPR 2026

Diffusion models have demonstrated remarkable success in image and video generation, yet their practical deployment remains hindered by the substantial computational overhead of multi-step iterative sampling. Among acceleration strategies, caching-based methods offer a training-free and effective so

Cited by 0SourcecodeScholar
2026

Evolving Semantic Propagation for Aerial Semantic 3D Gaussian Splatting

AAAI 2026technical

Semantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal

Cited by 0SourcePDFScholar
2026

FELP: Fast and Effective Autonomous Flight on Large-Scale and Cluttered Environments Based on Unified Linear Parametric Map

ICRA 2026poster

Current AAV autonomous flights exhibit efficient performance in both indoor and field environments. However, they often face significant challenges in large-scale and cluttered environments, where the vast amount of captured data can lead to computation and storage bottlenecks. Additionally, the existin…

Cited by 0SourceScholar
2026

Foreground-Aware Token Routing Vision Transformer for Real-Time Satellite Video Tracking

ICML 2026poster

Real-time satellite video tracking poses distinct challenges, including accommodating high spatial-temporal resolution, dynamic backgrounds, and constrained onboard computational resources. While Discriminative Correlation Filter (DCF)-based methods offer high-speed inference, they suffer from limit…

Cited by 0SourceScholar
2026

HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video Tracking

AAAI 2026technical

In recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to h

Cited by 0SourcePDFScholar
2026

IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments

AAAI 2026technical

Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains unde

Cited by 0SourcePDFScholar
2026

Let’s Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts

AAAI 2026technical

Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thou

Cited by 0SourcePDFScholar
2026

RECS4R: Bridging Semantics and Geometry for Referring Remote Sensing Interpretation

CVPR 2026

Referring expression comprehension and segmentation (RECS) task plays a vital role in remote sensing due to its high efficiency in multi-tasking. However, RECS has reached a performance bottleneck rooted in representational insufficiency, primarily due to cross-task representational fragmentation in

Cited by 0SourcecodeScholar
2026

ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization

IJCAI 2026

Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lies in the non-stationary distribution shift between offline datasets and the evolving online policy. Common approaches often rely on static mixing rati

Cited by 0Scholar
2026

SOAR: Semi-Supervised Open-Vocabulary Aerial Object Detection via Dual-Aware Enhanced Prior Denoising

AAAI 2026technical

Open-Vocabulary Object Detection (OVOD) shows promise in remote sensing (RS), but due to its unique value, there are challenges such as the predominance of background regions, sparse labels, limited semantic information, and difficulties in semi-supervised training. To tackle these challenges, we pr

Cited by 0SourcePDFScholar
2026

VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiency

CVPR 2026

Text-driven 3D editing, enabled by advancements in 3D reconstruction techniques such as NeRF and 3D Gaussian Splatting, aims to provide intuitive scene customization. However, existing methods frequently exhibit limitations in controllability and consistency. To address these shortcomings, we propos

Cited by 0SourceScholar
2025

Adaptive Merchant-Centric Risk Control via Unbiased Decision-Making and Dynamic Optimization in E-Commerce

AAAI 2025technical

In the domain of merchant-oriented risk control decisions within e-commerce, balancing the effectiveness of risk management with merchant satisfaction remains a critical challenge. Strict risk control strategies, while effectively mitigating risks, often lead to increased merchant dissatisfaction. C…

Cited by 0SourcePDFScholar
2025

Asynchronous Collaborative Graph Representation for Frames and Events

CVPR 2025poster

Integrating frames and events has become a widely accepted solution for various tasks in challenging scenarios. However, most multimodal methods directly convert events into image-like formats synchronized with frames and process each stream through separate two-branch backbones, making it difficult…

2025

Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs

COLING 2025main

Chain-of-Thought (CoT) has been a widely adopted prompting method, eliciting impressive reasoning abilities of Large Language Models (LLMs). Inspired by the sequential thought structure of CoT, a number of Chain-of-X (CoX) methods have been developed to address challenges across diverse domains and…

Cited by 22SourcePDFScholar
2025

Bridging Gait Recognition and Large Language Models Sequence Modeling

CVPR 2025poster

Gait sequences exhibit sequential structures and contextual relationships similar to those in natural language, where each element--whether a word or a gait step--is connected to its predecessors and successors. This similarity enables the transformation of gait sequences into "texts" containing ide…

Cited by 1SourcePDFScholar
2025

CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models

ACL 2025long

Investigating hallucination issues in large language models (LLMs) within cross-lingual and cross-modal scenarios can greatly advance the large-scale deployment in real-world applications. Nevertheless, the current studies are limited to a single scenario, either cross-lingual or cross-modal, leavin…

2025

Domain-aware Category-level Geometry Learning Segmentation for 3D Point Clouds

ICCV 2025poster

Domain generalization in 3D segmentation is a critical challenge in deploying models to unseen environments. Current methods mitigate the domain shift by augmenting the data distribution of point clouds. However, the model learns global geometric patterns in point clouds while ignoring the category-…

2025

FELP:Fast and Effective Autonomous Flight on Large-Scale and Cluttered Environments Based on Unified Linear Parametric Map

RA-L 2025

Current UAV autonomous flights exhibit efficient performance in both indoor and field environments. However, they often face significant challenges in large-scale and cluttered environments, where the vast amount of captured data can lead to computation and storage bottlenecks. Additionally, the exi

Cited by 1SourceScholar
2025

FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching

ICCV 2025poster

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live videos with facial retouching. However, previous FAP datasets are either small or closed-source. Moreover, the corresponding FAP models exhibit limited generalization and ada…

2025

Foresight in Motion: Reinforcing Trajectory Prediction with Reward Heuristics

ICCV 2025poster

Motion forecasting for on-road traffic agents presents both a significant challenge and a critical necessity for ensuring safety in autonomous driving systems. In contrast to most existing data-driven approaches that directly predict future trajectories, we rethink this task from a planning perspect…

Cited by 0SourcePDFScholar
2025

F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations

EMNLP 2025

With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified. Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from real-world open-ended interactions. These formats are prone to position bias and

2025

Gait-X: Exploring X modality for Generalized Gait Recognition

ICCV 2025poster

Modality exploration has been repeatedly mentioned in gait recognition, evolving from silhouette to parsing, mesh, point clouds, etc. These latest modalities agree that silhouette is less affected by background and clothing noises, but argue it loses too much valuable discriminative information. The…

Cited by 0SourcePDFScholar
2025

Hierarchical Variational Test-Time Prompt Generation for Zero-Shot Generalization

ICCV 2025poster

Vision-language models like CLIP have demonstrated strong zero-shot generalization, making them valuable for various downstream tasks through prompt learning. However, existing test-time prompt tuning methods, such as entropy minimization, treat both text and visual prompts as fixed learnable parame…

Cited by 0SourcePDFScholar
2025

HyperMST: Multi-scale Spatio-Temporal Hypercorrelation Network for POI Recommendation

ICASSP 2025accepted

Point-of-Interest (POI) recommendation has become increasingly important in the trajectory prediction domain. However, most existing approaches focus on a single scale and tend to overemphasize either spatial or temporal aspects. These methods often overlook the temporal dependencies in movement beh…

Cited by 0SourceScholar
2025

Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing Images

IJCAI 2025

Visual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scen

Cited by 0SourcePDFScholar
2025

Latent Theory of Mind: A Decentralized Diffusion Architecture for Cooperative Manipulation

CoRL 2025oral

We present Latent Theory of Mind (LatentToM), a decentralized diffusion policy architecture for collaborative robot manipulation. Our policy allows multiple manipulators with their own perception and computation to collaborate with each other towards a common task goal with or without explicit commu…

Cited by 0SourceScholar
2025

Logits DeConfusion with CLIP for Few-Shot Learning

CVPR 2025poster

With its powerful visual-language alignment capability, CLIP performs well in zero-shot and few-shot learning tasks. However, we found in experiments that CLIP's logits suffer from serious inter-class confusion problems in downstream tasks, and the ambiguity between categories seriously affects the…

2025

McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

ACL 2025finding

As large language models (LLMs) are increasingly applied to various NLP tasks, their inherent biases are gradually disclosed. Therefore, measuring biases in LLMs is crucial to mitigate its ethical risks. However, most existing bias evaluation datasets are focus on English andNorth American culture,…

Cited by 0SourcePDFScholar
2025

Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

ICML 2025poster

Achieving effective unified pretraining on large time series corpora remains an open challenge in developing time series foundation models. Existing methods, such as Moirai, introduce multiple projection layers for time series of different frequencies to account for high data heterogeneity. We ident…

Cited by 0SourcePDFScholar
2025

Mutual-View Contrastive Generative Framework for Attribute-Missing Graph Clustering

ICASSP 2025accepted

Attribute-Missing Graph Clustering addresses the challenging problem of incomplete node attribute information in graphs. Recent advancements in self-supervised learning techniques, particularly contrastive learning and generative approaches, have shown effectiveness in tackling tasks involving missi…

Cited by 0SourceScholar
2025

Offline-to-Online Reinforcement Learning with Classifier-Free Diffusion Generation

ICML 2025poster

Offline-to-online Reinforcement Learning (O2O RL) aims to perform online fine-tuning on an offline pre-trained policy to minimize costly online interactions. Existing work used offline datasets to generate data that conform to the online data distribution for data augmentation. However, generated da…

Cited by 0SourcePDFScholar
2024

Autost: Training-Free Neural Architecture Search For Spiking Transformers

ICASSP 2024accepted

Spiking Transformers have gained considerable attention because they achieve both the energy efficiency of Spiking Neural Networks (SNNs) and the high capacity of Transformers. However, the existing Spiking Transformer architectures, derived from Artificial Neural Networks (ANNs), exhibit a notable…

Cited by 0SourceScholar
2024

CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading

EMNLP 2024main

The utilization of Large Language Models (LLMs) in financial trading has primarily been concentrated within the stock market, aiding in economic and financial decisions. Yet, the unique opportunities presented by the cryptocurrency market, noted for its on-chain data’s transparency and the critical…

2024

Cut out the Middleman: Revisiting Pose-based Gait Recognition

ECCV 2024poster

"Recent pose-based gait recognition methods, which utilize human skeletons as the model input, have demonstrated significant potential in handling variations in clothing and occlusions. However, methods relying on such skeleton to encode pose are constrained mainly by two problems: (1) poor performa…

2024

Free Lunch for Gait Recognition: A Novel Relation Descriptor

ECCV 2024poster

"Gait recognition is to seek correct matches for query individuals by their unique walking patterns. However, current methods focus solely on extracting individual-specific features, overlooking “interpersonal” relationships. In this paper, we propose a novel Relation Descriptor that captures not on…

Cited by 3SourcePDFScholar
2024

Hallucination Diversity-Aware Active Learning for Text Summarization

NAACL 2024long

Large Language Models (LLMs) have shown propensity to generate hallucinated outputs, i.e., texts that are factually incorrect or unsupported. Existing methods for alleviating hallucinations typically require costly human annotations to identify and correct hallucinations in LLM outputs. Moreover, mo…

Cited by 6SourcePDFScholar
2024

Improving Neural Logic Machines via Failure Reflection

ICML 2024poster

Reasoning is a fundamental ability towards artificial general intelligence (AGI). Fueled by the success of deep learning, the neural logic machines models (NLMs) have introduced novel neural-symbolic structures and demonstrate great performance and generalization on reasoning and decision-making tas…

Cited by 3SourcePDFScholar
2024

Multiplane Prior Guided Few-Shot Aerial Scene Rendering

CVPR 2024poster

Neural Radiance Fields (NeRF) have been successfully applied in various aerial scenes yet they face challenges with sparse views due to limited supervision. The acquisition of dense aerial views is often prohibitive as unmanned aerial vehicles (UAVs) may encounter constraints in perspective range an…

Cited by 3SourcePDFScholar
2024

Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective

ECCV 2024poster

"Extensive occlusions in real-world scenarios pose challenges to gait recognition due to missing and noisy information, as well as body misalignment in position and scale. We argue that rich dynamic contextual information within a gait sequence inherently possesses occlusion-solving traits: 1) Adjac…

2024

Purpose Enhanced Reasoning through Iterative Prompting: Uncover Latent Robustness of ChatGPT on Code Comprehension

IJCAI 2024poster

Code comments are crucial for gaining in-depth insights to facilitate code comprehension. The key to obtaining these insights lies in precisely summarizing the main purpose of the code. Recent approaches on code comment generation lie in prompting large language models (LLMs) such as ChatGPT, instea…

2024

QAGait: Revisit Gait Recognition from a Quality Perspective

AAAI 2024technical

Gait recognition is a promising biometric method that aims to identify pedestrians from their unique walking patterns. Silhouette modality, renowned for its easy acquisition, simple structure, sparse representation, and convenient modeling, has been widely employed in controlled in-the-lab research.…

2024

Time-FFM: Towards LM-Empowered Federated Foundation Model for Time Series Forecasting

NeurIPS 2024poster

Unlike natural language processing and computer vision, the development of Foundation Models (FMs) for time series forecasting is blocked due to data scarcity. While recent efforts are focused on building such FMs by unlocking the potential of language models (LMs) for time series analysis, dedicat…

2024

TreeScope: An Agricultural Robotics Dataset for LiDAR-Based Mapping of Trees in Forests and Orchards

ICRA 2024poster

Data collection for forestry, timber, and agriculture relies on manual techniques which are labor-intensive and time-consuming. We seek to demonstrate that robotics offers improvements over these techniques and can accelerate agricultural research, beginning with semantic segmentation and diameter e…

Cited by 11SourcecodeScholar
2024

ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided Optimization

AAAI 2024technical

Pre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring th…

Cited by 16SourcePDFScholar
2023

Active Collaborative Localization in Heterogeneous Robot Teams

RSS 2023poster

Accurate and robust state estimation is critical for autonomous navigation of robot teams. This task is especially challenging for large groups of size, weight, and power (SWAP) constrained aerial robots operating in perceptually-degraded GPS-denied environments. We can, however, actively increase t…

2023

Active Metric-Semantic Mapping by Multiple Aerial Robots

ICRA 2023poster

Traditional approaches for active mapping focus on building geometric maps. For most real-world applications, however, actionable information is related to semantically meaningful objects in the environment. We propose an approach to the active metric-semantic mapping problem that enables multiple h…

Cited by 24SourceScholar
2023

An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions

CVPR 2023poster

The target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is ofte…

2023

Cooperative Exploration of Heterogeneous UAVs in Mountainous Environments by Constructing Steady Communication

RA-L 2023

Unmanned aerial vehicles (UAVs) must fly at low altitudes to execute certain missions when operating in complex mountainous areas. However, in these environments, UAVs lose their line-of-sight (LOS) communication with the ground station (GS) due to the obstruction of the mountains and are unable to

Cited by 8SourceScholar
2023

Coupling Artificial Neurons in BERT and Biological Neurons in the Human Brain

AAAI 2023technical

Linking computational natural language processing (NLP) models and neural responses to language in the human brain on the one hand facilitates the effort towards disentangling the neural representations underpinning language perception, on the other hand provides neurolinguistics evidence to evaluat…

2023

Curvature-Balanced Feature Manifold Learning for Long-Tailed Classification

CVPR 2023poster

To address the challenges of long-tailed classification, researchers have proposed several approaches to reduce model bias, most of which assume that classes with few samples are weak classes. However, recent studies have shown that tail classes are not always hard to learn, and model bias has been…

Cited by 57SourcePDFScholar
2023

Deciphering Spatio-Temporal Graph Forecasting: A Causal Lens and Treatment

NeurIPS 2023poster

Spatio-Temporal Graph (STG) forecasting is a fundamental task in many real-world applications. Spatio-Temporal Graph Neural Networks have emerged as the most popular method for STG forecasting, but they often struggle with temporal out-of-distribution (OoD) issues and dynamic spatial causation. In t…

2023

Fine-grained Artificial Neurons in Audio-transformers for Disentangling Neural Auditory Encoding

ACL 2023findings

The Wav2Vec and its variants have achieved unprecedented success in computational auditory and speech processing. Meanwhile, neural encoding studies that integrate the superb representation capability of Wav2Vec and link those representations to brain activities have provided novel insights into a f…

2023

LargeST: A Benchmark Dataset for Large-Scale Traffic Forecasting

NeurIPS 2023poster

Road traffic forecasting plays a critical role in smart city initiatives and has experienced significant advancements thanks to the power of deep learning in capturing non-linear patterns of traffic data. However, the promising results achieved on current public datasets may not be applicable to pra…

2023

Robust Localization of Aerial Vehicles via Active Control of Identical Ground Vehicles

IROS 2023poster

This paper addresses the problem of active collaborative localization in heterogeneous robot teams with unknown data association. It involves positioning a small number of identical unmanned ground vehicles (UGVs) at desired positions so that an unmanned aerial vehicle (UAV) can, through unlabelled…

Cited by 4SourceScholar
2022

Experiments in Adaptive Replanning for Fast Autonomous Flight in Forests

ICRA 2022poster

Fast, autonomous flight in unstructured, cluttered environments such as forests is challenging because it requires the robot to compute new plans in realtime on a computationally-constrained platform. In this paper, we enable this capability with a search-based planning framework that adapts samplin…

Cited by 20SourcecodeScholar
2022

Large-Scale Autonomous Flight With Real-Time Semantic SLAM Under Dense Forest Canopy

RA-L 2022

Semantic maps represent the environment using a set of semantically meaningful objects. This representation is storage-efficient, less ambiguous, and more informative, thus facilitating large-scale autonomy and the acquisition of actionable information in highly unstructured, GPS-denied environments

Cited by 99SourceScholar
2021

ASCNet: Self-Supervised Video Representation Learning With Appearance-Speed Consistency

ICCV 2021poster

We study self-supervised video representation learning, which is a challenging task due to 1) sufficient labels for supervision; 2) unstructured and noisy visual information. Existing methods mainly use contrastive loss with video clips as the instances and learn visual representation by discriminat…

Cited by 56PDFScholar
2021

Closed-Loop Pose Control and Automated Suturing of Continuum Surgical Manipulators With Customized Wrist Markers Under Stereo Vision

RA-L 2021

The use of continuum manipulators in surgical applications can be beneficial because of their inherent safety from their structural compliance. However, the tip pose (position and orientation) accuracy of a continuum manipulator can be low when an external load or disturbance is applied. Closed-loop

Cited by 31SourceScholar
2021

Exploring Imitation Learning for Autonomous Driving with Feedback Synthesizer and Differentiable Rasterization

IROS 2021poster

We present a learning-based planner that aims to robustly drive a vehicle by mimicking human drivers’ driving behavior. We leverage a mid-to-mid approach that allows us to manipulate the input to our imitation learning network freely. With that in mind, we propose a novel feedback synthesizer for da…

Cited by 42SourceScholar
2021

Multi-Scale Progressive Attention Network for Video Question Answering

ACL 2021short

Understanding the multi-scale visual information in a video is essential for Video Question Answering (VideoQA). Therefore, we propose a novel Multi-Scale Progressive Attention Network (MSPAN) to achieve relational reasoning between cross-scale video information. We construct clips of different leng…

Cited by 23SourcePDFScholar
2021

Multiresolution Representations for Large-Scale Terrain with Local Gaussian Process Regression

ICRA 2021poster

To address the problem of building accurate and coherent models for large-scale terrains from incomplete and noisy sensor data, this paper proposes a novel framework that can efficiently infer terrain structures by divisionally providing the best linear unbiased estimates for the elevation values. T…

Cited by 3SourceScholar
2021

Place Recognition in Forests With Urquhart Tessellations

RA-L 2021

In this letter, we present a novel descriptor based on Urquhart tessellations derived from the position of trees in a forest. We propose a framework that uses these descriptors to detect previously seen observations and landmark correspondences, even with partial overlap and noise. We run loop closu

Cited by 19SourcecodeScholar
2020

Design and Kinematic Modeling of a Novel Steerable Needle for Image-Guided Insertion

ICRA 2020poster

Needle-based procedures, such as biopsy and percutaneous tumor ablation, highly depend on the accuracy of needle placement. The accuracy is significantly affected by the needle-tissue interaction no matter what needles (straight or steerable) are used. Due to the unknown tissue mechanics, it is chal…

Cited by 2SourceScholar
2020

Gait Lateral Network: Learning Discriminative and Compact Representations for Gait Recognition

ECCV 2020poster

Gait recognition aims at identifying different people by the walking patterns, which can be conducted at a long distance without the cooperation of subjects. A key challenge for gait recognition is to learn representations from the silhouettes that are invariant to the factors such as clothing, carr…

Cited by 231SourcePDFScholar
2020

GaitPart: Temporal Part-Based Model for Gait Recognition

CVPR 2020poster

Gait recognition, applied to identify individual walking patterns in a long-distance, is one of the most promising video-based biometric technologies. At present, most gait recognition methods take the whole human body as a unit to establish the spatio-temporal representations. However, we have obse…

Cited by 529PDFcodeScholar
2020

Group Contextual Encoding for 3D Point Clouds

NeurIPS 2020poster

Global context is crucial for 3D point cloud scene understanding tasks. In this work, we extended the contextual encoding layer that was originally designed for 2D tasks to 3D Point Cloud scenarios. The encoding layer learns a set of code words in the feature space of the 3D point cloud to characte…

2020

SLOAM: Semantic Lidar Odometry and Mapping for Forest Inventory

RA-L 2020

This letter describes an end-to-end pipeline for tree diameter estimation based on semantic segmentation and lidar odometry and mapping. Accurate mapping of this type of environment is challenging since the ground and the trees are surrounded by leaves, thorns and vines, and the sensor typically exp

Cited by 165SourceScholar
2019

Monocular Camera Based Fruit Counting and Mapping With Semantic Data Association

RA-L 2019

In this letter, we present a cheap, lightweight, and fast fruit counting pipeline. Our pipeline relies only on a monocular camera, and achieves counting performance comparable to a state-of-the-art fruit counting system that utilizes an expensive sensor suite including a monocular camera, LiDAR and

Cited by 81SourceScholar
2018

Distributed Submodular Maximization for Large Vocabulary Continuous Speech Recognition

ICASSP 2018accepted

Huge training datasets for automatic speech recognition (ASR) typically contain redundant information so that a subset of data is generally enough to obtain similar ASR performance to that obtained when the entire dataset is employed for training. Although the centralized submodular-based data selec…

Cited by 0SourceScholar
2018

Robust Fruit Counting: Combining Deep Learning, Tracking, and Structure from Motion

IROS 2018poster

We present a novel fruit counting pipeline that combines deep segmentation, frame to frame tracking, and 3D localization to accurately count visible fruits across a sequence of images. Our pipeline works on image streams from a monocular camera, both in natural light, as well as with controlled illu…

Cited by 157SourceScholar