← Search

Yan Huang

85 accepted papers

2026

A Wearable Isokinetic Training Robot for Enhanced Bedside Knee Rehabilitation

ICRA 2026poster

Knee pain is prevalent in over 20% of the population, limiting the mobility of those affected. In turn, isokinetic dynamometers and robots have been used to facilitate rehabilitation for those still capable of ambulation. However, there are at most only a few wearable robots capable of delivering is…

Cited by 0SourceScholar
2026

Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

CVPR 2026

Modern industrial quality control heavily relies on automated anomaly detection. While few-shot anomaly detection addresses the challenge of limited labeled data, real-world inspection faces a vast diversity of anomaly types, sizes, and shapes. We identify the primary cause for the anomaly detection

Cited by 0SourceScholar
2026

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

AAAI 2026technical

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further

Cited by 0SourcePDFScholar
2026

Gait Transformer: End-to-End Transformer Backbone for Gait Recognition

AAAI 2026technical

Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modali

Cited by 0SourcePDFScholar
2026

OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding

ICLR 2026poster

We introduce spatio-temporal omni-object video grounding, dubbed $\textbf{OmniSTVG}$, a new STVG task aiming to localize spatially and temporally all targets mentioned in the textual query within videos. Compared to classic STVG locating only a single target, OmniSTVG enables localization of not onl…

Cited by 0SourcecodeScholar
2026

VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation

RA-L 2026

When performing 3D manipulation tasks, robots have to execute action planning based on perceptions from multiple fixed cameras. The multi-camera setup introduces substantial redundancy and irrelevant information, which increases computational costs and forces the model to spend extra training time e

Cited by 2SourcecodeScholar
2026

VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation

ICRA 2026poster

When performing 3D manipulation tasks, robots have to execute action planning based on perceptions from multiple fixed cameras. The multi-camera setup introduces substantial redundancy and irrelevant information, which increases computational costs and forces the model to spend extra training time e…

2025

AirTouch: A Low-Cost Versatile Visuotactile Feedback System for Enhanced Robotic Teleoperation

IROS 2025

Vision-based teleoperation systems are widely used due to their cost-effectiveness and intuitive operation. However, these systems often suffer from challenges such as hand occlusions, environmental variability, and the lack of tactile feedback, limiting their precision and applicability in complex

Cited by 0SourceScholar
2025

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

NeurIPS 2025poster

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully levera…

Cited by 0SourcecodeScholar
2025

Chemistry3D: Robotic Interaction Toolkit for Chemistry Experiments

ICRA 2025

The advent of simulation engines has revolutionized learning and operational efficiency for robots, offering cost-effective and swift pipelines. However, the lack of a universal simulation platform tailored for chemical scenarios impedes progress in robotic manipulation and visualization of reaction

Cited by 2SourcecodeScholar
2025

CoCoL: A Communication Efficient Decentralized Collaborative Learning Method for Multi-Robot Systems

IROS 2025

Collaborative learning enhances the performance and adaptability of multi-robot systems in complex tasks but faces significant challenges due to high communication overhead and data heterogeneity inherent in multi-robot tasks. To this end, we propose CoCoL, a Communication efficient decentralized Co

Cited by 1SourceScholar
2025

DAM: Dynamic Attention Mask for Long-Context Large Language Model Inference Acceleration

ACL 2025finding

Long-context understanding is crucial for many NLP applications, yet transformers struggle with efficiency due to the quadratic complexity of self-attention. Sparse attention methods alleviate this cost but often impose static, predefined masks, failing to capture heterogeneous attention patterns. T…

2025

DATA: Domain-And-Time Alignment for High-Quality Feature Fusion in Collaborative Perception

ICCV 2025poster

Feature-level fusion shows promise in collaborative perception (CP) through balanced performance and communication bandwidth trade-off. However, its effectiveness critically relies on input feature quality. The acquisition of high-quality features faces domain gaps from hardware diversity and deploy…

Cited by 0SourcePDFScholar
2025

Depth Restoration of Hand-Held Transparent Objects for Human-to-Robot Handover

ICRA 2025

Transparent objects are common in daily life, while their optical properties pose challenges for RGB-D cameras to capture accurate depth information. This issue is further amplified when these objects are hand-held, as hand occlusions further complicate depth estimation. For assistant robots, howeve

Cited by 3SourceScholar
2025

Dyn-D^2P: Dynamic Differentially Private Decentralized Learning with Provable Utility Guarantee

IJCAI 2025

Most existing decentralized learning methods with differential privacy (DP) guarantee rely on constant gradient clipping bounds and fixed-level DP Gaussian noises for each node throughout the training process, leading to a significant accuracy degradation compared to non-private counterparts. In thi

Cited by 0SourcePDFScholar
2025

EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow

ICCV 2025poster

Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate this issue, they remain limited to scenarios involving rigid objects with clear displacement and minimal occlusion. In th…

2025

Enhanced Visual-Semantic Interaction with Tailored Prompts for Pedestrian Attribute Recognition

CVPR 2025highlight

Pedestrian attribute recognition (PAR) seeks to predict multiple semantic attributes associated with a specific pedestrian. There are two types of approaches for PAR: unimodal framework and bimodal framework. The former one is to seek a robust visual feature. However, the lack of exploiting semantic…

Cited by 0SourcePDFScholar
2025

GR-MG: Leveraging Partially-Annotated Data via Multi-Modal Goal-Conditioned Policy

RA-L 2025

The robotics community has consistently aimed to achieve generalizable robot manipulation with flexible natural language instructions. One primary challenge is that obtaining robot trajectories fully annotated with both actions and texts is time-consuming and labor-intensive. However, partially-anno

Cited by 39SourcecodeScholar
2025

Glance2Gaze: Efficient Vision-Language Models from Glance Fusion to Gaze Compression

NeurIPS 2025poster

Vision-language models heavily rely on visual representations, yet ensuring its efficiency remains a critical challenge. Most existing approaches focus on reducing visual tokens either at the visual encoder phase or during the LLM decoder stage. Inspired by human visual cognition, where an initial g…

Cited by 0SourceScholar
2025

HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection

AAAI 2025technical

Millimeter-wave radar plays a vital role in 3D object detection for autonomous driving due to its all-weather and all-lighting-condition capabilities for perception. However, radar point clouds suffer from pronounced sparsity and unavoidable angle estimation errors. To address these limitations, inc…

2025

Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

ICLR 2025oral

Transformer has attracted increasing interest in spatio-temporal video grounding, or STVG, owing to its end-to-end pipeline and promising result. Existing Transformer-based STVG approaches often leverage a set of object queries, which are initialized simply using zeros and then gradually learn targe…

2025

Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation

AAAI 2025technical

Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks…

2025

Multimodal Dialogue Emotion Recognition Based on Label Optimization and Coarse-Grained Assisted Fine-Grained

ICASSP 2025accepted

Multimodal dialogue emotion recognition integrates data from multiple modalities to accurately identify emotional states in conversations. However, differences in expression and information density across modalities complicate the fusion of features. Traditional methods may introduce redundant infor…

Cited by 0SourceScholar
2025

RoLocMe: A Robust Multi-agent Source Localization System with Learning-based Map Estimation

IJCAI 2025

This paper addresses the source localization problem by introducing RoLocMe, a multi-agent reinforcement learning system that integrates SkipNet - a skip-connection-based RSS estimation model - with parallel Q-learning. SkipNet predicts RSS propagation of the entire search region, enabling agents to

Cited by 0SourcePDFScholar
2025

Target word activity detector: An approach to obtain ASR word boundaries without lexicon

ICASSP 2025accepted

Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalabi…

Cited by 0SourceScholar
2025

UltraTac: Integrated Ultrasound-Augmented Visuotactile Sensor for Enhanced Robotic Perception

IROS 2025

Visuotactile sensors provide high-resolution tactile information but are incapable of perceiving the material features of objects. We present UltraTac, an integrated sensor that combines visuotactile imaging with ultrasound sensing through a coaxial optoacoustic architecture. The design shares struc

Cited by 1SourceScholar
2025

Zero-Shot Low-Light Image Enhancement via Latent Diffusion Models

AAAI 2025technical

Low-light image enhancement (LLIE) aims to improve visibility and signal-to-noise ratio in images captured under poor lighting conditions. While deep learning has shown promise in this domain, current approaches require extensive paired training data, limiting their practical utility. We present a n…

2024

A Riemannian-Based Joint Design Framework of Mimo Radar Transmit Waveform And Receive Filter Via Information Theory

ICASSP 2024accepted

In this paper, we explore the joint design of a transmit waveform and receive filter to enhance the detection performance of multiple-input multiple-output (MIMO) radar. Target echoes are assumed to be embedded in signal-dependent interference and colored Gaussian noise. As design metrics, we exploi…

Cited by 0SourceScholar
2024

A mmWave Radar SLAM Method in Subterranean Tunnel for Low Visibility and Degradation

RA-L 2024

A novel mmWave radar SLAM method is proposed to integrate multi-dimensional information, including velocity, spatial, RCS, and semantics to enable autonomous navigation in subterranean tunnel environments, which are characterized by low visibility and degraded conditions. By combining doppler odomet

Cited by 8SourceScholar
2024

Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes

NeurIPS 2024poster

In this paper, we show that applying adaptive methods directly to distributed minimax problems can result in non-convergence due to inconsistency in locally computed adaptive stepsizes. To address this challenge, we propose D-AdaST, a Distributed Adaptive minimax method with Stepsize Tracking. The k…

Cited by 0SourcePDFScholar
2024

Analyzing Large Language Models’ Capability in Location Prediction

COLING 2024main

In this paper, we investigate and evaluate large language models’ capability in location prediction. We present experimental results with four models—FLAN-T5, FLAN-UL2, FLAN-Alpaca, and ChatGPT—in various instruction finetuning and exemplar settings. We analyze whether taking into account the contex…

Cited by 21SourcePDFScholar
2024

Attribute-Guided Pedestrian Retrieval: Bridging Person Re-ID with Internal Attribute Variability

CVPR 2024poster

In various domains such as surveillance and smart retail pedestrian retrieval centering on person re-identification (Re-ID) plays a pivotal role. Existing Re-ID methodologies often overlook subtle internal attribute variations which are crucial for accurately identifying individuals with changing ap…

Cited by 9SourcePDFScholar
2024

Dual-modal Tactile E-skin: Enabling Bidirectional Human-Robot Interaction via Integrated Tactile Perception and Feedback

ICRA 2024poster

To foster an immersive and natural human-robot interaction (HRI), the implementation of tactile perception and feedback becomes imperative, effectively bridging the conventional sensory gap. In this paper, we propose a dual-modal electronic skin (e-skin) that integrates magnetic tactile sensing and…

Cited by 2SourceScholar
2024

Efficient Multimodal Semantic Segmentation via Dual-Prompt Learning

IROS 2024

Multimodal (e.g., RGB-Depth/RGB-Thermal) fusion has shown great potential for improving semantic segmentation in complex scenes (e.g., indoor/low-light conditions). Existing approaches often fully fine-tune a dual-branch encoder-decoder framework with a complicated feature fusion strategy for achiev

Cited by 45SourcecodeScholar
2024

Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor

NeurIPS 2024poster

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serio…

Cited by 0SourcePDFScholar
2024

Free Lunch for Gait Recognition: A Novel Relation Descriptor

ECCV 2024poster

"Gait recognition is to seek correct matches for query individuals by their unique walking patterns. However, current methods focus solely on extracting individual-specific features, overlooking “interpersonal” relationships. In this paper, we propose a novel Relation Descriptor that captures not on…

Cited by 3SourcePDFScholar
2024

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

CVPR 2024highlight

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performance gains contributed by large vision and language pre-training we find that state-of…

2024

PrivSGP-VR: Differentially Private Variance-Reduced Stochastic Gradient Push with Tight Utility Bounds

IJCAI 2024poster

In this paper, we propose a differentially private decentralized learning method (termed PrivSGP-VR) which employs stochastic gradient push with variance reduction and guarantees (epsilon, delta)-differential privacy (DP) for each node. Our theoretical analysis shows that, under DP Gaussian noise wi…

Cited by 0SourcePDFScholar
2024

Selective and Orthogonal Feature Activation for Pedestrian Attribute Recognition

AAAI 2024technical

Pedestrian Attribute Recognition (PAR) involves identifying the attributes of individuals in person images. Existing PAR methods typically rely on CNNs as the backbone network to extract pedestrian features. However, CNNs process only one adjacent region at a time, leading to the loss of long-range…

Cited by 5SourcePDFScholar
2024

TDeLTA: A Light-Weight and Robust Table Detection Method Based on Learning Text Arrangement

AAAI 2024technical

The diversity of tables makes table detection a great challenge, leading to existing models becoming more tedious and complex. Despite achieving high performance, they often overfit to the table style in training set, and suffer from significant performance degradation when encountering out-of-distr…

2023

Bag of Tricks for Training Data Extraction from Language Models

ICML 2023poster

With the advance of language models, privacy protection is receiving more attention. Training data extraction is therefore of great importance, as it can serve as a potential tool to assess privacy leakage. However, due to the difficulty of this task, most of the existing methods are proof-of-concep…

2023

Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation

NeurIPS 2023poster

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through complex environments based on natural language instructions. In contrast to conventional approaches, which primarily focus on the spatial domain exploration, we propose a paradigm shift toward the F…

2023

Gender-tuning: Empowering Fine-tuning for Debiasing Pre-trained Language Models

ACL 2023findings

Recent studies have revealed that the widely-used Pre-trained Language Models (PLMs) propagate societal biases from the large unmoderated pre-training corpora. Existing solutions require debiasing training processes and datasets for debiasing, which are resource-intensive and costly. Furthermore, th…

2023

PlanarTrack: A Large-scale Challenging Benchmark for Planar Object Tracking

ICCV 2023poster

Planar object tracking is a critical computer vision problem and has drawn increasing interest owing to its key roles in robotics, augmented reality, etc. Despite rapid progress, its further development, especially in the deep learning era, is largely hindered due to the lack of large-scale challeng…

Cited by 5PDFScholar
2022

Generalizable Person Re-identification via Self-Supervised Batch Norm Test-Time Adaption

AAAI 2022technical

In this paper, we investigate the generalization problem of person re-identification (re-id), whose major challenge is the distribution shift on an unseen domain. As an important tool of regularizing the distribution, batch normalization (BN) has been widely used in existing methods. However, they n…

Cited by 26SourcePDFScholar
2022

MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text Matching

NeurIPS 2022accept

Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which are trained on millions or billions of paired images and texts. Different from them, this paper studies a new scenario as unpaired image-text matching, in which paired images and tex…

Cited by 26SourcePDFScholar
2022

Regularized Graph Structure Learning with Semantic Knowledge for Multi-variates Time-Series Forecasting

IJCAI 2022poster

Multivariate time-series forecasting is a critical task for many applications, and graph time-series network is widely studied due to its capability to capture the spatial-temporal correlation simultaneously. However, most existing works focus more on learning with the explicit prior graph structure…

2022

Tackling Data Heterogeneity: A New Unified Framework for Decentralized SGD with Sample-induced Topology

ICML 2022spotlight

We develop a general framework unifying several gradient-based stochastic optimization methods for empirical risk minimization problems both in centralized and distributed scenarios. The framework hinges on the introduction of an augmented graph consisting of nodes modeling the samples and edges mod…

Cited by 19SourcePDFScholar
2021

Clothing Status Awareness for Long-Term Person Re-Identification

ICCV 2021poster

Long-Term person re-identification (LT-reID) exposes extreme challenges because of the longer time gaps between two recording footages where a person is likely to change clothing. There are two types of approaches for LT-reID: biometrics-based approach and data adaptation based approach. The former…

Cited by 130PDFScholar
2021

FWB-Net: Front White Balance Network for Color Shift Correction in Single Image Dehazing Via Atmospheric Light Estimation

ICASSP 2021accepted

In recent years, single image dehazing deep models based on Atmospheric Scattering Model (ASM) have achieved remarkable results. But the dehazing outputs of those models suffer from color shift. Analyzing the ASM model shows that the atmospheric light factor (ALF) is set as a scalar which indicates…

Cited by 0SourceScholar
2021

Hierarchical Bit-Wise Differential Coding (HBDC) of Point Cloud Attributes

ICASSP 2021accepted

Targeting both computing and coding efficiencies, we propose in this work a novel hierarchical bit-wise differential coding scheme to compress point cloud attributes. The encoder firstly quantizes and organizes the points into an octree structure and, for each internal node, picks its attribute(s) f…

Cited by 0SourceScholar
2021

Knowledge-aware Leap-LSTM: Integrating Prior Knowledge into Leap-LSTM towards Faster Long Text Classification

AAAI 2021technical

While widely used in industry, recurrent neural networks (RNNs) are known to have deficiencies in dealing with long sequences (e.g. slow inference, vanishing gradients etc.). Recent research has attempted to accelerate RNN models by developing mechanisms to skip irrelevant words in input. Due to th…

2021

Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision

NeurIPS 2021poster

In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper,…

2021

Rethinking the Heatmap Regression for Bottom-Up Human Pose Estimation

CVPR 2021poster

Heatmap regression has become the most prevalent choice for nowadays human pose estimation methods. The ground-truth heatmaps are usually constructed by covering all skeletal keypoints by 2D gaussian kernels. The standard deviations of these kernels are fixed. However, for bottom-up methods, which n…

Cited by 219PDFcodeScholar
2021

Riemannian Geometric Optimization Methods for Joint Design of Transmit Sequence and Receive Filter of MIMO Radar

ICASSP 2021accepted

To maximize the signal-to-interference-plus-noise ratio (SINR) under a constant-envelope constraint, an efficient joint design of the transmit waveform and the receive filter for multipleinput multiple-output (MIMO) radars is essential. In this paper, we propose a novel optimization framework to sol…

Cited by 0SourceScholar
2020

Acoustic Model Adaptation for Presentation Transcription and Intelligent Meeting Assistant Systems

ICASSP 2020accepted

We present our solution for unsupervised rapid speaker adaptation in a state-of-art presentation and intelligent meeting transcription system. We adopt the Kullback-Leibler (KL) divergence regularized model adaptation paradigm. For the adaptation architecture, we found that the linear projection lay…

Cited by 0SourceScholar
2020

L-Vector: Neural Label Embedding for Domain Adaptation

ICASSP 2020accepted

We propose a novel neural label embedding (NLE) scheme for the domain adaptation of a deep neural network (DNN) acoustic model with unpaired data samples from source and target domains. With NLE method, we distill the knowledge from a powerful source-domain DNN into a dictionary of label embeddings,…

Cited by 0SourceScholar
2020

Pointing to Select: A Fast Pointer-LSTM for Long Text Classification

COLING 2020main

Recurrent neural networks (RNNs) suffer from well-known limitations and complications which include slow inference and vanishing gradients when processing long sequences in text classification. Recent studies have attempted to accelerate RNNs via various ad hoc mechanisms to skip irrelevant words in…

2020

Prediction and Recovery for Adaptive Low-Resolution Person Re-Identification

ECCV 2020poster

Low-resolution person re-identification (LR re-id) is a challenging task with low-resolution probes and high-resolution gallery images. To address the resolution mismatch, existing methods typically recover missing details for low-resolution probes by super-resolution. However, they usually pre-spec…

Cited by 32SourcePDFScholar
2020

Towards Part-aware Monocular 3D Human Pose Estimation: An Architecture Search Approach

ECCV 2020poster

Even though most existing monocular 3D pose estimation approaches achieve very competitive results, they ignore the heterogeneity among human body parts by estimating them with the same network architecture. To accurately estimate 3D poses of different body parts, we attempt to build a part-aware 3D…

Cited by 32SourcePDFScholar
2020

Unfolding the Alternating Optimization for Blind Super Resolution

NeurIPS 2020poster

Previous methods decompose blind super resolution (SR) problem into two sequential steps: \textit{i}) estimating blur kernel from given low-resolution (LR) image and \textit{ii}) restoring SR image based on estimated kernel. This two-step solution involves two independently trained models, which may…

2020

Using Personalized Speech Synthesis and Neural Language Generator for Rapid Speaker Adaptation

ICASSP 2020accepted

We propose to use the personalized speech synthesis and the neural language generator to synthesize content relevant personalized speech for rapid speaker adaptation. It has two distinct aspects: First, it relieves the general data sparsity issue in rapid adaptation via making use of additional synt…

Cited by 32SourceScholar
2019

Box-Driven Class-Wise Region Masking and Filling Rate Guided Loss for Weakly Supervised Semantic Segmentation

CVPR 2019poster

Semantic segmentation has achieved huge progress via adopting deep Fully Convolutional Networks (FCN). However, the performance of FCN based models severely rely on the amounts of pixel-level annotations which are expensive and time-consuming. To address this problem, it is a good choice to learn to…

Cited by 283PDFcodeScholar
2019

Language-Driven Temporal Activity Localization: A Semantic Matching Reinforcement Learning Model

CVPR 2019oral

Current studies on action detection in untrimmed videos are mostly designed for action classes, where an action is described at word level such as jumping, tumbling, swing, etc. This paper focuses on a rarely investigated problem of localizing an activity via a sentence query which would be more cha…

Cited by 215PDFScholar
2019

Local Relationship Learning With Person-Specific Shape Regularization for Facial Action Unit Detection

CVPR 2019poster

Encoding individual facial expressions via action units (AUs) coded by the Facial Action Coding System (FACS) has been found to be an effective approach in resolving the ambiguity issue among different expressions. While a number of methods have been proposed for AU detection, robust AU detection in…

Cited by 171PDFScholar
2019

SBSGAN: Suppression of Inter-Domain Background Shift for Person Re-Identification

ICCV 2019poster

Cross-domain person re-identification (re-ID) is challenging due to the bias between training and testing domains. We observe that if backgrounds in the training and testing datasets are very different, it dramatically introduces difficulties to extract robust pedestrian features, and thus compromis…

Cited by 139PDFcodeScholar
2018

Aligning Infinite-Dimensional Covariance Matrices in Reproducing Kernel Hilbert Spaces for Domain Adaptation

CVPR 2018poster

Domain shift, which occurs when there is a mismatch between the distributions of training (source) and testing (target) datasets, usually results in poor performance of the trained model on the target domain. Existing algorithms typically solve this issue by reducing the distribution discrepancy in…

Cited by 61SourcePDFScholar
2018

Cross-Modal Ranking with Soft Consistency and Noisy Labels for Robust RGB-T Tracking

ECCV 2018poster

Due to the complementary benefits of visible (RGB) and thermal infrared (T) data, RGB-T object tracking attracts more and more attention recently for boosting the performance under adverse illumination conditions. Existing RGB-T tracking methods usually localize a target object with a bounding box,…

Cited by 163SourcePDFScholar
2018

Learning Semantic Concepts and Order for Image and Sentence Matching

CVPR 2018poster

Image and sentence matching has made great progress recently, but it remains challenging due to the large visual semantic discrepancy. This mainly arises from that the representation of pixel-level image usually lacks of high-level semantic information as in its matched sentence. In this work, we pr…

Cited by 409SourcePDFScholar
2018

Mask-Guided Contrastive Attention Model for Person Re-Identification

CVPR 2018poster

Person Re-identification (ReID) is an important yet challenging task in computer vision. Due to the diverse background clutters, variations on viewpoints and body poses, it is far from solved. How to extract discriminative and robust features invariant to background clutters is the core problem. In…

2018

RetGK: Graph Kernels based on Return Probabilities of Random Walks

NeurIPS 2018poster

Graph-structured data arise in wide applications, such as computer vision, bioinformatics, and social networks. Quantifying similarities among graphs is a fundamental problem. In this paper, we develop a framework for computing graph kernels, based on return probabilities of random walks. The advant…

Cited by 124SourcePDFScholar
2017

Improved cepstra minimum-mean-square-error noise reduction algorithm for robust speech recognition

ICASSP 2017accepted

In the era of deep learning, although beam-forming multi-channel signal processing is still very helpful, it was reported that single-channel robust front-ends usually cannot benefit deep learning models because the layer-by-layer structure of deep learning models provides a feature extraction strat…

Cited by 0SourceScholar
2017

See the Forest for the Trees: Joint Spatial and Temporal Recurrent Neural Networks for Video-Based Person Re-Identification

CVPR 2017poster

Surveillance cameras have been widely used in different scenes. Accordingly, a demanding need is to recognize a person under different cameras, which is called person re-identification. This topic has gained increasing interests in computer vision recently. However, less attention has been paid to v…

Cited by 395PDFScholar
2016

Cat-inspired mechanical design of self-adaptive toes for a legged robot

IROS 2016poster

Cats have protractible claws to fold their tips to keep them sharp. They protract claws while hunting and pawing on slippery surfaces. Protracted claws by tendons and muscles of toes can help cats anchoring themselves steady while their locomotion trends to slip and releasing the hold while they ret…

Cited by 10SourceScholar
2015

Bidirectional Recurrent Convolutional Networks for Multi-Frame Super-Resolution

NeurIPS 2015poster

Super resolving a low-resolution video is usually handled by either single-image super-resolution (SR) or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video super-res…

Cited by 328SourcePDFScholar
2015

Conditional High-Order Boltzmann Machine: A Supervised Learning Model for Relation Learning

ICCV 2015poster

Relation learning is a fundamental operation in many computer vision tasks. Recently, high-order Boltzmann machine and its variants have exhibited the great power of modelling various data relation. However, most of them are unsupervised learning models which are not very discriminative and thus can…

Cited by 9PDFScholar