← Search

Hong Liu

130 accepted papers

2026

An Enhanced Soft Growing Robot with Mixed-Layer Jamming for Superior Load Capacity and Improved Mobility

ICRA 2026poster

Soft robots have gained widespread attention due to their lightweight nature and inherent safety. Among them, soft growing robots (SGRs) are inspired by the growth mechanism of vines, achieving movement through tip eversion. However, their load-bearing capacity remains a significant challenge due to…

Cited by 0SourceScholar
2026

COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning

CVPR 2026

Recent advances in text-to-image (T2I) generation can produce highly resembling images of copyrighted characters, often indistinguishable from official depictions, raising serious concerns about intellectual property infringement. Consequently, robust detection of copyright character infringement is

Cited by 0SourceScholar
2026

Debiased Multiplex Tokenizer for Efficient Map-Free Visual Relocalization

AAAI 2026technical

Image-based feature representation plays a critical role in visual localization, enabling robots to estimate their position and orientation in GPS-denied environments. However, this task is often undermined by significant variations in camera viewpoints and scene appearances. Recently, map-free visu

Cited by 0SourcePDFScholar
2026

ERTACache: Error Rectification and Timesteps Adjustment for Efficient Diffusion

ICLR 2026poster

Diffusion models suffer from substantial computational overhead due to their inherently iterative inference process. While feature caching offers a promising acceleration strategy by reusing intermediate outputs across timesteps, naive reuse often incurs noticeable quality degradation. In this work…

Cited by 0SourceScholar
2026

Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization

ICLR 2026poster

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across different degradation types. Existing approaches either sacrifice efficiency for versatility or fail to capture the distin…

Cited by 0SourcecodeScholar
2026

Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments

ICRA 2026poster

Motion generation in dynamic environments is crucial for human-machine interaction with redundant manipulators. In this context, we propose the Enhancing Safety and Manipulability (ESM) scheme, which integrates geometry-based dynamic obstacle avoidance, manipulability optimization,trajectory trackin…

Cited by 0SourceScholar
2026

Implicit LiDAR SLAM with Confidence-Guided SDF and Normal-Driven Sampling

ICRA 2026poster

Implicit representations for LiDAR-based Simultaneous Localization and Mapping (SLAM) offer significant advantages in storage efficiency and expressive power over traditional explicit maps. However, a critical limitation for implicit SLAM is their deterministic nature, which prevents the quantificat…

Cited by 0Scholar
2026

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

RSS 2026poster

Achieving safe manipulation-oriented navigation for humanoid robots is fundamentally challenged by two factors: locomotion-induced perceptual distortion (causing semantic-geometry distortion) and changes within the environment (causing map-reality mismatches). Existing static scene graphs often fail…

Cited by 0SourceScholar
2026

Listening Between the Frames: Bridging Temporal Gaps in Large Audio-Language Models

AAAI 2026technical

Recent Large Audio-Language Models (LALMs) exhibit impressive capabilities in understanding audio content for conversational QA tasks. However, these models struggle to accurately understand timestamps for temporal localization (e.g., Temporal Audio Grounding) and are restricted to short audio perce

Cited by 0SourcePDFScholar
2026

Masked Clustering Prediction for Unsupervised Point Cloud Pre-training

AAAI 2026technical

Vision transformers (ViTs) have recently been widely applied to 3D point cloud understanding, with masked autoencoding as the predominant pre-training paradigm. However, the challenge of learning dense and informative semantic features from point clouds via standard ViTs remains underexplored. We pr

Cited by 0SourcePDFScholar
2026

QuMAB: Query-based Multi-annotator Behavior Pattern Learning

AAAI 2026technical

Multi-annotator learning traditionally aggregates diverse annotations to approximate a single “ground truth”, treating disagreements as noise. However, this paradigm faces fundamental challenges: subjective tasks often lack absolute ground truth, and sparse annotation coverage makes aggregation stat

Cited by 0SourcePDFScholar
2026

SimLabel: Similarity-Weighted Semi-supervision for Multi-annotator Learning with Missing Labels

AAAI 2026technical

Multi-annotator learning (MAL) aims to model annotator-specific labeling patterns. However, existing methods face a critical challenge: they simply skip updating annotator-specific model parameters when encountering missing labels—a common scenario in real-world crowdsourced datasets where each anno

Cited by 0SourcePDFScholar
2026

Synthetic Bootstrapped Pretraining

ICLR 2026poster

We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dataset and then leverages it to synthesize a vast new corpus for joint training. While the standard pretraining teaches LM…

Cited by 0SourceScholar
2025

A Hybrid Motion Optimization Framework for the Humanoid Upper-Body Robot: Safe and Dexterous Object Carrying

RA-L 2025

In this letter, a novel hybrid motion optimization framework is proposed for a humanoid upper-body robot with two 7-DOF arms and a 2-DOF waist, to flexibly and safely carry objects in constrained and dynamic environments. The framework consists of <bold xmlns:mml="http://www.w3.org/1998/Math/MathML"

Cited by 1SourceScholar
2025

A Unified End-to-End Network for Category-Level and Instance-Level Object Pose Estimation from RGB Images

ICRA 2025

Accurately estimating the 6-DoF pose of objects is a fundamental challenge in computer vision and robotics. While category-level pose estimation based on RGBD data has achieved good performance in recent years, estimating poses solely from RGB images remains a significant challenge. Existing RGB-bas

Cited by 0SourcecodeScholar
2025

ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language Models

AAAI 2025technical

Large Language Models (LLMs) have revolutionized natural language processing tasks. However, their practical application is constrained by substantial memory and computational demands. Post-training quantization (PTQ) is considered an effective method to accelerate LLM inference. Despite its growing…

2025

An Enhanced Soft Growing Robot with Mixed-Layer Jamming for Superior Load Capacity and Improved Mobility

RA-L 2025

Soft robots have gained widespread attention due to their lightweight nature and inherent safety. Among them, soft growing robots (SGRs) are inspired by the growth mechanism of vines, achieving movement through tip eversion. However, their load-bearing capacity remains a significant challenge due to

Cited by 2SourceScholar
2025

CPRM: A LLM-based Continual Pre-training Framework for Relevance Modeling in Commercial Search

NAACL 2025industry

Relevance modeling between queries and items stands as a pivotal component in commercial search engines, directly affecting the user experience. Given the remarkable achievements of large language models (LLMs) in various natural language processing (NLP) tasks, LLM-based relevance modeling is gradu…

2025

D-VST: Diffusion Transformer for Pathology-Correct Tone-Controllable Cross-Dye Virtual Staining of Whole Slide Images

NeurIPS 2025poster

Diffusion-based virtual staining methods of histopathology images have demonstrated outstanding potential for stain normalization and cross-dye staining (e.g., hematoxylin-eosin to immunohistochemistry). However, achieving pathology-correct cross-dye virtual staining with versatile tone controls pos…

Cited by 0SourceScholar
2025

DexMGNet: Multi-Mode Dexterous Grasping in Cluttered Scenes With Generative Models

RA-L 2025

Dexterous grasping is a crucial technique in humanoid robot manipulation. However, existing methods still fall short in effectively detecting dexterous grasps in cluttered environments. In this work, we propose DexMGNet, a novel multi-mode dexterous grasping framework designed for such challenging s

Cited by 1SourceScholar
2025

Enhancing Safety and Manipulability of Redundant Manipulators: Accelerated Motion Generation in Dynamic Environments

RA-L 2025

Motion generation in dynamic environments is crucial for human-machine interaction with redundant manipulators. In this context, we propose the Enhancing Safety and Manipulability (ESM) scheme, which integrates geometry-based dynamic obstacle avoidance, manipulability optimization, trajectory tracki

Cited by 1SourceScholar
2025

FusionClassNet: A Multi-Scale Feature Fusion Network with Contrastive Loss-Driven Classification for Enhanced Lung Tumor Image Representations

ICASSP 2025accepted

Lung cancer remains a leading cause of mortality, making accurate subtype identification crucial for effective lung cancer diagnosis and treatment. Recent advancements in medical image classification, particularly through Convolutional Neural Networks (CNNs) and Transformers, have significantly impr…

Cited by 0SourceScholar
2025

GraphGPT: Generative Pre-trained Graph Eulerian Transformer

ICML 2025poster

We introduce *GraphGPT*, a novel self-supervised *generative pre-trained* model for graph learning based on the *Graph Eulerian Transformer* (**GET**). First, we propose **GET**, which combines a standard transformer encoder or decoder architecture with an innovative graph-to-sequence transformation…

2025

Hierarchical Trajectory Planning Method for Piano-Playing Robot

IROS 2025

Piano-playing tasks, which effectively demonstrate bimanual coordination capabilities in humanoid robots, are increasingly becoming a research focus. However, prior research has predominantly focused on Cartesian space trajectory planning without adequately addressing real-world obstacle avoidance c

Cited by 0SourceScholar
2025

Learning Perceptive Humanoid Locomotion over Challenging Terrain

IROS 2025

Humanoid robots are engineered to navigate terrains akin to those encountered by humans, which necessitates human-like locomotion and perceptual abilities. Currently, the most reliable controllers for humanoid motion rely exclusively on proprioception, a reliance that becomes both dangerous and unre

Cited by 23SourceScholar
2025

Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification

ICASSP 2025accepted

Cognitive impairment detection through spontaneous speech is a promising avenue for early diagnosis of Alzheimer’s disease (AD) and mild cognitive impairment (MCI), where timely intervention can significantly improve patient outcomes. The PROCESS Grand Challenge at ICASSP 2025 addresses these tasks…

Cited by 0SourceScholar
2025

LiD-FL: Towards List-Decodable Federated Learning

AAAI 2025technical

Federated learning is often used in environments with many unverified participants. Therefore, federated learning under adversarial attacks receives significant attention. This paper proposes an algorithmic framework for list-decodable federated learning, where a central server maintains a list of m…

2025

Multi-Sector Overlap Loss: A Universal Framework for One-Shot 6DoF Global Localization Across Heterogeneous LiDARs

RA-L 2025

This paper presents a universal LiDAR point cloud global localization framework based on multi-sector overlapping loss to address the localization challenges caused by heterogeneous LiDAR point clouds with varying resolutions, scanning formats, and field of view differences. The proposed method firs

Cited by 0SourceScholar
2025

One-shot Global Localization through Semantic Distribution Feature Retrieval and Semantic Topological Histogram Registration

IROS 2025

One-shot global localization is crucial in many robotic applications, providing significant advantages during initialization and relocalization processes. However, LiDAR-based one-shot global localization methods encounter challenges, including local feature matching errors, sensitivity to dynamic o

Cited by 0SourcecodeScholar
2025

PDDM: Pseudo Depth Diffusion Model for RGB-PD Semantic Segmentation Based in Complex Indoor Scenes

AAAI 2025technical

The integration of RGB and depth modalities significantly enhances the accuracy of segmenting complex indoor scenes, with depth data from RGB-D cameras playing a crucial role in this improvement. However, collecting an RGB-D dataset is more expensive than an RGB dataset due to the need for specializ…

2025

Recognizing Actions from Robotic View for Natural Human-Robot Interaction

ICCV 2025poster

Natural Human-Robot Interaction (N-HRI) requires robots to recognize human actions at varying distances and states, regardless of whether the robot itself is in motion or stationary. This setup is more flexible and practical than conventional human action recognition tasks. However, existing benchma…

2025

SGTD: A Semantic-Guided Triangle Descriptor for One-Shot LiDAR-Based Global Localization

RA-L 2025

This paper presents a novel one-shot global localization algorithm based on semantic-guided triangle descriptors to address initialization and global localization challenges in GNSSdenied environments. By encoding semantic geometric information into triangle descriptors, the proposed approach achiev

Cited by 1SourcecodeScholar
2025

SVTformer: Spatial-View-Temporal Transformer for Multi-View 3D Human Pose Estimation

AAAI 2025technical

Recently, transformer-based methods have been introduced to estimate 3D human pose from multiple views by aggregating the spatial-temporal information of human joints to achieve the lifting of 2D to 3D. However, previous approaches cannot model the inter-frame correspondence of each view's joint ind…

2025

TCNet: A Temporally Consistent Network for Self-supervised Monocular Depth Estimation

IROS 2025

Despite significant advances in self-supervised monocular depth estimation methods, achieving temporally consistent and accurate depth maps from frame sequences remains a formidable challenge. Existing approaches often estimate depth maps for individual frames in isolation, neglecting the rich geome

Cited by 0SourceScholar
2025

TCPFormer: Learning Temporal Correlation with Implicit Pose Proxy for 3D Human Pose Estimation

AAAI 2025technical

Recent multi-frame lifting methods have dominated the 3D human pose estimation. However, previous methods ignore the intricate dependence within the 2D pose sequence and learn single temporal correlation. To alleviate this limitation, we propose TCPFormer, which leverages an implicit pose proxy as a…

2025

USD: Unsupervised Soft Contrastive Learning for Fault Detection in Multivariate Time Series

ICASSP 2025accepted

Unsupervised fault detection in multivariate time series is critical for maintaining the integrity and efficiency of complex systems, with current methodologies largely focusing on statistical and machine learning techniques. However, these approaches often rest on the assumption that data distribut…

Cited by 0SourceScholar
2024

AttA-NET: Attention Aggregation Network for Audio-Visual Emotion Recognition

ICASSP 2024accepted

In video-based emotion recognition, effective multi-modal fusion techniques are essential to leverage the complementary relationship between audio and visual modalities. Recent attention-based fusion methods are widely leveraged for capturing modal-shared properties. However, they often ignore the m…

Cited by 0SourceScholar
2024

Audio Generation with Multiple Conditional Diffusion Model

AAAI 2024technical

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the controllability of existing pre-trained text-to-audio models…

2024

Chain of Thought Empowers Transformers to Solve Inherently Serial Problems

ICLR 2024poster

Generating a sequence of intermediate steps, \emph{a.k.a.}, a chain of thought (CoT), is a highly effective method to improve the accuracy of large language models (LLMs) on arithmetics and symbolic reasoning tasks. However, the mechanism behind CoT remains unclear. This work provides a theoretical…

Cited by 109SourcePDFScholar
2024

DiffusionFake: Enhancing Generalization in Deepfake Detection via Guided Stable Diffusion

NeurIPS 2024poster

The rapid progress of Deepfake technology has made face swapping highly realistic, raising concerns about the malicious use of fabricated facial content. Existing methods often struggle to generalize to unseen domains due to the diverse nature of facial manipulations. In this paper, we revisit the g…

2024

Dual-Branch Graph Transformer Network for 3D Human Mesh Reconstruction from Video

IROS 2024poster

Human Mesh Reconstruction (HMR) from monocular video plays an important role in human-robot interaction and collaboration. However, existing video-based human mesh reconstruction methods face a trade-off between accurate reconstruction and smooth motion. These methods design networks based on either…

Cited by 0SourcecodeScholar
2024

Federated Modality-Specific Encoders and Multimodal Anchors for Personalized Brain Tumor Segmentation

AAAI 2024technical

Most existing federated learning (FL) methods for medical image analysis only considered intramodal heterogeneity, limiting their applicability to multimodal imaging applications. In practice, it is not uncommon that some FL participants only possess a subset of the complete imaging modalities, posi…

2024

Hourglass Tokenizer for Efficient Transformer-Based 3D Human Pose Estimation

CVPR 2024highlight

Transformers have been successfully applied in the field of video-based 3D human pose estimation. However the high computational costs of these video pose transformers (VPTs) make them impractical on resource-constrained devices. In this paper we present a plug-and-play pruning-and-recovering framew…

2024

Mask-Homo: Pseudo Plane Mask-Guided Unsupervised Multi-Homography Estimation

AAAI 2024technical

Homography estimation is a fundamental problem in computer vision. Previous works mainly focus on estimating either a single homography, or multiple homographies based on mesh grid division of the image. In practical scenarios, single homography is inadequate and often leads to a compromised result…

2024

Mitigating robust overfitting via self-residual-calibration regularization (Abstract Reprint)

IJCAI 2024poster

Overfitting in adversarial training has attracted the interest of researchers in the community of artificial intelligence and machine learning in recent years. To address this issue, in this paper we begin by evaluating the defense performances of several calibration methods on various robust models…

Cited by 0SourcePDFScholar
2024

Model Design and Concept of Operations of Standard Interface for On-orbit Construction

ICRA 2024poster

The construction of large-scale space facilities requires the use of on-orbit construction technology. However, several of its key components, such as standard interface design, compliant control methods, and path planning for multi-branch robots, still need improvement before practical application.…

Cited by 1SourceScholar
2024

NEDS-SLAM: A Neural Explicit Dense Semantic SLAM Framework Using 3D Gaussian Splatting

RA-L 2024

We propose NEDS-SLAM, a dense semantic SLAM system based on 3D Gaussian representation, that enables robust 3D semantic mapping, accurate camera tracking, and high-quality rendering in real-time. In the system, we propose a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www

Cited by 38SourceScholar
2024

Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers

AAAI 2024technical

Byzantine machine learning has garnered considerable attention in light of the unpredictable faults that can occur in large-scale distributed learning systems. The key to secure resilience against Byzantine machines in distributed learning is resilient aggregation mechanisms. Although abundant resil…

2024

Position: Towards Implicit Prompt For Text-To-Image Models

ICML 2024poster

Recent text-to-image (T2I) models have had great success, and many benchmarks have been proposed to evaluate their performance and safety. However, they only consider explicit prompts while neglecting implicit prompts (hint at a target without explicitly mentioning it). These prompts may get rid of…

Cited by 4SourcePDFScholar
2024

Semi-Supervised Sound Event Detection with Local and Global Consistency Regularization

ICASSP 2024accepted

Learning meaningful frame-wise features on a partially labeled dataset is crucial to semi-supervised sound event detection. Prior works either maintain consistency on frame-level predictions or seek feature-level similarity among neighboring frames, which cannot exploit the potential of unlabeled da…

Cited by 0SourceScholar
2024

Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

ICLR 2024poster

Given the massive cost of language model pre-training, a non-trivial improvement of the optimization algorithm would lead to a material reduction on the time and cost of training. Adam and its variants have been state-of-the-art for years, and more sophisticated second-order (Hessian-based) optimize…

Cited by 151SourcePDFScholar
2024

Transformer-Enhanced Motion Planner: Attention-Guided Sampling for State-Specific Decision Making

RA-L 2024

Sampling-based motion planning (SBMP) algorithms are renowned for their robust global search capabilities. However, the inherent randomness in their sampling mechanisms often results in inconsistent path quality and limited search efficiency. In response to these challenges, this work proposes a nov

Cited by 5SourceScholar
2023

A Multi-Configuration Track-Legged Humanoid Robot for Dexterous Manipulation and High Mobility: Design and Development

RA-L 2023

Various applications (e.g., disaster response) put forward higher requirements for the capabilities of robots. However, to date, few platforms can be compatible with the required mobility, stability and dexterity. A major challenge is mobile stability may severely limit the variety of manipulation s

Cited by 7SourceScholar
2023

An Efficient Trajectory Planner for Car-Like Robots on Uneven Terrain

IROS 2023poster

Autonomous navigation of ground robots on uneven terrain is being considered in more and more tasks. However, uneven terrain will bring two problems to motion planning: how to assess the traversability of the terrain and how to cope with the dynamics model of the robot associated with the terrain. T…

Cited by 17SourcecodeScholar
2023

Body Prior Guided Graph Convolutional Neural Network for Skeleton-Based Action Recognition

ICASSP 2023accepted

Graph Convolutional Network (GCN) has achieved high success in the skeleton-based human action recognition task by modeling the human skeleton as a graph. However, it remains a problem for GCN-based methods to learn distinctive action features from a limited number of training samples. Via taking fu…

Cited by 0SourceScholar
2023

Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial Training

ICASSP 2023accepted

Person re-identification (ReID) aims at retrieving a person of interest across multiple cameras. Despite significant progress in person ReID, viewpoint variation remains an obstacle to extracting discriminative features for retrieval. To address this problem, we propose a Viewpoint-Robust Network (V…

Cited by 0SourceScholar
2023

Co-Evolution of Pose and Mesh for 3D Human Body Estimation from Video

ICCV 2023poster

Despite significant progress in single image-based 3D human mesh recovery, accurately and smoothly recovering 3D human motion from a video remains challenging. Existing video-based methods generally recover human mesh by estimating the complex pose and shape parameters from coupled image features, w…

Cited by 22PDFcodeScholar
2023

FSAR: Federated Skeleton-based Action Recognition with Adaptive Topology Structure and Knowledge Distillation

ICCV 2023poster

Existing skeleton-based action recognition methods typically follow a centralized learning paradigm, which can pose privacy concerns when exposing human-related videos. Federated Learning (FL) has attracted much attention due to its outstanding advantages in privacy-preserving. However, directly app…

Cited by 8PDFScholar
2023

Gator: Graph-Aware Transformer with Motion-Disentangled Regression for Human Mesh Recovery from a 2D Pose

ICASSP 2023accepted

3D human mesh recovery from a 2D pose plays an important role in various applications. However, it is hard for existing methods to simultaneously capture the multiple relations during the evolution from skeleton to mesh, including joint-joint, joint-vertex and vertex-vertex relations, which often le…

Cited by 0SourceScholar
2023

HTNet: Human Topology aware network for 3d Human pose estimation

ICASSP 2023accepted

3D human pose estimation errors would propagate along the human body topology and accumulate at the end joints of limbs. Inspired by the backtracking mechanism in automatic control systems, we design an Intra-Part Constraint module that utilizes the parent nodes as the reference to build topological…

Cited by 0SourceScholar
2023

Improving Adversarial Robustness via Information Bottleneck Distillation

NeurIPS 2023poster

Previous studies have shown that optimizing the information bottleneck can significantly improve the robustness of deep neural networks. Our study closely examines the information bottleneck principle and proposes an Information Bottleneck Distillation approach. This specially designed, robust disti…

2023

Inferential Knowledge-Enhanced Integrated Reasoning for Video Question Answering

AAAI 2023technical

Recently, video question answering has attracted growing attention. It involves answering a question based on a fine-grained understanding of video multi-modal information. Most existing methods have successfully explored the deep understanding of visual modality. We argue that a deep understanding…

Cited by 1SourcePDFScholar
2023

Interweaved Graph and Attention Network for 3D Human Pose Estimation

ICASSP 2023accepted

Despite substantial progress in 3D human pose estimation from a single-view image, prior works rarely explore global and local correlations, leading to insufficient learning of human skeleton representations. To address this issue, we propose a novel Interweaved Graph and Attention Network (IGANet)…

Cited by 0SourceScholar
2023

Learning Concordant Attention via Target-aware Alignment for Visible-Infrared Person Re-identification

ICCV 2023poster

Owing to the large distribution gap between the heterogeneous data in Visible-Infrared Person Re-identification (VI Re-ID), we point out that existing paradigms often suffer from the inter-modal semantic misalignment issue and thus fail to align and compare local details properly. In this paper, we…

Cited by 37PDFScholar
2023

M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing Modalities

AAAI 2023technical

Multimodal magnetic resonance imaging (MRI) provides complementary information for sub-region analysis of brain tumors. Plenty of methods have been proposed for automatic brain tumor segmentation using four common MRI modalities and achieved remarkable performance. In practice, however, it is common…

2023

STAR Loss: Reducing Semantic Ambiguity in Facial Landmark Detection

CVPR 2023poster

Recently, deep learning-based facial landmark detection has achieved significant improvement. However, the semantic ambiguity problem degrades detection performance. Specifically, the semantic ambiguity causes inconsistent annotation and negatively affects the model's convergence, leading to worse a…

2023

Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models

ICML 2023oral

Language modeling on large-scale datasets improves performance of various downstream tasks. The validation pre-training loss is often used as the evaluation metric for language models since the pre-training loss tends to be well-correlated with downstream performance (which is itself hard to evaluat…

Cited by 52SourcePDFScholar
2023

Uniform Sequence Better: Time Interval Aware Data Augmentation for Sequential Recommendation

AAAI 2023technical

Sequential recommendation is an important task to predict the next-item to access based on a sequence of interacted items. Most existing works learn user preference as the transition pattern from the previous item to the next one, ignoring the time interval between these two items. However, we obser…

2022

Adaptive Weighted Network With Edge Enhancement Module For Monocular Self-Supervised Depth Estimation

ICASSP 2022accepted

Monocular self-supervised depth estimation can be easily applied in many areas since only a single camera is required. However, current methods do not predict well in depth borders. Besides, factors such as occlusion and texture sparsity can lead to the failure of the photometric consistency, affect…

Cited by 0SourceScholar
2022

An Information Theoretic Approach for Attention-Driven Face Forgery Detection

ECCV 2022poster

"Recently, Deepfakes arises as a powerful tool to fool the existing real-world face detection systems, which has received wide attention in both academia and society. Most existing forgery face detection methods use heuristic clues to build a binary forgery detector, which mainly takes advantage of…

Cited by 42SourcePDFScholar
2022

Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack

ECCV 2022poster

"Previous studies have verified that the functionality of black-box models can be stolen with full probability outputs. However, under the more practical hard-label setting, we observe that existing methods suffer from catastrophic performance degradation. We argue this is due to the lack of rich in…

2022

Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action Recognition

AAAI 2022technical

In recent years, self-supervised representation learning for skeleton-based action recognition has been developed with the advance of contrastive learning methods. The existing contrastive learning methods use normal augmentations to construct similar positive samples, which limits the ability to ex…

2022

Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering

NAACL 2022long

Existing video question answering (video QA) models lack the capacity for deep video understanding and flexible multistep reasoning. We propose for video QA a novel model which performs dynamic multistep reasoning between questions and videos. It creates video semantic representation based on the vi…

Cited by 12SourcePDFScholar
2022

Explainable Question Answering based on Semantic Graph by Global Differentiable Learning and Dynamic Adaptive Reasoning

EMNLP 2022main

Multi-hop Question Answering is an agent task for testing the reasoning ability. With the development of pre-trained models, the implicit reasoning ability has been surprisingly improved and can even surpass human performance. However, the nature of the black box hinders the construction of explaina…

Cited by 3SourcePDFScholar
2022

Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question Answering

COLING 2022main

Recently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by computing confidence among questions, options and evidence sentences, which is insuff…

2022

MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation

CVPR 2022poster

Estimating 3D human poses from monocular videos is a challenging task due to depth ambiguity and self-occlusion. Most existing works attempt to solve both issues by exploiting spatial and temporal relationships. However, those works ignore the fact that it is an inverse problem where multiple feasib…

Cited by 415PDFcodeScholar
2022

Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker Tracking

AAAI 2022technical

Multi-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of multi-modal signals remains a challenging issue. In this paper,…

2022

Pose-Guided Feature Disentangling for Occluded Person Re-identification Based on Transformer

AAAI 2022technical

Occluded person re-identification is a challenging task as human body parts could be occluded by some obstacles (e.g. trees, cars, and pedestrians) in certain scenes. Some existing pose-guided methods solve this problem by aligning body parts according to graph matching, but these graph-based method…

2022

SRP-DNN: Learning Direct-Path Phase Difference for Multiple Moving Sound Source Localization

ICASSP 2022accepted

Multiple moving sound source localization in real-world scenarios remains a challenging issue due to interaction between sources, time-varying trajectories, distorted spatial cues, etc. In this work, we propose to use deep learning techniques to learn competing and time-varying direct-path phase dif…

Cited by 0SourceScholar
2022

Self-supervised Learning is More Robust to Dataset Imbalance

ICLR 2022spotlight

Self-supervised learning (SSL) is a scalable way to learn general visual representations since it learns without labels. However, large-scale unlabeled datasets in the wild often have long-tailed label distributions, where we know little about the behavior of SSL. In this work, we systematically inv…

Cited by 204SourcePDFScholar
2022

Solving the Real-Time Motion Planning Problem for Non-Holonomic Robots With Collision Avoidance in Dynamic Scenes

RA-L 2022

Collision-free motion planning allows robots to perform tasks in complex environments, but the conservative bounding curves during collision detection will reduce the robot's dexterity. This paper presents a novel algorithm for real-time motion planning of non-holonomic robots in dynamic scenes. To

Cited by 4SourceScholar
2021

Adversarial Feature Disentanglement for Long-Term Person Re-identification

IJCAI 2021poster

Most existing person re-identification methods are effective in short-term scenarios because of their appearance dependencies. However, these methods may fail in long-term scenarios where people might change their clothes. To this end, we propose an adversarial feature disentanglement network (AFD-N…

Cited by 62SourcePDFScholar
2021

Domain General Face Forgery Detection by Learning to Weight

AAAI 2021technical

In this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making…

2021

Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-Learning

AAAI 2021technical

Recent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re-ID systems face the domain shift issue that training and testing…

2021

Modality-aware Style Adaptation for RGB-Infrared Person Re-Identification

IJCAI 2021poster

RGB-infrared (IR) person re-identification is a challenging task due to the large modality gap between RGB and IR images. Many existing methods bridge the modality gap by style conversion, requiring high-similarity images exchanged by complex CNN structures, like GAN. In this paper, we propose a hig…

Cited by 22SourcePDFScholar
2021

Multi-Scale Spatial Temporal Graph Convolutional Network for Skeleton-Based Action Recognition

AAAI 2021technical

Graph convolutional networks have been widely used for skeleton-based action recognition due to their excellent modeling ability of non-Euclidean data. As the graph convolution is a local operation, it can only utilize the short-range joint dependencies and short-term trajectory but fails to directl…

2021

Supervised Direct-Path Relative Transfer Function Learning for Binaural Sound Source Localization

ICASSP 2021accepted

Direct-path relative transfer function (DP-RTF) refers to the ratio between the direct-path acoustic transfer functions of two channels. Though DP-RTF fully encodes the sound directional cues and serves as a reliable localization feature, it is often erroneously estimated in the presence of noise an…

Cited by 0SourceScholar
2021

Towards Robustness Against Natural Language Word Substitutions

ICLR 2021spotlight

Robustness against word substitutions has a well-defined and widely acceptable form, i.e., using semantically similar words as substitutions, and thus it is considered as a fundamental stepping-stone towards broader robustness in natural language processing. Previous defense methods capture word sub…

2020

A Fast and Accurate Super-Resolution Network Using Progressive Residual Learning

ICASSP 2020accepted

Single-image super-resolution (SISR) task has witnessed great strides in the past few years with the development of deep learning. However, most existing studies concentrate on exploiting much deeper super-resolution networks, which are not friendly to the constrained computation resources. In this…

Cited by 0SourceScholar
2020

API-Net: Robust Generative Classifier via a Single Discriminator

ECCV 2020poster

Robustness of deep neural network classifiers has been attracting increased attention. As for the robust classification problem, a generative classifier typically models the distribution of inputs and labels, and thus can better handle off-manifold examples at the cost of a concise structure. On the…

2020

Anti-Bandit Neural Architecture Search for Model Defense

ECCV 2020poster

Deep convolutional neural networks (DCNNs) have dominated as the best performers in machine learning, but can be challenged by adversarial attacks. In this paper, we defend against adversarial attacks using neural architecture search (NAS) which is based on a comprehensive search of denoising blocks…

Cited by 43SourcePDFScholar
2020

GFNet: A Lightweight Group Frame Network for Efficient Human Action Recognition

ICASSP 2020accepted

Human action recognition aims at assigning an action label to a well-segmented video. Recent work using two-stream or 3D convolutional neural networks achieves high recognition rates at the cost of huge computation complexity, memory footprint, and parameters. In this paper, we propose a lightweight…

Cited by 0SourceScholar
2020

Guided Learning for Weakly-Labeled Semi-Supervised Sound Event Detection

ICASSP 2020accepted

We propose a simple but efficient method termed Guided Learning for weakly-labeled semi-supervised sound event detection (SED). There are two sub-targets implied in weakly-labeled SED: audio tagging and boundary detection. Instead of designing a single model by considering a trade-off between the tw…

Cited by 0SourceScholar
2020

Multi-Branch Learning for Weakly-Labeled Sound Event Detection

ICASSP 2020accepted

There are two sub-tasks implied in the weakly-supervised SED: audio tagging and event boundary detection. Current methods which combine multi-task learning with SED requires annotations both for these two sub-tasks. Since there are only annotations for audio tagging available in weakly-supervised SE…

Cited by 0SourceScholar
2020

Projection & Probability-Driven Black-Box Attack

CVPR 2020poster

Generating adversarial examples in a black-box setting retains a significant challenge with vast practical application prospects. In particular, existing black-box attacks suffer from the need for excessive queries, as it is non-trivial to find an appropriate direction to optimize in the high-dimens…

Cited by 58PDFcodeScholar
2020

Spatial Pyramid Based Graph Reasoning for Semantic Segmentation

CVPR 2020poster

The convolution operation suffers from a limited receptive filed, while global modeling is fundamental to dense prediction tasks, such as semantic segmentation. In this paper, we apply graph convolution into the semantic segmentation task and propose an improved Laplacian. The graph reasoning is dir…

Cited by 227PDFScholar
2020

Spatio-Temporal and Geometry Constrained Network for Automobile Visual Odometry

ICASSP 2020accepted

Visual odometry (VO) is an essence of vision-based localization and mapping system where existing learning-based approaches utilize CNN and RNN to model camera motion and gain promising results. However, these methods lack full use of the relationship between spatial characteristics and temporal clu…

Cited by 0SourceScholar
2020

Unsupervised Monocular Visual-inertial Odometry Network

IJCAI 2020poster

Recently, unsupervised methods for monocular visual odometry (VO), with no need for quantities of expensive labeled ground truth, have attracted much attention. However, these methods are inadequate for long-term odometry task, due to the inherent limitation of only using monocular visual data and t…

2019

A Weight-shared Dual-branch Convolutional Neural Network for Unsupervised Dense Depth Prediction and Camera Motion Estimation

ICASSP 2019accepted

Convolutional Neural Network (CNN) can be used to indiscriminately predict dense depth and camera motion from images, however, ignoring the relationship between depth map and camera motion increases the computational burden to label the datasets and limits the accuracy of the results. In this paper,…

Cited by 0SourceScholar
2019

Expectation-Maximization Attention Networks for Semantic Segmentation

ICCV 2019oral

Self-attention mechanism has been widely used for various tasks. It is designed to compute the representation of each position by a weighted sum of the features at all positions. Thus, it can capture long-range relations for computer vision tasks. However, it is computationally consuming. Since the…

Cited by 783PDFScholar
2019

Separate to Adapt: Open Set Domain Adaptation via Progressive Separation

CVPR 2019poster

Domain adaptation has become a resounding success in leveraging labeled data from a source domain to learn an accurate classifier for an unlabeled target domain. When deployed in the wild, the target domain usually contains unknown classes that are not observed in the source domain. Such setting is…

Cited by 386PDFScholar
2019

Transferable Adversarial Training: A General Approach to Adapting Deep Classifiers

ICML 2019oral

Domain adaptation enables knowledge transfer from a labeled source domain to an unlabeled target domain. A mainstream approach is adversarial feature adaptation, which learns domain-invariant representations through aligning the feature distributions of both domains. However, a theoretical prerequis…

Cited by 314SourcePDFScholar
2019

Universal Adversarial Perturbation via Prior Driven Uncertainty Approximation

ICCV 2019oral

Deep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupe…

Cited by 119PDFScholar
2018

A Discriminatively Learned Feature Embedding Based on Multi-Loss Fusion For Person Search

ICASSP 2018accepted

Person search is a challenging task that requires to address pedestrian detection and person re- identification simultaneously. Though significant progress has been made in detection and re-identification respectively, the similar appearances of persons, pedestrian misdetections and false alarms sti…

Cited by 0SourceScholar
2018

An End-To-End Siamese Convolutional Neural Network for Loop Closure Detection in Visual Slam System

ICASSP 2018accepted

Loop closure detection is essential and important in visual simultaneous localization and mapping (SLAM) systems. Most existing methods typically utilize a separate feature extraction part and a similarity metric part. Compared to these methods, an end-to-end network is proposed in this paper to joi…

Cited by 0SourceScholar
2018

Learning Explicit Shape and Motion Evolution Maps for Skeleton-Based Human Action Recognition

ICASSP 2018accepted

Human action recognition based on skeleton sequences has wide applications in human-computer interaction and intelligent surveillance. Although previous methods have successfully applied Long Short-Term Memory(LSTM) networks to model shape evolution of human actions, it still remains a problem to ef…

Cited by 0SourceScholar
2018

Online Initialization and Automatic Camera-IMU Extrinsic Calibration for Monocular Visual-Inertial SLAM

ICRA 2018poster

Most of the existing monocular visual-inertial SLAM techniques assume that the camera-IMU extrinsic parameters are known, therefore these methods merely estimate the initial values of velocity, visual scale, gravity, biases of gyroscope and accelerometer in the initialization stage. However, it's us…

Cited by 78SourceScholar
2018

Recurrent Squeeze-and-Excitation Context Aggregation Net for Single Image Deraining

ECCV 2018poster

Rain streaks can severely degrade the visibility, which causes many current computer vision algorithms fail to work. So it is necessary to remove the rain from images. We propose a novel deep network architecture based on deep convolutional and recurrent neural networks for single image deraining. A…

Cited by 1039SourcePDFScholar
2018

Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation

CVPR 2018poster

Recent works have shown the benefit of integrating Conditional Random Fields (CRFs) models into deep architectures for improving pixel-level prediction tasks. Following this line of research, in this paper we introduce a novel approach for monocular depth estimation. Similarly to previous works, our…

2017

A novel actuation configuration of robotic hand and the mechanical implementation via postural synergies

ICRA 2017poster

How to design a robotic hand for reproducing the move characteristics of human hand joints is a big challenge in robotics. In this paper, we present an approach to determine the actuation configuration based on the statistical results of hand joint angle in different grasps. A relationship between t…

Cited by 8SourceScholar
2017

Cross-Modality Binary Code Learning via Fusion Similarity Hashing

CVPR 2017poster

Binary code learning has been emerging topic in large-scale cross-modality retrieval recently. It aims to map features from multiple modalities into a common Hamming space, where the cross-modality similarity can be approximated efficiently via Hamming distance. To this end, most existing works lear…

Cited by 255PDFScholar
2017

Human action recognition using Adaptive Hierarchical Depth Motion Maps and Gabor filter

ICASSP 2017accepted

Depth motion maps (DMMs) have shown effectiveness in human action recognition, however, they lose the temporal information and suffer from intra-class variations caused by action speed variations. To address these challenges, we propose a novel method for human action recognition. Firstly, Adaptive…

Cited by 0SourceScholar
2017

Multiple sound source localization based on TDOA clustering and multi-path matching pursuit

ICASSP 2017accepted

Multiple sound source localization in wireless acoustic sensor networks (WASNs) is a challenging problem. Although compressive sensing based methods have shown effectiveness in uncorrelated sources localization, their performance degrades significantly when they are used to locate multiple speech so…

Cited by 0SourceScholar
2016

Probabilistic binaural multiple sources localization based on time-delay compensation estimator and clustering analysis

IROS 2016poster

Sound source localization (SSL) is an essential technique in many applications, such as robot audition, human-robot interaction and speech capturing. However, SSL from a binaural input is still a challenging problem, particularly when multiple sources are active simultaneously. In this work, we prop…

Cited by 0SourceScholar
2015

A predictive model for narrow passage path planner by using Support Vector Machine in changing environments

ICRA 2015poster

Narrow passages in changing environments create huge difficulties, since locations and shapes of narrow passages in Configuration Space(C-space) change frequently. It is very important for a planner to identify narrow passages in real time and boost valid points within them effectively. A novel narr…

Cited by 5SourceScholar
2015

Binaural sound source localization based on generalized parametric model and two-layer matching strategy in complex environments

ICRA 2015poster

Binaural sound source localization is an important technique involving Human-Robot Interaction (HRI), video conference, speech enhancement, etc. In many real application scenarios, especially for closed environments, the affect of reverberation and noise would degrade the precision of position estim…

Cited by 7SourceScholar