← Search

Tao Liu

62 accepted papers

2026

A Multi-Mode Motion Polar Robot: Energy-Saving through Foldable Sail and Transformable Tracks

ICRA 2026poster

Existing polar robots are constrained by limited energy supply, making it difficult to carry out long-term scientific exploration missions, which highlights an urgent demand for energy conservation. An energy-efficient multi-mode motion polar robot is proposed to address this challenge. Both increas…

Cited by 1SourceScholar
2026

Discrete-Periodic Ambiguity Function of Random Communication Signals

ICASSP 2026oral

This paper investigates the ambiguity function (AF) of communication signals carrying random data payloads, which is a fundamental metric characterizing sensing capability in ISAC systems. We first develop a unified analytical framework to evaluate the AF of communication-centric ISAC signals constr…

Cited by 0SourcePDFScholar
2026

Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

ICLR 2026poster

Recent advances in large multimodal models have leveraged image-based tools with reinforcement learning to tackle visual problems. However, existing open-source approaches often exhibit monotonous reasoning patterns and allow only a limited number of interaction turns, making them inadequate for dif…

Cited by 0SourcecodeScholar
2026

No Labels, No Look-Ahead: Unsupervised Online Video Stabilization with Classical Priors

CVPR 2026

We propose a novel unsupervised framework for online video stabilization. Unlike deep learning-based stabilizers that require paired stable/unstable datasets, our method models the classical three-stage stabilization pipeline and integrates a multithreaded buffering mechanism, effectively addressing

Cited by 0SourcecodeScholar
2026

Plantar Compensation Via Dynamic Control of Pneumatic Insoles for Flatfoot Deformity

ICRA 2026poster

Human feet are crucial for supporting body weight and adapting to complex terrains. Adult-acquired flatfoot deformity (AAFD) arises from congenital or acquired causes, impairing the foot's ability to transition between flexible and rigid states, known as the lock-unlock mechanism during the stance a…

Cited by 0Scholar
2026

Position: Early-Stage Quality Assurance in Annotation Pipelines Is More Cost-Effective Than Late-Stage Validation

ICML 2026poster

This position paper argues that the machine learning community should prioritize early-stage quality assurance in annotation pipelines over the prevailing practice of late-stage validation. Data quality bottlenecks increasingly limit foundation model improvement, yet quality assurance research focus…

Cited by 0SourceScholar
2026

Realistic Curriculum Reinforcement Learning for Autonomous and Sustainable Marine Vessel Navigation

AAAI 2026technical

Sustainability is becoming increasingly critical in the maritime transport, encompassing both environmental and social impacts, such as Greenhouse Gas (GHG) emissions and navigational safety. Traditional vessel navigation heavily relies on human experience, often lacking autonomy and emission awaren

Cited by 0SourcePDFScholar
2026

STARK: Strategic Team of Agents for Refining Kernels

ICLR 2026poster

The efficiency of GPU kernels is central to the progress of modern AI, yet optimizing them remains a difficult and labor-intensive task due to complex interactions between memory hierarchies, thread scheduling, and hardware-specific characteristics. While recent advances in large language models (LL…

Cited by 0SourceScholar
2025

Collapsing Sequence-Level Data-Policy Coverage via Poisoning Attack in Offline Reinforcement Learning

UAI 2025

Offline reinforcement learning (RL) heavily relies on the coverage of pre-collected data over the target policy’s distribution. Existing studies aim to improve data-policy coverage to mitigate distributional shifts, but overlook security risks from insufficient coverage, and the single-step analysis

Cited by 0SourcePDFScholar
2025

Dynamic Arch Compensation Based on Controllable Pneumatic Insoles for Flatfoot Deformity

RA-L 2025

Human feet are crucial for supporting body weight and adapting to complex terrains. Adult-acquired flatfoot deformity (AAFD) arises from congenital or acquired causes, impairing the foot's ability to transition between flexible and rigid states, known as the lock-unlock mechanism during the stance a

Cited by 0SourceScholar
2025

Efficient Multi-Robot Task and Path Planning in Large-Scale Cluttered Environments

RA-L 2025

As the potential of multi-robot systems continues to be explored and validated across various real-world applications, such as package delivery, search and rescue, and autonomous exploration, the need to improve the efficiency and quality of task and path planning has become increasingly urgent, par

Cited by 4SourceScholar
2025

Error-Subspace Transform Kalman Filter Based Real-Time Gait Prediction for Rehabilitation Exoskeletons

ICRA 2025

With the rapid development of rehabilitation robotics, there is a pressing need for efficient and accurate gait prediction methods. However, due to the complexity and variability of individual gait characteristics and external disturbances, accurately predicting gait in real time remains a significa

Cited by 0SourceScholar
2025

FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering

CVPR 2025poster

Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated significant potential over traditional rendering techniques, attracting widespread attention from both industry and academia. However, real-time rendering with 3DGS remains a challenging problem, particularly in large-scale, high-reso…

2025

Frequency-Biased Synergistic Design for Image Compression and Compensation

CVPR 2025poster

Compression artifacts removal (CAR), an effective post-processing method to reduce compression distortion in edge-side codecs, demonstrates remarkable results by utilizing convolutional neural networks (CNNs) on high computational power cloud side. Traditional image compression reduces redundancy in…

Cited by 0SourcePDFScholar
2025

From Cradle to Cane: A Two-Pass Framework for High-Fidelity Lifespan Face Aging

NeurIPS 2025poster

Face aging has become a crucial task in computer vision, with applications ranging from entertainment to healthcare. However, existing methods struggle with achieving a realistic and seamless transformation across the entire lifespan, especially when handling large age gaps or extreme head poses. Th…

Cited by 0SourcecodeScholar
2025

GS-EVT: Cross-Modal Event Camera Tracking Based on Gaussian Splatting

ICRA 2025

Reliable self-localization is a foundational skill for many intelligent mobile platforms. This paper explores the use of event cameras for motion tracking thereby providing a solution with inherent robustness under difficult dynamics and illumination. In order to circumvent the challenge of event ca

Cited by 3SourceScholar
2025

GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation

NeurIPS 2025poster

While Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction,…

Cited by 0SourceScholar
2025

One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

ICLR 2025spotlight

Text-to-image generation models can create high-quality images from input prompts. However, they struggle to support the consistent generation of identity-preserving requirements for storytelling. Existing approaches to this problem typically require extensive training in large datasets or additiona…

2025

One-Way Ticket: Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models

CVPR 2025poster

Text-to-Image (T2I) diffusion models have made remarkable advancements in generative modeling; however, they face a trade-off between inference speed and image quality, posing challenges for efficient deployment. Existing distilled T2I models can generate high-fidelity images with fewer sampling ste…

2025

Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph Generation

AAAI 2025technical

Open-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual representations. This enables the identification of novel visual relationships, making it applicable to real-world scena…

Cited by 1SourcePDFScholar
2025

Robust Adversarial Training for Industrial Defect Classification with Long-Tailed Data

ICASSP 2025accepted

Deep neural networks are vulnerable to adversarial examples which fool model predictions by adding imperceptible perturbations to natural examples. Adversarial training is effective in defending against adversarial attacks but faces a challenge with long-tailed data, where the over-compression of ta…

Cited by 0SourceScholar
2025

VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization

AAAI 2025technical

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinc…

Cited by 0SourcePDFScholar
2024

Beyond Traditional Threats: A Persistent Backdoor Attack on Federated Learning

AAAI 2024technical

Backdoors on federated learning will be diluted by subsequent benign updates. This is reflected in the significant reduction of attack success rate as iterations increase, ultimately failing. We use a new metric to quantify the degree of this weakened backdoor effect, called attack persistence. Give…

2024

DiffDub: Person-Generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-Encoder

ICASSP 2024accepted

Generating high-quality and person-generic visual dubbing remains a challenge. Recent innovation has seen the advent of a two-stage paradigm, decoupling the rendering and lip synchronization process facilitated by intermediate representation as a conduit. Still, previous methodologies rely on rough…

Cited by 0SourceScholar
2024

EVIT: Event-based Visual-Inertial Tracking in Semi-Dense Maps Using Windowed Nonlinear Optimization

IROS 2024poster

Event cameras are an interesting visual exteroceptive sensor that reacts to brightness changes rather than integrating absolute image intensities. Owing to this design, the sensor exhibits strong performance in situations of challenging dynamics and illumination conditions. While event-based simulta…

Cited by 3SourceScholar
2024

Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference

NeurIPS 2024poster

One of the main drawback of diffusion models is the slow inference time for image generation. Among the most successful approaches to addressing this problem are distillation methods. However, these methods require considerable computational resources. In this paper, we take another approach to diff…

2024

Foot Arch Stiffness-Based Dynamic Plantar Support Control of Human Walking Gait with Active Pneumatic Insoles

IROS 2024poster

The human foot arch plays a significant role in bearing weight, keeping balance, and walking efficiently. In this study, we present a pneumatic arch support insole (PASI) and a foot arch stiffness-based dynamic plantar support control to reduce the metabolic cost of walking. We first obtain the foot…

Cited by 1SourceScholar
2024

Foot Shape-Dependent Resistive Force Model for Bipedal Walkers on Granular Terrains

ICRA 2024poster

Legged robots have demonstrated high efficiency and effectiveness in unstructured and dynamic environments. However, it is still challenging for legged robots to achieve rapid and efficient locomotion on deformable, yielding substrates, such as granular terrains. We present an enhanced resistive for…

Cited by 3SourceScholar
2024

HMD-Poser: On-Device Real-time Human Motion Tracking from Scalable Sparse Observations

CVPR 2024poster

It is especially challenging to achieve real-time human motion tracking on a standalone VR Head-Mounted Display (HMD) such as Meta Quest and PICO. In this paper we propose HMD-Poser the first unified approach to recover full-body motions using scalable sparse observations from HMD and body-worn IMUs…

2024

SemanticHuman-HD: High Resolution Semantic disentangled 3D Human Generation

ECCV 2024poster

"With the development of neural radiance fields and generative models, numerous methods have been proposed for learning 3D human generation from 2D images. These methods allow control over the pose of the generated 3D human and enable rendering from different viewpoints. However, none of these metho…

2024

Towards Open Domain Text-Driven Synthesis of Multi-Person Motions

ECCV 2024poster

"This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one or two subjects from in-the-wild prompts, mainly due to the…

Cited by 10SourcePDFScholar
2023

Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation

ICASSP 2023accepted

Audio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive and costly. The lack of data resource results in poor synthesis effect. To alleviate this issue, we propose to use TTS (…

Cited by 0SourceScholar
2023

Multi-Speaker End-to-End Multi-Modal Speaker Diarization System for the MISP 2022 Challenge

ICASSP 2023accepted

This paper presents the design and implementation of our system for Track 1 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We design an end-to-end transformer-based multi-talker system. The transformer backbone is well-suited to capture long-term features, which is cru…

Cited by 0SourceScholar
2023

Natural Actor-Critic for Robust Reinforcement Learning with Function Approximation

NeurIPS 2023poster

We study robust reinforcement learning (RL) with the goal of determining a well-performing policy that is robust against model mismatch between the training simulator and the testing environment. Previous policy-based robust RL algorithms mainly focus on the tabular setting under uncertainty sets th…

2023

Penguin: Parallel-Packed Homomorphic Encryption for Fast Graph Convolutional Network Inference

NeurIPS 2023poster

The marriage of Graph Convolutional Network (GCN) and Homomorphic Encryption (HE) enables the inference of graph data on the cloud with significantly enhanced client data privacy. However, the tremendous computation and memory overhead associated with HE operations challenges the practicality of HE-…

2023

Provably Fast Convergence of Independent Natural Policy Gradient for Markov Potential Games

NeurIPS 2023poster

This work studies an independent natural policy gradient (NPG) algorithm for the multi-agent reinforcement learning problem in Markov potential games. It is shown that, under mild technical assumptions and the introduction of the \textit{suboptimality gap}, the independent NPG method with an oracle…

2023

Score Priors Guided Deep Variational Inference for Unsupervised Real-World Single Image Denoising

ICCV 2023poster

Real-world single image denoising is crucial and practical in computer vision. Bayesian inversions combined with score priors now have proven effective for single image denoising but are limited to white Gaussian noise. Moreover, applying existing score-based methods for real-world denoising require…

Cited by 15PDFcodeScholar
2023

SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network Inference

ICML 2023poster

Homomorphic Encryption (HE) is a promising technology to protect clients' data privacy for Machine Learning as a Service (MLaaS) on public clouds. However, HE operations can be orders of magnitude slower than their counterparts for plaintexts and thus result in prohibitively high inference latency,…

2022

Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement Learning

NeurIPS 2022accept

We study policy optimization for Markov decision processes (MDPs) with multiple reward value functions, which are to be jointly optimized according to given criteria such as proportional fairness (smooth concave scalarization), hard constraints (constrained MDP), and max-min trade-off. We propose an…

2022

Falconn++: A Locality-sensitive Filtering Approach for Approximate Nearest Neighbor Search

NeurIPS 2022accept

We present Falconn++, a novel locality-sensitive filtering (LSF) approach for approximate nearest neighbor search on angular distance. Falconn++ can filter out potential far away points in any hash bucket before querying, which results in higher quality candidates compared to other hashing-based so…

2022

Learning from Few Samples: Transformation-Invariant SVMs with Composition and Locality at Multiple Scales

NeurIPS 2022accept

Motivated by the problem of learning with small sample sizes, this paper shows how to incorporate into support-vector machines (SVMs) those properties that have made convolutional neural networks (CNNs) successful. Particularly important is the ability to incorporate domain knowledge of invariances,…

2021

A Passive Hydraulic Auxiliary System Designed for Increasing Legged Robot Payload and Efficiency

ICRA 2021poster

Load-carrying capability is an essential criterion in legged robots' practical application. This paper proposes an unpowered hydraulic auxiliary system to improve the legged robot's loading capability and energy efficiency. For humans, it has been widely hypothesized that intra-abdominal pressure ca…

Cited by 6SourceScholar
2021

Design and Clinical Validation of a Robotic Ankle-Foot Simulator With Series Elastic Actuator for Ankle Clonus Assessment Training

RA-L 2021

To fulfill the need for reliable and consistent medical training of the neurological examination technique to assess ankle clonus, a series elastic actuator (SEA) based haptic training simulator was proposed and developed. The simulator's mechanism (a hybrid of belt and linkage drive) and controller

Cited by 10SourceScholar
2021

Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPs

NeurIPS 2021poster

We address the issue of safety in reinforcement learning. We pose the problem in an episodic framework of a constrained Markov decision process. Existing results have shown that it is possible to achieve a reward regret of $\tilde{\mathcal{O}}(\sqrt{K})$ while allowing an $\tilde{\mathcal{O}}(\sqrt{…

Cited by 95SourcePDFScholar
2021

Real-Time Human Lower Limbs Motion Estimation and Feedback for Potential Applications in Robotic Gait Aid and Training

ICRA 2021poster

Real-time lower limbs motion or gait measurement is an important part in human-robotic interaction for the control of robotic walkers and rehabilitation devices. Laser range finder or infrared sensor that is mounted on the device has been widely used in applications. Although these sensors can provi…

Cited by 4SourceScholar
2021

Sliding Mode Control of the Semi-active Hover Backpack Based on the Bioinspired Skyhook Damper Model

ICRA 2021poster

It is inevitable for human to bear the gravitational and inertial force when carrying loads. The impact force exerted on human body is originated from the inertial force which can increase the energy expenditure and cause injury to human body. This paper proposes a semi-active hover backpack with co…

Cited by 6SourceScholar
2021

What If We Could Not See? Counterfactual Analysis for Egocentric Action Anticipation

IJCAI 2021poster

Egocentric action anticipation aims at predicting the near future based on past observation in first-person vision. While future actions may be wrongly predicted due to the dataset bias, we present a counterfactual analysis framework for egocentric action anticipation (CA-EAA) to enhance the capacit…

Cited by 16SourcePDFScholar
2020

Towards Accurate Scene Text Recognition With Semantic Reasoning Networks

CVPR 2020poster

Scene text image contains two levels of contents: visual texture and semantic information. Although the previous scene text recognition methods have made great progress over the past few years, the research on mining semantic information to assist text recognition attracts less attention, only RNN-l…

Cited by 424PDFScholar
2020

Two Shank-Mounted IMUs-Based Gait Analysis and Classification for Neurological Disease Patients

RA-L 2020

Automatic gait measurement and analysis is an enabling tool for intelligent healthcare and robotics-assisted rehabilitation. This letter proposes a novel two shank-mounted inertial measurement units (IMU)-based method on gait analysis and classification for three different neurological diseases. The

Cited by 71SourceScholar
2019

Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples

CVPR 2019poster

Image compression-based approaches for defending against the adversarial-example attacks, which threaten the safety use of deep neural networks (DNN), have been investigated recently. However, prior works mainly rely on directly tuning parameters like compression rate, to blindly reduce image featur…

Cited by 337PDFScholar
2019

Machine Vision Guided 3D Medical Image Compression for Efficient Transmission and Accurate Segmentation in the Clouds

CVPR 2019poster

Cloud based medical image analysis has become popular recently due to the high computation complexities of various deep neural network (DNN) based frameworks and the increasingly large volume of medical images that need to be processed. It has been demonstrated that for medical images the transmissi…

Cited by 48PDFScholar
2015

A robotic bipedal model for human walking with slips

ICRA 2015poster

Slip is the major cause of falls in human locomotion. We present a new bipedal modeling approach to capture and predict human walking locomotion with slips. Compared with the existing bipedal models, the proposed slip walking model includes the human foot rolling effects, the existence of the double…

Cited by 43SourceScholar