← Search

Chao Liu

57 accepted papers

2026

FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning

AAAI 2026technical

Representation learning is fundamental to modern machine learning, powering applications such as text retrieval and multimodal understanding. However, learning robust and generalizable representations remains challenging. While prior work has demonstrated that active noise injection, a form of data

Cited by 0SourcePDFScholar
2026

Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

ICML 2026spotlight

Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models excel at global coherence and visual fidelity but suffer from slow inference, while autoregressive model…

Cited by 0SourceScholar
2026

InstEmb: Instruction-Following Embeddings through Glimpses of the Future

ICML 2026poster

Recent advances have empowered large language models (LLMs) with remarkable fine-grained instruction-following capabilities in text generation tasks. However, embedding methods typically rely solely on the hidden state of the input's last token, limiting their ability to capture complete semantic si…

Cited by 0SourceScholar
2026

Mode Seeking meets Mean Seeking for Long Video Generation

ICML 2026poster

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to narrow domains. While multi-resolution image training works because higher resolution is largely an interpolation of th…

Cited by 7SourceScholar
2026

NeuralActuator: Neural Actuation Modeling for Robot Dynamics and External Force Perception

RSS 2026poster

Differentiable simulators have advanced policy learning and model-based control across diverse robotic tasks. To date, actuator dynamics remain underexplored and are a major source of sim-to-real error, especially on low-cost platforms where the linear current–torque model τ = K_tI breaks down under…

Cited by 0SourceScholar
2026

Stable Vision-Based Robot Kinematic Control With Deep Learning-Based Oriented Object Detector

RA-L 2026

Recent advances in machine learning and deep learning have significantly enhanced robot control by improving object detection and visual feature extraction. However, ensuring theoretical guarantees of stability and convergence in learning-enabled control systems remains a major challenge. In this pa

Cited by 0SourceScholar
2026

Transition Matching Distillation for Fast Video Generation

CVPR 2026

Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel

Cited by 0SourceScholar
2025

A Novel Path Following Method Based on Whole-Body Deviation Evaluation for Hyper-Redundant Robots

RA-L 2025

The accuracy of path following is crucial for collision-free navigation of hyper-redundant robots, especially in narrow environments. However, the existing path following methods only consider the deviations of joints and the end effector, while ignoring the deviations of the robot body. In this let

Cited by 1SourceScholar
2025

BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations

CVPR 2025poster

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual primitives -- blob video representation, a general represen…

Cited by 3SourcePDFScholar
2025

Coherent 3D Portrait Video Reconstruction via Triplane Fusion

CVPR 2025poster

Recent breakthroughs in single-image 3D portrait reconstruction have enabled telepresence systems to stream 3D portrait videos from a single camera in real-time, democratizing telepresence. However, per-frame 3D reconstruction exhibits temporal inconsistency and forgets the user's appearance. On the…

Cited by 1SourcePDFScholar
2025

DASA-Trans-STM: Adaptive Efficient Transformer for Short Text Matching using Data Augmentation and Semantic Awareness

EMNLP 2025

Rencent advancements in large language models (LLM) have shown impressive versatility across various tasks. Short text matching is one of the fundamental technologies in natural language processing. In previous studies, the common approach to applying them to Chinese is segmenting each sentence into

Cited by 0SourcePDFScholar
2025

Delicate Operation of a Microneedle-Forceps Mechanism for Ultra-Flexible Probe Implantation

RA-L 2025

The implantation of ultra-flexible neural probes is one of the key technical challenges in the field of brain-computer interfaces. Existing implantation techniques are constrained in both the overall surgical procedure and operational reliability. This study introduces a novel microneedle-forceps me

Cited by 2SourceScholar
2025

Hybrid Tension and Configuration Control of Cable-Driven Hyper-Redundant Robots for High Accuracy and Stability

RA-L 2025

Due to the advantages of dexterity and adaptability, cable-driven hyper-redundant robots (CDHRRs) are promising for detection in confined spaces like narrow internal cavities. However, due to redundant degrees of freedom (DoFs), CDHRRs are susceptible to the singular configuration, which aggravates

Cited by 1SourceScholar
2025

Learning Object Properties Using Robot Proprioception via Differentiable Robot-Object Interaction

ICRA 2025

Differentiable simulation has become a powerful tool for system identification. While prior work has focused on identifying robot properties using robot-specific data or object properties using object-specific data, our approach calibrates object properties by using information from the robot, witho

Cited by 5SourceScholar
2025

MM-Tracker: Motion Mamba for UAV-platform Multiple Object Tracking

AAAI 2025technical

Multiple object tracking (MOT) from unmanned aerial vehicle (UAV) platforms requires efficient motion modeling. This is because UAV-MOT faces both local object motion and global camera motion. Motion blur also increases the difficulty of detecting large moving objects. Previous UAV motion modeling a…

2025

MambaRF: A Bi-directional Mamba Structure for Radio Frequency Signal Classification of Unmanned Aerial Vehicle

ICASSP 2025accepted

The rapid development of Unmanned Aerial Vehicle (UAV) technology has facilitated the widespread use of UAVs in daily life, this advancement brings huge regulatory demand for UAVs. Automatic identification of UAVs using radio frequency (RF) signals can effectively reduce regulatory costs. Previous s…

Cited by 0SourceScholar
2025

NETracer: A Topology-Aware Iterative Tracing Approach for Tubular Structure Extraction

ICCV 2025poster

Extracting tubular structures from images is a widespread and challenging task in computer vision. To explore these continuous structures, iterative tracing methods offer a promising direction. However, in scenes with dense and blurred branches, existing tracing methods tend to jump to adjacent bran…

2025

Not-So-Optimal Transport Flows for 3D Point Cloud Generation

ICLR 2025poster

Learning generative models of 3D point clouds is one of the fundamental problems in 3D generative learning. One of the key properties of point clouds is their permutation invariance, i.e., changing the order of points in a point cloud does not change the shape they represent. In this paper, we analy…

Cited by 0SourcePDFScholar
2025

Passivity Filters for Bilateral Teleoperation with Variable Impedance Control

ICRA 2025

In robotic teleoperation, it is crucial to be able to dynamically adjust interactions with the environment. Drawing inspiration from human behavior during interactions, Variable Impedance Control (VIC) has been widely adopted to enhance robotic flexibility and adaptability. However, maintaining the

Cited by 0SourceScholar
2025

Similarity Memory Prior is All You Need for Medical Image Segmentation

ICCV 2025poster

In recent years, it has been found that "grandmother cells" in the primary visual cortex (V1) of macaques can directly recognize visual input with complex shapes. This inspires us to examine the value of these cells in promoting the research of medical image segmentation. In this paper, we design a…

2024

A Novel Cascade Instruction Tuning Method for Biomedical NER

ICASSP 2024accepted

Large language models(LLMs) have achieved remarkable performance on various tasks. However, LLMs suffer from severe limitations in domain generalisation, primarily due to inherent limitations. Closed-source LLMs face constraints in fine-tuning, while open-source LLMs contend with the scarcity of dom…

Cited by 0SourceScholar
2024

Application of SNNS Model Based On Multi-Dimensional Attention In Drone Radio Frequency Signal Classification

ICASSP 2024accepted

Spiking Neural Networks (SNNs) are attracting attention due to their energy efficiency and importance in neuromorphic computing. Therefore, we propose an SNN-based method for classifying drone RF signals in complex electromagnetic environments. Specifically, we designed a new SNNs model called Spiki…

Cited by 0SourceScholar
2024

Compositional Text-to-Image Generation with Dense Blob Representations

ICML 2024poster

Existing text-to-image models struggle to follow complex text prompts, raising the need for extra grounding inputs for better controllability. In this work, we propose to decompose a scene into visual primitives - denoted as dense blob representations - that contain fine-grained details of the scene…

Cited by 16SourcePDFScholar
2024

DIFFTACTILE: A Physics-based Differentiable Tactile Simulator for Contact-rich Robotic Manipulation

ICLR 2024poster

We introduce DIFFTACTILE, a physics-based differentiable tactile simulation system designed to enhance robotic manipulation with dense and physically accurate tactile feedback. In contrast to prior tactile simulators which primarily focus on manipulating rigid bodies and often rely on simplified app…

2024

DeepBranchTracer: A Generally-Applicable Approach to Curvilinear Structure Reconstruction Using Multi-Feature Learning

AAAI 2024technical

Curvilinear structures, which include line-like continuous objects, are fundamental geometrical elements in image-based applications. Reconstructing these structures from images constitutes a pivotal research area in computer vision. However, the complex topology and ambiguous image evidence render…

2024

Learning to Jointly Understand Visual and Tactile Signals

ICLR 2024poster

Modeling and analyzing object and shape has been well studied in the past. However, manipulation of these complex tools and articulated objects remains difficult for autonomous agents. Our human hands, however, are dexterous and adaptive. We can easily adapt a manipulation skill on one object to all…

Cited by 6SourcePDFScholar
2024

Liquids Identification and Manipulation via Digitally Fabricated Impedance Sensors

ICRA 2024poster

Despite recent exponential advancements in computer vision and reinforcement learning, it remains challenging for robots to interact with liquids. These challenges are particularly pronounced due to the limitations imposed by opaque containers, transparent liquids, fine-grained splashes, and visual…

Cited by 2SourceScholar
2024

Search for Gravitational Wave Probes - A Self-Supervised Learning for Pulsars Based on Signal Contexts

ICASSP 2024accepted

The recent successful detection of gravitational waves (GWs) at nanohertz based on pulsar timing arrays has underscored the growing significance of searching for new pulsars, which serve as valuable probes for GWs. However, one of the challenges in this endeavor is the lack of labeled data, which ca…

Cited by 0SourceScholar
2024

Thin-Shell Object Manipulations With Differentiable Physics Simulations

ICLR 2024spotlight

In this work, we aim to teach robots to manipulate various thin-shell materials. Prior works studying thin-shell object manipulation mostly rely on heuristic policies or learn policies from real-world video demonstrations, and only focus on limited material types and tasks (e.g., cloth unfolding).…

Cited by 5SourcePDFScholar
2023

A Method of Constructing and Automatically Labeling Radio Frequency Signal Training Dataset for UAV

ICASSP 2023accepted

The problem of signal detection and classification of multiple UAVs can be solved using object detection techniques in computer vision. However, this requires collecting and labeling a large amount of reliable raw data. Since the UAV signal dataset cannot be directly applied to object detection, we…

Cited by 0SourceScholar
2023

Adaptive Robust Model Predictive Control for Bilateral Teleoperation

IROS 2023poster

In this work, we use recent developments in the field of adaptive robust Model Predictive Control (MPC) to build a controller for bilateral teleoperation systems. To guarantee robust constraint satisfaction, we incorporate polytopic tube controllers in the MPC design. In addition, we use online lear…

Cited by 2SourceScholar
2023

LADA-Trans-NER: Adaptive Efficient Transformer for Chinese Named Entity Recognition Using Lexicon-Attention and Data-Augmentation

AAAI 2023technical

Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the semantic relationship before and after the sentence after integ…

Cited by 10SourcePDFScholar
2023

Precognition in Contextual Spoken Language Understanding via Knowledge Distillation

ICASSP 2023accepted

Task-oriented dialogue systems have become overwhelmingly popular in recent researches. Spoken Language Understanding (SLU) is widely used to extract the semantics frame of user queries and comprehend users’ intent/emotion/dialogue state in task-oriented dialogue systems. Most previous works on such…

Cited by 0SourceScholar
2023

Towards Cooperative Flight Control Using Visual-Attention

IROS 2023poster

The cooperation of a human pilot with an autonomous agent during flight control realizes parallel autonomy. We propose an air-guardian system that facilitates cooperation between a pilot with eye tracking and a parallel end-to-end neural control system. Our vision-based air-guardian system combines…

Cited by 7SourceScholar
2022

ActionSense: A Multimodal Dataset and Recording Framework for Human Activities Using Wearable Sensors in a Kitchen Environment

NeurIPS 2022accept

This paper introduces ActionSense, a multimodal dataset and recording framework with an emphasis on wearable sensing in a kitchen environment. It provides rich, synchronized data streams along with ground truth data to facilitate learning pipelines that could extract insights about how humans inter…

Cited by 59SourcePDFScholar
2022

LEGO-ABSA: A Prompt-based Task Assemblable Unified Generative Framework for Multi-task Aspect-based Sentiment Analysis

COLING 2022main

Aspect-based sentiment analysis (ABSA) has received increasing attention recently. ABSA can be divided into multiple tasks according to the different extracted elements. Existing generative methods usually treat the output as a whole string rather than the combination of different elements and only…

Cited by 76SourcePDFScholar
2022

Neural Interferometry: Image Reconstruction from Astronomical Interferometers Using Transformer-Conditioned Neural Fields

AAAI 2022technical

Astronomical interferometry enables a collection of telescopes to achieve angular resolutions comparable to that of a single, much larger telescope. This is achieved by combining simultaneous observations from pairs of telescopes such that the signal is mathematically equivalent to sampling the Four…

2022

Online Adaptive Identification and Switching of Soft Contact Model Based on ART-II Method

ICRA 2022poster

In order to obtain a high-precision contact model that can properly describe the target soft tissue, this paper proposes a hybrid soft contact model based on a clustering algorithm ART-II, which selects the most suitable soft contact model according to the surgical environment. The least-square meth…

Cited by 3SourceScholar
2022

TranSHER: Translating Knowledge Graph Embedding with Hyper-Ellipsoidal Restriction

EMNLP 2022main

Knowledge graph embedding methods are important for the knowledge graph completion (or link prediction) task.One state-of-the-art method, PairRE, leverages two separate vectors to model complex relations (i.e., 1-to-N, N-to-1, and N-to-N) in knowledge graphs. However, such a method strictly restrict…

2021

Self-Supervised Learning on 3D Point Clouds by Learning Discrete Generative Models

CVPR 2021poster

While recent pre-training tasks on 2D images have proven very successful for transfer learning, pre-training for 3D data remains challenging. In this work, we introduce a general method for 3D self-supervised representation learning that 1) remains agnostic to the underlying neural network architect…

Cited by 74PDFScholar
2019

Characterizing Nanoparticle Swarms With Tuneable Concentrations for Enhanced Imaging Contrast

RA-L 2019

Microrobots capable of performing targeted delivery are promising for biomedical applications. Due to the restriction of their small size and volume, however, the real-time <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">in vivo</i> imaging strategy

Cited by 28SourceScholar
2019

Neural RGB(r)D Sensing: Depth and Uncertainty From a Video Camera

CVPR 2019oral

Depth sensing is crucial for 3D reconstruction and scene understanding. Active depth sensors provide dense metric measurements, but often suffer from limitations such as restricted operating ranges, low spatial resolution, sensor interference, and high power consumption. In this paper, we propose a…

Cited by 170PDFScholar
2019

Spiral Zipper Manipulator for Aerial Grasping and Manipulation

IROS 2019poster

This paper presents a novel manipulator for aerial vehicles to perform grasping and manipulation tasks. The goal is to design a low-cost, relatively light but strong manipulator with a large workspace and compact storage space that can be mounted on an unmanned aerial system. A novel design solution…

Cited by 9SourceScholar
2019

Toward Lateral Aerial Grasping & Manipulation Using Scalable Suction

ICRA 2019poster

This paper is an initial step toward the realization of an aerial robot that can perform lateral physical work, such as drilling a hole or fastening a screw in a wall. Aerial robots are capable of high maneuverability and can provide access to locations that would be difficult or impossible for grou…

Cited by 12SourceScholar
2017

PaintPots: Low cost, accurate, highly customizable potentiometers for position sensing

ICRA 2017poster

The PaintPot manufacturing process is a new way to create low-cost, low-profile, highly customizable potentiometers for position sensing in robotic applications. It uses widely accessible materials, requires no special expertise, and creates custom potentiometers in a variety of shapes and sizes, in…

Cited by 12SourceScholar
2015

Real-Time Visual Analysis of Microvascular Blood Flow for Critical Care

CVPR 2015poster

Microcirculatory monitoring plays an important role in diagnosis and treatment of critical care patients. Sidestream Dark Field (SDF) imaging devices have been used to visualize and support interpretation of the micro-vascular blood flow. However, due to subsurface scattering within the tissue that…

Cited by 16SourcePDFScholar
2015

Scalable clustering based on enhanced-SMART for large-scale FMRI datasets

ICASSP 2015accepted

In this paper, we propose a scalable clustering paradigm to address the problems of excessive computational load and limited clustering performance in large-scale data. The proposed method employs the enhanced splitting merging awareness tactics (E-SMART) algorithm. The large-scale dataset is divide…

Cited by 0SourceScholar