← Search

Gang Liu

35 accepted papers

2026

A Centerline-Aligned Frenet Graph Framework for Surface-Based Path Planning in Pipeline Environments

ICRA 2026poster

Pipeline inspection is essential for maintaining the safety of critical infrastructure, but manual inspection is dangerous and inefficient, and existing robotic solutions struggle to handle curved and constrained surfaces. Traditional planning methods are either computationally expensive or prone to…

Cited by 0Scholar
2026

Bayes-inspired Integration of Pretrained Priors and Few-Shot Evidence for Few-Shot Classification

ICML 2026poster

Few-shot classification aims to adapt a pretrained model to novel classes with limited examples. While current methods often heuristically combine pretrained knowledge and few-shot evidence, we seek a more principled understanding of their relationship. In this paper, we propose a Bayesian-inspired …

Cited by 0SourceScholar
2026

Design of an Active Haptic Interface Using Proprioception Feedback for Continuous Endovascular Teleoperation

RA-L 2026

Force feedback is essential for safe endovascular teleoperation, yet typically constrained by complex sensor integration. This article presents a compact active haptic interface system designed for robotic catheterization. Leveraging the intrinsic proprioception of a Permanent Magnet Synchronous Mot

Cited by 0SourceScholar
2026

Graph Diffusion Transformers are In-Context Molecular Designers

ICLR 2026poster

In-context learning lets large models adapt to new tasks from a few demonstrations, but it has shown limited success in molecular design, where labeled data are scarce and properties span millions of biological assays and material measurements. We introduce demonstration-conditioned diffusion models…

Cited by 0SourcecodeScholar
2026

IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning

ICLR 2026poster

Large Vision-Language Models (LVLMs) achieve impressive performance across multiple tasks. A significant challenge, however, is their prohibitive inference cost when processing high-resolution visual inputs. While visual token pruning has emerged as a promising solution, existing methods that primar…

Cited by 0SourceScholar
2026

Next Generation Active Learning: Mixture of LLMs in the Loop

AAAI 2026technical

With the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering the annotation quality, labels generated by LLMs often fall sh

Cited by 0SourcePDFScholar
2026

Protein Structure Tokenization via Geometric Byte Pair Encoding

ICLR 2026poster

Protein structure is central to biological function, and enabling multimodal protein models requires joint reasoning over sequence, structure, and function. A key barrier is the lack of principled protein structure tokenizers (PSTs): existing approaches fix token size or rely on continuous vector co…

Cited by 0SourcecodeScholar
2025

CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object Detection

NeurIPS 2025poster

With the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive ins…

Cited by 0SourcecodeScholar
2025

Directed Graph Grammars for Sequence-based Learning

ICML 2025poster

Directed acyclic graphs (DAGs) are a class of graphs commonly used in practice, with examples that include electronic circuits, Bayesian networks, and neural architectures. While many effective encoders exist for DAGs, it remains challenging to decode them in a principled manner, because the nodes o…

2025

Foundation Molecular Grammar: Multi-Modal Foundation Models Induce Interpretable Molecular Graph Languages

ICML 2025poster

Recent data-efficient molecular generation approaches exploit graph grammars to introduce interpretability into the generative models. However, grammar learning therein relies on expert annotation or unreliable heuristics for algorithmic inference. We propose Foundation Molecular Grammar (FMG), whic…

2025

IMM-MOT: A Novel 3D Multi-object Tracking Framework with Interacting Multiple Model Filter

IROS 2025

3D Multi-Object Tracking (MOT) provides the trajectories of surrounding objects, assisting robots or vehicles in smarter path planning and obstacle avoidance. Existing 3D MOT methods based on the Tracking-by-Detection framework typically use a single motion model to track an object throughout its en

Cited by 1SourcecodeScholar
2025

In-Pipe Navigation Development Environment and a Smooth Path Planning Method on Pipeline Surface

ICRA 2025

Autonomous in-pipe inspection robots can automatically navigate through complex pipeline networks and detect potential risks from corrosion and defects, demonstrating great potential for replacing costly manual inspections. However, there is no publicly available simulation environment where researc

Cited by 6SourceScholar
2025

Learning Molecular Representation in a Cell

ICLR 2025poster

Predicting drug efficacy and safety in vivo requires information on biological responses (e.g., cell morphology and gene expression) to small molecule perturbations. However, current molecular representation learning methods do not provide a comprehensive view of cell states under these perturbation…

2025

Learning Repetition-Invariant Representations for Polymer Informatics

NeurIPS 2025poster

Polymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit…

Cited by 0SourceScholar
2025

Multimodal Large Language Models for Inverse Molecular Design with Retrosynthetic Planning

ICLR 2025poster

While large language models (LLMs) have integrated images, adapting them to graphs remains challenging, limiting their applications in materials and drug design. This difficulty stems from the need for coherent autoregressive generation across texts and graphs. To address this, we introduce Llamole,…

2024

An Attention-Enhanced Retentive Broad Learning System for Subject-Generic Emotion Recognition from EEG Signals

ICASSP 2024accepted

Emotion recognition (ER) utilizing electroencephalography (EEG) is significant in affective brain-computer interface research. Recent advances have underscored the supremacy of deep learning-based ER techniques over traditional statistical methods. Still, challenges persist in extracting subject-spe…

Cited by 0SourceScholar
2024

Enhancing Generalization in Medical Visual Question Answering Tasks Via Gradient-Guided Model Perturbation

ICASSP 2024accepted

Leveraging pre-trained visual language models has become a widely adopted approach for improving performance in downstream visual question answering (VQA) applications. However, in the specialized field of medical VQA, the scarcity of available data poses a significant barrier to achieving reliable…

Cited by 0SourceScholar
2024

Graph Diffusion Transformers for Multi-Conditional Molecular Generation

NeurIPS 2024oral

Inverse molecular design with diffusion models holds great potential for advancements in material and drug discovery. Despite success in unconditional molecule generation, integrating multiple properties such as synthetic score and gas permeability as condition constraints into diffusion models rema…

2024

Modeling Personalized Retweeting Behaviors for Multi-Stage Cascade Popularity Prediction

IJCAI 2024poster

Predicting the size of message cascades is critical in various applications, such as online advertising and early detection of rumors. However, most existing deep learning approaches rely on cascade observation, which hinders accurate cascade prediction before message posting. Besides, these approac…

2024

PECR: Parameter-Efficient Transfer Learning with Cross-Modal Representation Learning for Remote Sensing Visual Question Answering

ICASSP 2024accepted

Remote sensing (RS) visual question answering (VQA) aims to provide accurate answers to questions related to RS images. Transformer-based models have gradually become popular to solve RS VQA tasks. Due to the ever-growing model size, full-parameter training of the model becomes prohibitively costly.…

Cited by 0SourceScholar
2024

SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark

ACL 2024short

The paper introduces SceMQA, a novel benchmark for scientific multimodal question answering at the college entrance level. It addresses a critical educational phase often overlooked in existing benchmarks, spanning high school to pre-college levels. SceMQA focuses on core science subjects including…

Cited by 5SourcePDFScholar
2023

Capturing the Motion of Every Joint: 3D Human Pose and Shape Estimation with Independent Tokens

ICLR 2023top-25%

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due to its inherent ambiguity. To improve precision, existing m…

2023

Data-Centric Learning from Unlabeled Graphs with Diffusion Model

NeurIPS 2023poster

Graph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then…

2023

Electromagnetic Clutch-Based Ankle Exosuit for Assisting Stroke Survivors With Different Body Sizes

RA-L 2023

Exosuits can be effective in aiding stroke rehabilitation. However, single-motor exosuits are <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">challenging</b> to adapt to users with different body sizes. <italic xmlns:mml="http://www.w3.org/1998/Math/Ma

Cited by 5SourceScholar
2023

Improving Transformer-Based Networks with Locality for Automatic Speaker Verification

ICASSP 2023accepted

Recently, Transformer-based architectures have been explored for speaker embedding extraction. Although the Transformer employs the self-attention mechanism to efficiently model the global interaction between token embeddings, it is inadequate for capturing short-range local context, which is essent…

Cited by 17SourceScholar
2022

D&D: Learning Human Dynamics from Dynamic Camera

ECCV 2022poster

"3D human pose estimation from a monocular video has recently seen significant improvements. However, most state-of-the-art methods are kinematics-based, which are prone to physically implausible motions with pronounced artifacts. Current dynamics-based methods can predict physically plausible motio…

2022

Learning from Counterfactual Links for Link Prediction

ICML 2022spotlight

Learning to predict missing links is important for many graph-based applications. Existing methods were designed to learn the association between observed graph structure and existence of link between a pair of nodes. However, the causal relationship between the two variables was largely ignored for…

2021

Microsoft Speaker Diarization System for the Voxceleb Speaker Recognition Challenge 2020

ICASSP 2021accepted

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge (VoxSRC) 2020. We will first explain our system design to address issues in handling real multi-talker reco…

Cited by 0SourceScholar
2020

CP-GAN: Context Pyramid Generative Adversarial Network for Speech Enhancement

ICASSP 2020accepted

The topic of speech enhancement has been largely improved recently, especially with the development of generative adversarial networks (GANs). However prior methods simply follow the GAN architectures from computer vision tasks without specific designs for the speech enhancement according to the aud…

Cited by 0SourceScholar
2018

An Ensemble Framework of Voice-Based Emotion Recognition System for Films and TV Programs

ICASSP 2018accepted

Employing voice-based emotion recognition function in artificial intelligence (AI) product will improve the user experience. Most of researches that have been done only focus on the speech collected under controlled conditions. The scenarios evaluated in these research were well controlled. The conv…

Cited by 0SourceScholar
2016

Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach

ICASSP 2016accepted

Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used…

Cited by 20SourceScholar
2015

Weighted training for speech under Lombard Effect for speaker recognition

ICASSP 2015accepted

The presence of Lombard Effect in speech is proven to have severe effects on the performance of speech systems, especially speaker recognition. Varying kinds of Lombard speech are produced by speakers under influence of varying noise types [1]. This study proposes a high-accuracy classifier using de…

Cited by 9SourceScholar