← Search

Tao Jiang

52 accepted papers

2026

DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning

ICRA 2026poster

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to limited spatial reasoning inherited from Vision-Language Models (VLMs). Existing VL…

2026

EvoGM: Learning to Merge LLMs via Evolutionary Generative Optimization

ICML 2026poster

Evolutionary model merging provides a powerful framework for the automated, training-free composition of LLMs through parameter-space search. However, existing methods predominantly rely on stochastic, hand-crafted operators that overlook the underlying performance landscape of the coefficient space…

Cited by 0SourceScholar
2026

FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding

ICLR 2026poster

Autoregressive vision-language-action (VLA) models have recently demonstrated strong capabilities in robotic manipulation. However, their core process of action tokenization often involves a trade-off between reconstruction fidelity and inference efficiency. We introduce \textbf{FASTer}, a unified f…

Cited by 0SourceScholar
2026

FedSkeleton: Secure Multi-Party Graph Skeleton Construction for Privacy-Preserving Federated Time-Series Forecasting

AAAI 2026technical

In real-world time-series modelling, graph structures are widely adopted because they explicitly encode node topology and capture complex network dynamics. In practice, however, a complete graph is often partitioned across multiple parties; each party can access only its local sub-graph and, owing t

Cited by 0SourcePDFScholar
2026

Multi-agent In-context Coordination via Decentralized Memory Retrieval

AAAI 2026technical

Large transformer models, trained on diverse datasets, have demonstrated impressive few-shot performance on previously unseen tasks without requiring parameter updates. This capability has also been explored in Reinforcement Learning (RL), where agents interact with the environment to retrieve conte

Cited by 0SourcePDFScholar
2026

PACE: Proactive UAV-Assisted Collaborative Motion Planning for UGV Navigation Enhancement

RA-L 2026

Existing air–ground collaborative systems often rely on passive or reactive coordination, which can cause discontinuous motion and underuse the UAV's ability for anticipatory support. To address these limitations, we propose <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://

Cited by 0SourceScholar
2026

ReFAct: Empowering Multimodal Web Agents with Visual and Context Focusing

CVPR 2026

Multimodal Web Search Agents demonstrate a practically valuable capability by fusing information from diverse modalities (e.g., text and vision), retrieved iteratively from the internet, to address complex user queries. However, the visual modality is prone to information overload, and the noise con

Cited by 0SourceScholar
2026

SA-MPPI: Sensitivity-Aware Model Predictive Path Integral Control for Robust and Agile Quadrotor Flight

ICRA 2026poster

Reliable quadrotor control in dynamic environments remains challenging due to external disturbances and internal uncertainties. While Model Predictive Path Integral (MPPI) control enables agile maneuvers through samplingbased optimization, its performance often degrades under such unmodeled uncertai…

Cited by 0Scholar
2026

SIMPC: Learning Self-Induced Mirror-Point Consistency for Unsupervised Point Cloud Denoising

ICML 2026poster

In point clouds, noise directly perturbs point coordinates that encode both spatial location and geometry, making one-to-one correspondence construction more challenging than in images. Existing methods impose statistical mappings across noisy variants via noise or optimal transport, but suffer from…

Cited by 0SourceScholar
2026

UQ-ViT: Harmonizing Extreme Activations with Hardware-Friendly Uniform Quantization in Vision Transformers

AAAI 2026technical

Post-Training Quantization enables efficient Vision Transformer (ViTs) deployment with a small calibration data, and its prevalent use of uniform quantization harnesses AI accelerator matrix cores for high-speed inference. However, the application of uniform quantization is fundamentally challenged

Cited by 0SourcePDFScholar
2025

Combined Modal Robust Cascade Control for Wheeled Self-Reconfigurable Robots Under Drive Failure and Safety Threat

ICRA 2025

Wheeled self-reconfigurable robots (WSRRs), a new type of multi-robot system with flexible configurations and task adaptability, have an extensive application prospects in unstructured mission environments. In this paper, based on the nonholonomic constraints and Lagrange method, the combinatorial m

Cited by 0SourceScholar
2025

Design, Manufacturing, and Experiments of an Origami-based Parallel-Legged Structure for Insect-scale Robots

IROS 2025

Aiming to address the challenges associated with complex manufacturing processes and the difficulties in batch production of insect-scale robots. A mechatronic origami mechanism applied to an insect-scale parallel-legged structure is designed, manufactured, and tested. The origami mechanism is const

Cited by 0SourceScholar
2025

Efficient 7-DoF Grasp for Target-Driven Object in Dense Cluttered Scenes

ICRA 2025

Achieving a real-time precise grasp of a specified target object in densely cluttered environments is an essential capability for autonomous robot operation. Recently, considerable investigations on planar and spatial grasp have been carried out, and significant results have been obtained. However,

Cited by 2SourcecodeScholar
2025

Heteroscedastic Bayesian Optimization-Based Dynamic PID Tuning for Accurate and Robust UAV Trajectory Tracking

IROS 2025

Unmanned Aerial Vehicles (UAVs) play an important role in various applications, where precise trajectory tracking is crucial. However, conventional control algorithms for trajectory tracking often exhibit limited performance due to the underactuated, nonlinear, and highly coupled dynamics of quadrot

Cited by 1SourceScholar
2025

Hierarchical Trajectory Planning Method for Piano-Playing Robot

IROS 2025

Piano-playing tasks, which effectively demonstrate bimanual coordination capabilities in humanoid robots, are increasingly becoming a research focus. However, prior research has predominantly focused on Cartesian space trajectory planning without adequately addressing real-world obstacle avoidance c

Cited by 0SourceScholar
2025

LLM-Assisted Semantically Diverse Teammate Generation for Efficient Multi-agent Coordination

ICML 2025poster

Training with diverse teammates is the key for learning generalizable agents. Typical approaches aim to generate diverse teammates by utilizing techniques like randomization, designing regularization terms, or reducing policy compatibility, etc. However, such teammates lack semantic information, res…

2025

MIA-Tuner: Adapting Large Language Models as Pre-training Text Detector

AAAI 2025technical

The increasing parameters and expansive dataset of large lan- guage models (LLMs) highlight the urgent demand for a technical solution to audit the underlying privacy risks and copyright issues associated with LLMs. Existing studies have partially addressed this need through an exploration of the pr…

2025

MoBA: Mixture of Block Attention for Long-Context LLMs

NeurIPS 2025spotlight

Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches eit…

Cited by 0SourcecodeScholar
2025

Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-Function

NeurIPS 2025poster

Learning from multi-agent expert demonstrations, known as Multi-Agent Imitation Learning (MAIL), provides a promising approach to sequential decision-making. However, existing MAIL methods including Behavior Cloning (BC) and Adversarial Imitation Learning (AIL) face significant challenges: BC suffer…

Cited by 0SourcecodeScholar
2025

Precision Autonomous Landing of UAV on High-Speed Vehicles Based on Enhanced Gimbal Stabilization and Smooth Trajectory Generation

IROS 2025

This paper proposes a precision autonomous landing system for unmanned aerial vehicles (UAVs) targeting high-speed moving platforms. By integrating gimbal-based precise positioning, smooth trajectory generation, and dynamically robust control, the system addresses key challenges in high-speed landin

Cited by 0SourceScholar
2025

Prompt-SID: Learning Structural Representation Prompt via Latent Diffusion for Single Image Denoising

AAAI 2025technical

Many studies have concentrated on constructing supervised models utilizing paired datasets for image denoising, which proves to be expensive and time-consuming. Current self-supervised and unsupervised approaches typically rely on blind-spot networks or sub-image pairs sampling, resulting in pixel i…

2025

TrackOcc: Camera-Based 4D Panoptic Occupancy Tracking

ICRA 2025

Comprehensive and consistent dynamic scene understanding from camera input is essential for advanced autonomous systems. Traditional camera-based perception tasks like 3D object tracking and semantic occupancy prediction lack either spatial comprehensiveness or temporal consistency. In this work, we

Cited by 4SourcecodeScholar
2024

CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction

ECCV 2024poster

"Vision-based 3D occupancy prediction is significantly challenged by the inherent limitations of monocular vision in depth estimation. This paper introduces CVT-Occ, a novel approach that leverages temporal fusion through the geometric correspondence of voxels over time to improve the accuracy of 3D…

2024

ChatMusician: Understanding and Generating Music Intrinsically with LLM

ACL 2024findings

While LLMs demonstrate impressive capabilities in musical knowledge, we find that music reasoning is still an unsolved task.We introduce ChatMusician, an open-source large language model (LLM) that integrates intrinsic musical abilities. It is based on continual pre-training and finetuning LLaMA2 on…

2024

FreeKD: Knowledge Distillation via Semantic Frequency Prompt

CVPR 2024poster

Knowledge distillation (KD) has been applied to various tasks successfully and mainstream methods typically boost the student model via spatial imitation losses. However the consecutive downsamplings induced in the spatial domain of teacher model is a type of corruption hindering the student from an…

2024

Generating Stereophonic Music with Single-Stage Language Models

ICASSP 2024accepted

The recent success of audio language models (LMs) has revolutionized the field of neural music generation. Among all audio LM approaches, MusicGen has demonstrated the success of a single-stage LMs based music generation framework, without needing to train multiple LMs. Despite its promising perform…

Cited by 0SourceScholar
2024

Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts

ICASSP 2024accepted

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation cap…

Cited by 0SourceScholar
2024

Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration

NeurIPS 2024poster

Membership Inference Attacks (MIA) aim to infer whether a target data record has been utilized for model training or not. Existing MIAs designed for large language models (LLMs) can be bifurcated into two types: reference-free and reference-based attacks. Although reference-based attacks appear prom…

2024

Multi-Agent Domain Calibration with a Handful of Offline Data

NeurIPS 2024poster

The shift in dynamics results in significant performance degradation of policies trained in the source domain when deployed in a different target domain, posing a challenge for the practical application of reinforcement learning (RL) in real-world scenarios. Domain transfer methods aim to bridge thi…

Cited by 0SourcePDFScholar
2024

RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation

CVPR 2024poster

Real-time multi-person pose estimation presents significant challenges in balancing speed and precision. While two-stage top-down methods slow down as the number of people in the image increases existing one-stage methods often fail to simultaneously deliver high accuracy and real-time performance.…

2024

SCNet: Sparse Compression Network for Music Source Separation

ICASSP 2024accepted

Deep learning-based methods have made significant achievements in music source separation. However, obtaining good results while maintaining a low model complexity remains challenging in super wide-band music source separation. Previous works either overlook the differences in subbands or inadequate…

Cited by 0SourceScholar
2024

SSCBench: A Large-Scale 3D Semantic Scene Completion Benchmark for Autonomous Driving

IROS 2024

Monocular scene understanding is a foundational component of autonomous systems. Within the spectrum of monocular perception topics, one crucial and useful task for holistic 3D scene understanding is semantic scene completion (SSC), which jointly completes semantic information and geometric details

Cited by 90SourcecodeScholar
2024

Segmented Safety Docking Control for Mobile Self-Reconfigurable Robots

IROS 2024poster

Mobile self-reconfigurable robots (MSRRs), as a novel multi-robot system with flexible configurations and task adaptability, hold promising applications in unstructured task environments. However, existing autonomous docking strategies are primarily applied in laboratory settings and face numerous c…

Cited by 0SourceScholar
2024

Tracking Control with Uncertainty Smoothing Estimation under Aggressive Maneuvers of Aerial Vehicles

IROS 2024poster

Aggressive maneuvering is crucial for aerial vehicles to execute adversarial and penetration missions. However, this challenges the accurate tracking control of drones due to uncertainties induced by high-speed flight. Therefore, firstly, a highly dynamic tracking control framework is proposed to ac…

Cited by 0SourceScholar
2023

Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving

NeurIPS 2023poster

Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D occupancy prediction, which estimates the detailed occupanc…

2022

Acceleration of Federated Learning with Alleviated Forgetting in Local Training

ICLR 2022poster

Federated learning (FL) enables distributed optimization of machine learning models while protecting privacy by independently training local models on each client and then aggregating parameters on a central server, thereby producing an effective global model. Although a variety of FL algorithms hav…

2022

PRNet: Point-Range Fusion Network for Real-Time LiDAR Semantic Segmentation

IJCAI 2022poster

Accurate and real-time LiDAR semantic segmentation is necessary for advanced autonomous driving systems. To guarantee a fast inference speed, previous methods utilize the highly optimized 2D convolutions to extract features on the range view (RV), which is the most compact representation of the LiDA…

Cited by 2SourcePDFScholar
2022

User Scheduling Using Graph Neural Networks for Reconfigurable Intelligent Surface Assisted Multiuser Downlink Communications

ICASSP 2022accepted

Reconfigurable intelligent surface (RIS) is capable of intelligently manipulating the phases of the incident electromagnetic wave to improve the wireless propagation environment between the base station (BS) and the users. This paper addresses the joint user scheduling, RIS configuration, and BS bea…

Cited by 0SourceScholar
2021

Conquering Textureless with RF-referenced Monocular Vision for MAV State Estimation

ICRA 2021poster

The versatile nature of agile micro aerial vehicles (MAVs) poses fundamental challenges to the design of robust state estimation in various complex environments. Achieving high-quality performance in textureless scenes is one of the missing pieces in the puzzle. Previously proposed solutions either…

Cited by 5SourcecodeScholar
2021

LREN: Low-Rank Embedded Network for Sample-Free Hyperspectral Anomaly Detection

AAAI 2021technical

Hyperspectral anomaly detection (HAD) is a challenging task because it explores the intrinsic structure of complex high-dimensional signals without any samples at training time. Deep neural networks (DNNs) can dig out the underlying distribution of hyperspectral data but are limited by the labeling…

2021

Litesing: Towards Fast, Lightweight and Expressive Singing Voice Synthesis

ICASSP 2021accepted

LiteSing proposed in this paper is a high-quality singing voice synthesis (SVS) system, which is fast, lightweight and expressive. This model mainly stacks several non-autoregressive WaveNet blocks in the encoder and decoder under a generative adversarial architecture, predicts full conditions from…

Cited by 0SourceScholar
2021

Look Before You Act: Boosting Pseudo-LiDAR with Online Semantic Embedding

IROS 2021poster

Vision-based 3D object detection is a research focus in the field of autonomous driving system. While recently proposed pseudo-LiDAR is a promising solution, its performance is severely restricted by the image-based depth estimator, leading to a considerable performance gap against the LiDAR-based c…

Cited by 0SourceScholar
2020

Intra-class Feature Variation Distillation for Semantic Segmentation

ECCV 2020poster

Current state-of-the-art semantic segmentation methods usually require high computational resources for accurate segmentation. One promising way to achieve a good trade-off between segmentation accuracy and efficiency is knowledge distillation. In this paper, different from previous methods performi…

2020

Ordinal Learning for Emotion Recognition in Customer Service Calls

ICASSP 2020accepted

Approaches toward ordinal speech emotion recognition (SER) tasks are commonly based on the categorical classification algorithms, where the rank-order emotions are arbitrarily treated as independent categories. To employ the ordinal information between emotional ranks, we propose to model the ordina…

Cited by 0SourceScholar
2020

Reinforced Molecular Optimization with Neighborhood-Controlled Grammars

NeurIPS 2020poster

A major challenge in the pharmaceutical industry is to design novel molecules with specific desired properties, especially when the property evaluation is costly. Here, we propose MNCE-RL, a graph convolutional policy network for molecular optimization with molecular neighborhood-controlled embeddin…

2019

Automatic Singing Evaluation without Reference Melody Using Bi-dense Neural Network

ICASSP 2019accepted

Automatic singing evaluation without reference melody has long been a difficult problem. This paper aims to pilot a novel data driven approach to tackle this artistic problem. We constructed a large scale dataset and designed an innovative Bi-Dense neural network which can address this task efficien…

Cited by 0SourceScholar
2019

Layer-wise Deep Neural Network Pruning via Iteratively Reweighted Optimization

ICASSP 2019accepted

The huge number of parameters of deep neural network makes it difficult to deploy on embedded devices with limited hardware, computation, storage and energy resources. In this paper, we shall propose a log-sum minimization approach to prune a trained network layer by layer thereby improving the netw…

Cited by 0SourceScholar
2019

Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding

ICASSP 2019accepted

Under noisy environments, to achieve the robust performance of speaker recognition is still a challenging task. Motivated by the promising performance of multi-task training in a variety of image processing tasks, we explore the potential of multitask adversarial training for learning a noise-robust…

Cited by 0SourceScholar