← Search

TAO SUN

77 accepted papers

2026

APT: Affine Prototype-Timestamp for Time Series Forecasting Under Distribution Shift

AAAI 2026technical

Time series forecasting under distribution shift remains challenging, as existing deep learning models often rely on local statistical normalization (e.g., mean and variance) that fails to capture global distribution shift. Methods like RevIN and its variants attempt to decouple distribution and pat

Cited by 0SourcePDFScholar
2026

Do 3D Large Language Models Really Understand 3D Spatial Relationships?

ICLR 2026poster

Recent 3D Large-Language Models (3D-LLMs) claim to understand 3D worlds, especially spatial relationships among objects. Yet, we find that simply fine-tuning a language model on text-only question-answer pairs can perform comparably or even surpass these methods on the SQA3D benchmark without using…

Cited by 0SourceScholar
2026

From Diagrams to Code: Multilingual Programming with Visual Design

ICML 2026poster

In modern software development, particularly in emerging ``vibe coding'' paradigms, project implementation increasingly begins with visual interactions between users and AI coding assistants, where system architectures are communicated through visual designs before coding. This visual-first approach…

Cited by 0SourceScholar
2026

P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

ICLR 2026poster

Academic posters are vital for scholarly communication, yet their manual creation is time-consuming. However, automated academic poster generation faces significant challenges in preserving intricate scientific details and achieving effective visual-textual integration. Existing approaches often str…

Cited by 0SourcecodeScholar
2026

Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio

IJCAI 2026

The impressive performance of large language models (LLMs) arises from their massive scale and heterogeneous module composition. However, this structural heterogeneity poses significant optimization challenges. While adaptive optimizers such as Adam(W) provide per-parameter adaptivity, they do not e

Cited by 0Scholar
2026

Semantic and Terrain-Aware Trajectory Optimization for Uniform Coverage in Obstacle-Laden Environments

ICRA 2026poster

Achieving efficient and uniform coverage in obstacle-laden unknown environments is essential for au- tonomous robots in cleaning, inspection and agricultural op- erations. Unlike most existing approaches that prioritize path length and time optimality, we propose the SHIFT planner framework, which i…

Cited by 0Scholar
2025

A Synchronous-Optimized and Safety-Improved Framework for Human-Robot Interaction in Robot-Assisted Knee Arthroplasty

RA-L 2025

Shared control is the most commonly used physical human-robot interaction (pHRI) in robot-assisted orthopedic surgery. However, challenges persist in tasks such as robotassisted knee arthroplasty, particularly in terms of motion delays and inadequate three-dimensional (3D) constraints. In this lette

Cited by 0SourceScholar
2025

A sEMG-Based Active-Passive Fusion Rehabilitation Method for Ankle Fracture Rehabilitation Robot after Surgery

RA-L 2025

In this letter, an active-passive fusion rehabilitation training method for the ankle fracture rehabilitation robot after surgery is developed. A subject-independent continuous estimation model of ankle torque is proposed based on surface electromyography (sEMG), an online adaptive algorithm for pas

Cited by 7SourceScholar
2025

Adaptive Wall-Following Control for Unmanned Ground Vehicles Using Spiking Neural Networks

IROS 2025

Unmanned ground vehicles operating in complex environments must adaptively adjust to modeling uncertainties and external disturbances to perform tasks such as wall following and obstacle avoidance. This paper introduces an adaptive control approach based on spiking neural networks for wall fitting a

Cited by 0SourceScholar
2025

Backdooring Vision-Language Models with Out-Of-Distribution Data

ICLR 2025poster

The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to generate detailed text descriptions from visual inputs. Despite their growing importance, the security of VLMs, particularly against backdoor attack…

Cited by 3SourcePDFScholar
2025

Investigating the Role of Weight Decay in Enhancing Nonconvex SGD

CVPR 2025poster

Weight decay is a widely used technique in training machine learning models, known to empirically enhance the generalization of Stochastic Gradient Descent (SGD). While intuitively weight decay allows SGD to train a regularized model rather than the original one, there is limited theoretical underst…

Cited by 0SourcePDFScholar
2025

McEval: Massively Multilingual Code Evaluation

ICLR 2025poster

Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard to evaluate the capability of different LLMs in such tasks.…

2025

Meta-Analogy Learning Based on Dynamic Graph Neural Networks for Inductive Knowledge Graph Link Prediction

ICASSP 2025accepted

For inductive link prediction in knowledge graphs, we address the problem by considering bridging links (i.e., links connecting discrete graphs). Although current research overcomes the traditional graph topological constraints, it tends to ignore the dynamic interactions between relations and entit…

Cited by 0SourceScholar
2025

OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation

NeurIPS 2025poster

Large Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature. Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all co…

Cited by 0SourcecodeScholar
2025

On the Stability and Generalization of Meta-Learning: the Impact of Inner-Levels

NeurIPS 2025poster

Meta-learning has achieved significant advancements, with generalization emerging as a key metric for evaluating meta-learning algorithms. While recent studies have mainly focused on training strategies, data-split methods, and tightening generalization bounds, they often ignore the impact of inner-…

Cited by 0SourceScholar
2025

Rectified Point Flow: Generic Point Cloud Pose Estimation

NeurIPS 2025spotlight

We present Rectified Point Flow, a unified parameterization that formulates pairwise point cloud registration and multi-part shape assembly as a single conditional generative problem. Given unposed point clouds, our method learns a continuous point-wise velocity field that transports noisy points to…

Cited by 0SourcecodeScholar
2025

Reward-Augmented Data Enhances Direct Preference Alignment of LLMs

ICML 2025poster

Preference alignment in Large Language Models (LLMs) has significantly improved their ability to adhere to human instructions and intentions. However, existing direct alignment algorithms primarily focus on relative preferences and often overlook the qualitative aspects of responses, despite having…

2025

SMARTraj$^2$: A Stable Multi-City Adaptive Method for Multi-View Spatio-Temporal Trajectory Representation Learning

NeurIPS 2025poster

Spatio-temporal trajectory representation learning plays a crucial role in various urban applications such as transportation systems, urban planning, and environmental monitoring. Existing methods can be divided into single-view and multi-view approaches, with the latter offering richer representati…

Cited by 0SourcecodeScholar
2025

Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks

ICASSP 2025accepted

Sharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated with the regularization to characterize the sharpness of the model. Despite its success, research on adaptive regularizat…

Cited by 0SourceScholar
2025

Targeted Low-rank Refinement: Enhancing Sparse Language Models with Precision

ICML 2025poster

Pruning is a widely used technique for compressing large neural networks that eliminates weights that have minimal impact on the model's performance. Current pruning methods, exemplified by magnitude pruning, assign an importance score to each weight based on its magnitude and remove weights with sc…

Cited by 0SourcePDFScholar
2025

Underwater Motions Analysis and Control of a Coupling-Tiltable Unmanned Aerial-Aquatic Vehicle

ICRA 2025

Coupling-Tiltable Unmanned Aerial-Aquatic Vehicles (UAAVs) have gained increasing importance, yet lack comprehensive analysis and suitable controllers. This paper analyzes the underwater motion characteristics of a self-designed UAAV, Mirs-Alioth, and designs a controller for it. The effectiveness o

Cited by 1SourceScholar
2025

WoMAP: World Models For Embodied Open-Vocabulary Object Localization

CoRL 2025poster

Active object localization remains a critical challenge for robots, requiring efficient exploration of partially observable environments. However, state-of-the-art robot policies either struggle to generalize beyond demonstration datasets (e.g., imitation learning methods) or fail to generate physic…

Cited by 0SourceScholar
2025

XCOT: Cross-lingual Instruction Tuning for Cross-lingual Chain-of-Thought Reasoning

AAAI 2025technical

Chain-of-thought (CoT) has emerged as a powerful technique to elicit reasoning in large language models and improve a variety of downstream tasks. CoT mainly demonstrates excellent performance in English, but its usage in low-resource languages is constrained due to poor language generalization. To…

Cited by 39SourcePDFScholar
2025

XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser

COLING 2025main

In the domain of Document AI, parsing semi-structured image form is a crucial Key Information Extraction (KIE) task. The advent of pre-trained multimodal models significantly empowers Document AI frameworks to extract key information from form documents in different formats such as PDF, Word, and im…

2024

$\mathcal{B}$-Coder: Value-Based Deep Reinforcement Learning for Program Synthesis

ICLR 2024spotlight

Program synthesis aims to create accurate, executable programs from problem specifications, specifically from natural language descriptions in our context. Recent studies have leveraged the power of reinforcement learning (RL) in conjunction with large language models (LLMs), significantly enhancin…

Cited by 2SourcePDFScholar
2024

Customizable Combination of Parameter-Efficient Modules for Multi-Task Learning

ICLR 2024poster

Modular and composable transfer learning is an emerging direction in the field of Parameter Efficient Fine-Tuning, as it enables neural networks to better organize various aspects of knowledge, leading to improved cross-task generalization. In this paper, we introduce a novel approach Customized Pol…

Cited by 7SourcePDFScholar
2024

Exploring the Inefficiency of Heavy Ball as Momentum Parameter Approaches 1

IJCAI 2024poster

The heavy ball momentum method is a commonly used technique for accelerating training processes in the machine learning community. However, empirical evidence suggests that the convergence of stochastic gradient descent (SGD) with heavy ball may slow down when the momentum hyperparameter approaches…

Cited by 0SourcePDFScholar
2024

MV-ROPE: Multi-view Constraints for Robust Category-level Object Pose and Size Estimation

IROS 2024poster

Recently there has been a growing interest in category-level object pose and size estimation, and prevailing methods commonly rely on single view RGB-D images. However, one disadvantage of such methods is that they require accurate depth maps which cannot be produced by consumer-grade sensors. Furth…

Cited by 2SourceScholar
2024

RoleAgent: Building, Interacting, and Benchmarking High-quality Role-Playing Agents from Scripts

NeurIPS 2024poster

Believable agents can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication. Recently, generative agents have been proposed to simulate believable human behavior by using Large Language Models. However, the existing method heavily re…

Cited by 1SourcePDFScholar
2024

Stability and Generalization for Stochastic Recursive Momentum-based Algorithms for (Strongly-)Convex One to $K$-Level Stochastic Optimizations

ICML 2024poster

STOchastic Recursive Momentum (STORM)-based algorithms have been widely developed to solve one to $K$-level ($K \geq 3$) stochastic optimization problems. Specifically, they use estimators to mitigate the biased gradient issue and achieve near-optimal convergence results. However, there is relativel…

Cited by 0SourcePDFScholar
2024

Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and Smoothness

NeurIPS 2024poster

Asynchronous stochastic gradient descent (ASGD) has evolved into an indispensable optimization algorithm for training modern large-scale distributed machine learning tasks. Therefore, it is imperative to explore the generalization performance of the ASGD algorithm. However, the existing results are…

Cited by 4SourcePDFScholar
2024

UniCoder: Scaling Code Large Language Model via Universal Code

ACL 2024long

Intermediate reasoning or acting steps have successfully improved large language models (LLMs) for handling various downstream natural language processing (NLP) tasks.When applying LLMs for code generation, recent works mainly focus on directing the models to articulate intermediate natural-language…

2023

A Modular Biological Neural Network-Based Neuro-Robotic System via Local Chemical Stimulation and Calcium Imaging

RA-L 2023

Embodying in vitro biological neural networks (BNNs) with robots to explore the rise of intelligence in these simpler models and to endow robots with biological intelligence has been attracting increasing attention in the fields of neuroscience and robotics. However, current research suffers from un

Cited by 11SourceScholar
2023

A Small-Scale Untethered Tensegrity Robot With High Velocity and Multi Locomotion Modes

RA-L 2023

Tensegrity mobile robots are well-appraised for their high stiffness-to-mass ratio and superior structural compliance. However, traditional untethered tensegrity mobile robots usually have low velocity due to the large actuation force required by coupling effects among stiff struts and soft cables.

Cited by 15SourceScholar
2023

Domain Adaptation with Adversarial Training on Penultimate Activations

AAAI 2023technical

Enhancing model prediction confidence on target data is an important objective in Unsupervised Domain Adaptation (UDA). In this paper, we explore adversarial training on penultimate activations, i.e., input features of the final linear classification layer. We show that this strategy is more efficie…

2023

Magnetically-Assisted Microfluidic Printing for the Fabrication of Anisotropic Skeletal Muscle Structure

RA-L 2023

Microfluidic printing provides a novel tool to facilitate the bulk assembly of cell-aligned microfibers for the fabrication of artificial skeletal muscle structure. However, due to the poor controllability for the deposition position of the microfiber, it is still difficult to realize the anisotropi

Cited by 2SourceScholar
2023

Scale Jump-Aware Pose Graph Relaxation for Monocular SLAM with Re-Initializations

IROS 2023poster

Pose graph relaxation has become an indispensable addition to SLAM enabling efficient global registration of sensor reference frames under the objective of satisfying pair-wise relative transformation constraints. The latter may be given by incremental motion estimation or global place recognition.…

Cited by 0SourceScholar
2023

Stability-Based Generalization Analysis of the Asynchronous Decentralized SGD

AAAI 2023technical

The generalization ability often determines the success of machine learning algorithms in practice. Therefore, it is of great theoretical and practical importance to understand and bound the generalization error of machine learning algorithms. In this paper, we provide the first generalization resul…

Cited by 19SourcePDFScholar
2022

Finite-Time Analysis of Adaptive Temporal Difference Learning with Deep Neural Networks

NeurIPS 2022accept

Temporal difference (TD) learning with function approximations (linear functions or neural networks) has achieved remarkable empirical success, giving impetus to the development of finite-time analysis. As an accelerated version of TD, the adaptive TD has been proposed and proved to enjoy finite-tim…

Cited by 6SourcePDFScholar
2022

On the Practicality of Deterministic Epistemic Uncertainty

ICML 2022spotlight

A set of novel approaches for estimating epistemic uncertainty in deep neural networks with a single forward pass has recently emerged as a valid alternative to Bayesian Neural Networks. On the premise of informative representations, these deterministic uncertainty methods (DUMs) achieve strong perf…

2022

SHIFT: A Synthetic Driving Dataset for Continuous Multi-Task Domain Adaptation

CVPR 2022poster

Adapting to a continuously evolving environment is a safety-critical challenge inevitably faced by all autonomous-driving systems. Existing image- and video-based driving datasets, however, fall short of capturing the mutable nature of the real world. In this paper, we introduce the largest syntheti…

Cited by 166PDFScholar
2022

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

IJCAI 2022poster

We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introd…

2021

Inertial Proximal Deep Learning Alternating Minimization for Efficient Neutral Network Training

ICASSP 2021accepted

In recent years, the Deep Learning Alternating Minimization (DLAM), which is actually the alternating minimization applied to the penalty form of the deep neutral networks training, has been developed as an alternative algorithm to overcome several drawbacks of Stochastic Gradient Descent (SGD) algo…

Cited by 0SourceScholar
2021

Micro Robotic Manipulation System for the Force Stimulation of Muscle Fiber-like Cell Structure

ICRA 2021poster

Many previous works have facilitated muscle cell (C2C12) alignment to form fiber-like cell structures. However, there still remains a challenge how to induce C2C12 myoblasts in the cell structures to differentiate into matured myocytes to form a functional muscle tissue, while external mechanical st…

Cited by 2SourceScholar
2021

REPAINT: Knowledge Transfer in Deep Reinforcement Learning

ICML 2021spotlight

Accelerating learning processes for complex tasks by leveraging previously learned tasks has been one of the most challenging problems in reinforcement learning, especially when the similarity between source and target tasks is low. This work proposes REPresentation And INstance Transfer (REPAINT) a…

Cited by 32SourcePDFScholar
2020

DeepRacer: Autonomous Racing Platform for Experimentation with Sim2Real Reinforcement Learning

ICRA 2020poster

DeepRacer is a platform for end-to-end experimentation with RL and can be used to systematically investigate the key challenges in developing intelligent control systems. Using the platform, we demonstrate how a 1/18th scale car can learn to drive autonomously using RL with a monocular camera. It is…

Cited by 76SourceScholar
2020

Magnetically Actuated Pick-and-place Operations of Cellular Micro-rings for High-speed Assembly of Micro-scale Biological Tube

IROS 2020poster

Tissue engineering is trying to use modular tissue micro-rings to construct artificial biological microtubes as substitute of autologous tissue tubes to alleviate the shortage of donor sources. However, because of the lack of effective assembly strategies, it is still challenging to achieve high-spe…

Cited by 0SourceScholar
2020

Robust Multi-Agent Reinforcement Learning with Model Uncertainty

NeurIPS 2020poster

In this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward func…

Cited by 111SourcePDFScholar
2019

Automated Sorting of Rare Cells Based on Autofocusing Visual Feedback in Fluorescence Microscopy

IROS 2019poster

The research on rare cells makes a significant contribution to biology research and medical treatment for the application of diagnostic operation as well as prognoses treatment. Therefore, sorting them from heterogeneous mixtures is crucial and valuable. Traditional cell sorting methods featured wit…

Cited by 6SourceScholar
2019

General Proximal Incremental Aggregated Gradient Algorithms: Better and Novel Results under General Scheme

NeurIPS 2019poster

The incremental aggregated gradient algorithm is popular in network optimization and machine learning research. However, the current convergence results require the objective function to be strongly convex. And the existing convergence rates are also limited to linear convergence. Due to the mathema…

Cited by 19SourcePDFScholar
2019

Iteratively Reweighted Penalty Alternating Minimization Methods with Continuation for Image Deblurring

ICASSP 2019accepted

In this paper, we consider a class of nonconvex problems with linear constraints appearing frequently in the area of image processing. We solve this problem by the penalty method and propose the iteratively reweighted alternating minimization algorithm. To speed up the algorithm, we also apply the c…

Cited by 0SourceScholar
2019

Leveraging Crowdsourced GPS Data for Road Extraction From Aerial Imagery

CVPR 2019poster

Deep learning is revolutionizing the mapping industry. Under lightweight human curation, computer has generated almost half of the roads in Thailand on Open- StreetMap (OSM) using high resolution aerial imagery. Bing maps are displaying 125 million computer generated building polygons in the U.S. Wh…

Cited by 119PDFScholar
2018

Design and Online Calibration of a Highly Compact Microgripper

ICRA 2018poster

Microgrippers play a significant role in manipulation of micro-objects. To achieve dexterous and precise manipulation, a microgripper is required to be compactly designed and embedded with sensing feedback. Meanwhile, to convert the sensor position into displacement of the microgripper, the embedded…

Cited by 0SourceScholar
2018

LAG: Lazily Aggregated Gradient for Communication-Efficient Distributed Learning

NeurIPS 2018spotlight

This paper presents a new class of gradient methods for distributed machine learning that adaptively skip the gradient calculations to learn with reduced communication and computation. Simple rules are designed to detect slowly-varying gradients and, therefore, trigger the reuse of outdated grad…

Cited by 381SourcePDFScholar
2017

Differentially Private Learning of Undirected Graphical Models Using Collective Graphical Models

ICML 2017poster

We investigate the problem of learning discrete graphical models in a differentially private way. Approaches to this problem range from privileged algorithms that conduct learning completely behind the privacy barrier to schemes that release private summary statistics paired with algorithms to learn…

Cited by 39SourcePDFScholar
2017

Non-contact transportation and rotation of micro objects by vibrating glass needle circularly under water

ICRA 2017poster

In micromanipulation, lots of methods have been developed to manipulate objects in microscale. However, few of them can be applied in both the transportation and the rotation of the micro objects. In this paper, we present a novel method to realize the non-contact transportation and rotation of the…

Cited by 4SourceScholar
2017

Robotics-based micro-reeling of magnetic microfibers to fabricate helical structure for smooth muscle cells culture

ICRA 2017poster

Helical structure assembled by hydrogel microfibers is significant for culture of smooth muscle cells. However, the helical structure is only fabricated at the macroscale, while the fabrication of helical microstructure is still a challenge due to the lack of assembly method. In this paper, we propo…

Cited by 2SourceScholar
2016

High-Speed Bioassembly of Cellular Microstructures With Force Characterization for Repeating Single-Step Contact Manipulation

RA-L 2016

In vitro tissues are significant biological substitute for drug test, cell morphogenesis exploration and organ transplantation. In this letter, a novel microrobotic bioassembly method is proposed to engineer 3-D cellular structure, which can be utilized to culture in vitro tissue with microstructura

Cited by 2SourceScholar
2016

Microbubbles for High-Speed Assembly of Cell-Laden Vascular-Like Microtube

RA-L 2016

Vascular-like microtube takes an important role in delivering oxygen and nutrient to keep cell alive in the generated tissue. In this paper, we present an automated micromanipulation system to assemble 2-D gel micro-rings to vascular-like microtubes by means of generating and controlling microbubble

Cited by 1SourceScholar
2016

Micromanipulation for Coiling Microfluidic Spun Alginate Microfibers by Magnetically Guided System

RA-L 2016

Alginate hydrogel microfibers are a promising cell-laden module for three-dimensional (3-D) assembly to build cellular structures. However, it is still a challenge to manipulate them for microassembly. In this letter, we report a novel magnetic control method to handle this challenge. To enhance the

Cited by 9SourceScholar
2015

Automated bubble-based assembly of cell-laden microgels into vascular-like microtubes

IROS 2015poster

Fabrication of artificial blood vessels in micro scale significantly benefits the regeneration of functional human vascular networks. In this paper, we develop an efficient multi-microrobotic system with an innovative motorized sample holder (MSH) and two manipulators. Air is injected into the solut…

Cited by 3SourceScholar
2015

Three-dimensional magnetic assembly of alginate microfibers using microfluidic “printing” method

ICRA 2015poster

Due to the poor controllability in hydrogels, Hydrogels-based assembly to form larger 3D complex shapes is still a big challenge. In this paper, we have reported a novel “bottom-up” method to fabricate three-dimensional (3D) magnetic alginate microfibers (MAMs) assemblies with complex shapes. Specif…

Cited by 2SourceScholar