← Search

Yao Li

46 accepted papers

2026

Let’s Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts

AAAI 2026technical

Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, current ICoT methods still suffer from two major limitations: (1) Static Visual Thou

Cited by 0SourcePDFScholar
2026

PGCSPose: Physics-Constrained Generation and Causal Semantic Fusion for Robust In-Hand Pose Estimation

RA-L 2026

Accurate in-hand pose estimation is essential for dexterous robotic manipulation but remains fragile under severe visual occlusion (<inline-formula><tex-math notation="LaTeX">$>$</tex-math></inline-formula>50%) and intermittent tactile contact. Existing visuo-tactile fusion methods treat vision and

Cited by 0SourceScholar
2026

Quantum Lipschitz Bandits

AAAI 2026technical

The Lipschitz bandit is a key variant of stochastic bandit problems where the expected reward function satisfies a Lipschitz condition with respect to an arm metric space. With its wide-ranging practical applications, various Lipschitz bandit algorithms have been developed, achieving the optimal reg

Cited by 0SourcePDFScholar
2026

RaCFusion: Improving Camera-Based 3D Object Detection via Radar-Assisted Hierarchical Refinement

RA-L 2026

Cameras and radar sensors are complementary in 3D object detection in that cameras specialize in capturing an object's visual information while radar provides spatial information and velocity hints. Existing radar-camera fusion methods often employ a symmetrical architecture that processes inputs fr

Cited by 0SourcecodeScholar
2026

Single Index Bandits: Generalized Linear Contextual Bandits with Unknown Reward Functions

ICLR 2026poster

Generalized linear bandits have been extensively studied due to their broad applicability in real-world online decision-making problems. However, these methods typically assume that the expected reward function is known to the users, an assumption that is often unrealistic in practice. Misspecificat…

Cited by 0SourceScholar
2025

Autonomous Suturing Method for Robot-Assisted Minimally Invasive Surgery

IROS 2025

Robot-assisted minimally invasive surgery is widely used because of its superior postoperative recovery outcomes. However, the workload for surgeons remains high. The development of autonomous suturing capabilities in surgical robots is poised to significantly reduce surgeon workload. In this study,

Cited by 0SourceScholar
2025

CELLmap: Enhancing LiDAR SLAM Through Elastic and Lightweight Spherical Map Representation

ICRA 2025

SLAM is a fundamental capability of unmanned systems, with LiDAR-based SLAM gaining widespread adoption due to its high precision. Current SLAM systems can achieve centimeter-level accuracy within a short period. However, there are still several challenges when dealing with largescale mapping tasks

Cited by 2SourceScholar
2025

Cockroach's Turning Strategy Enhanced Hexapod Robot with Flexible Torso

IROS 2025

The design and control of hexapod robots have become an active research field due to the ability to achieve adaptive and stable multi-terrain locomotion. However, existing hexapod robots focus on the integration of flexible pitch joints to enhance their obstacle-crossing and slope-climbing abilities

Cited by 0SourceScholar
2025

Counterfactual Knowledge Maintenance for Unsupervised Domain Adaptation

IJCAI 2025

Traditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic

2025

Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation

EMNLP 2025

End-to-end speech-to-speech (S2S) dialogue systems have recently garnered increasing research attention for their lower latency and more natural integration of nonverbal cues such as emotion and speaker identity. However, these systems face key challenges, particularly in incorporating external know

2025

Perception Helps Planning: Facilitating Multi-Stage Lane-Level Integration via Double-Edge Structures

RA-L 2025

When planning for autonomous driving, it is crucial to consider essential traffic elements such as lanes, intersections, traffic regulations, and dynamic agents. However, they are often overlooked by the traditional end-to-end planning methods, likely leading to inefficiencies and non-compliance wit

Cited by 1SourceScholar
2025

Zero-shot Quantization for Large-kernels via Shape-based Distribution and Diversity Self-distillation

ICASSP 2025accepted

Zero-shot quantization (ZSQ) has emerged as an effective method to reduce model complexity and memory footprint without using original training data, thereby mitigating data privacy and security concerns during model deployment. Recently, Large-Kernel Convolutional Neural Networks (LKCNNs) have achi…

Cited by 0SourceScholar
2024

AdaDiff: Accelerating Diffusion Models through Step-Wise Adaptive Computation

ECCV 2024poster

"Diffusion models achieve great success in generating diverse and high-fidelity images, yet their widespread application, especially in real-time scenarios, is hampered by their inherently slow generation speed. The slow generation stems from the necessity of multi-step network inference. While some…

Cited by 3SourcePDFScholar
2024

CRPlace: Camera-Radar Fusion with BEV Representation for Place Recognition

IROS 2024poster

The integration of complementary characteristics from camera and radar data has emerged as an effective approach in 3D object detection. However, such fusion-based methods remain unexplored for place recognition, an equally important task for autonomous systems. Given that place recognition relies o…

Cited by 4SourceScholar
2024

CalibFormer: A Transformer-based Automatic LiDAR-Camera Calibration Network

ICRA 2024poster

The fusion of LiDARs and cameras has been increasingly adopted in autonomous driving for perception tasks. The performance of such fusion-based algorithms largely depends on the accuracy of sensor calibration, which is challenging due to the difficulty of identifying common features across different…

Cited by 14SourceScholar
2024

DALD: Improving Logits-based Detector without Logits from Black-box LLMs

NeurIPS 2024poster

The advent of Large Language Models (LLMs) has revolutionized text generation, producing outputs that closely mimic human writing. This blurring of lines between machine- and human-written text presents new challenges in distinguishing one from the other – a task further complicated by the frequent…

2024

FARFusion: A Practical Roadside Radar-Camera Fusion System for Far-Range Perception

RA-L 2024

Far-range perception through roadside sensors is crucial to the effectiveness of intelligent transportation systems. The main challenge of far-range perception is due to the difficulty of performing accurate object detection and tracking under far distances <italic xmlns:mml="http://www.w3.org/1998/

Cited by 24SourceScholar
2024

SciCode: A Research Coding Benchmark Curated by Scientists

NeurIPS 2024poster

Since language models (LMs) now outperform average humans on many challenging tasks, it is becoming increasingly difficult to develop challenging, high-quality, and realistic evaluations. We address this by examining LM capabilities to generate code for solving real scientific research problems. Inc…

Cited by 18SourcePDFScholar
2024

The Continuous Jump Control of a Locust-Inspired Robot With Omnidirectional Trajectory Adjustment

RA-L 2024

Jumping is an effective way for small robots to overcome obstacles. After years of development, many miniature jumping robots have been proposed with various mechanisms, and they have achieved jump trajectory control, fall recovery, and even continuous jumps. However, most miniature jumping robots d

Cited by 7SourceScholar
2023

Accelerating Dataset Distillation via Model Augmentation

CVPR 2023highlight

Dataset Distillation (DD), a newly emerging field, aims at generating much smaller but efficient synthetic training datasets from large ones. Existing DD methods based on gradient matching achieve leading performance; however, they are extremely computationally intensive as they require continuously…

2023

An Origami-Based Miniature Jumping Robot with Adjustable Jumping Trajectory and Enhanced Intermittent Jumps

IROS 2023poster

A small-scale jumping robot can reach obstacles much larger than its size. It is important for a jumping robot to perform intermittent jumps to cross through rough terrains. However, the limitations of conventional structures hinder the further integration of functions to a miniature (sub-50 g) jump…

Cited by 1SourceScholar
2023

Bi-LRFusion: Bi-Directional LiDAR-Radar Fusion for 3D Dynamic Object Detection

CVPR 2023poster

LiDAR and Radar are two complementary sensing approaches in that LiDAR specializes in capturing an object's 3D shape while Radar provides longer detection ranges as well as velocity hints. Though seemingly natural, how to efficiently combine them for improved feature representation is still unclear.…

2023

CluB: Cluster Meets BEV for LiDAR-Based 3D Object Detection

NeurIPS 2023poster

Currently, LiDAR-based 3D detectors are broadly categorized into two groups, namely, BEV-based detectors and cluster-based detectors. BEV-based detectors capture the contextual information from the Bird's Eye View (BEV) and fill their center voxels via feature diffusion with a stack of convolution l…

Cited by 6SourcePDFScholar
2023

DSR: Dynamical Surface Representation as Implicit Neural Networks for Protein

NeurIPS 2023poster

We propose a novel neural network-based approach to modeling protein dynamics using an implicit representation of a protein’s surface in 3D and time. Our method utilizes the zero-level set of signed distance functions (SDFs) to represent protein surfaces, enabling temporally and spatially continuous…

2023

Perturbation Towards Easy Samples Improves Targeted Adversarial Transferability

NeurIPS 2023poster

The transferability of adversarial perturbations provides an effective shortcut for black-box attacks. Targeted perturbations have greater practicality but are more difficult to transfer between models. In this paper, we experimentally and theoretically demonstrated that neural networks trained on t…

2023

You Need Multiple Exiting: Dynamic Early Exiting for Accelerating Unified Vision Language Model

CVPR 2023poster

Large-scale transformer models bring significant improvements for various downstream vision language tasks with a unified architecture. The performance improvements come with increasing model size, resulting in slow inference speed and increased cost for severing. While some certain predictions bene…

2022

${\mathsf{EZFusion}}$: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking

RA-L 2022

A recent trend is to combine multiple sensors ( <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i.e.</i> , cameras, LiDARs and millimeter-wave Radars) to achieve robust multi-modal perception for autonomous systems such as self-driving vehicles. Alth

Cited by 11SourceScholar
2022

ADDMU: Detection of Far-Boundary Adversarial Examples with Data and Model Uncertainty Estimation

EMNLP 2022main

Adversarial Examples Detection (AED) is a crucial defense technique against adversarial attacks and has drawn increasing attention from the Natural Language Processing (NLP) community. Despite the surge of new AED methods, our studies show that existing methods heavily rely on a shortcut to achieve…

2022

Design and Experimental Validation of a Shock-Absorption Mechanism Inspired From the Frog's Forelimbs

RA-L 2022

Frogs reveal superior shock-absorption ability, which is mainly brought by the forelimbs. Frogs touch the ground with their forelimbs first in landing, followed by the elbow joints’ compression. In this process, the muscles and tendons of the forelimbs are pulled to store energy. However, muscles an

Cited by 0SourceScholar
2022

The Feedback Trajectory Control of a SMA-Driven Miniature Jumping Robot

ICRA 2022poster

Jumping motion is an effective way to overcome large obstacles, especially for the miniature robots. However, controlling of the jumping trajectory on a centimeter scale robot is not easy due to the limitation of size and payload. None of the jumping robots lighter than 90 g achieved the feedback co…

Cited by 8SourceScholar
2021

Linear Convergent Decentralized Optimization with Compression

ICLR 2021poster

Communication compression has become a key strategy to speed up distributed optimization. However, existing decentralized algorithms with compression mainly focus on compressing DGD-type algorithms. They are unsatisfactory in terms of convergence rate, stability, and the capability to handle heterog…

Cited by 61SourcePDFScholar
2021

Towards Robustness of Deep Neural Networks via Regularization

ICCV 2021poster

Recent studies have demonstrated the vulnerability of deep neural networks against adversarial examples. Inspired by the observation that adversarial examples often lie outside the natural image data manifold and the intrinsic dimension of image data is much smaller than its pixel space dimension, w…

Cited by 18PDFcodeScholar
2020

A Double Residual Compression Algorithm for Efficient Distributed Learning

AISTATS 2020poster

Large-scale machine learning models are often trained by parallel stochastic gradient descent algorithms. However, the communication cost of gradient aggregation and model synchronization between the master and worker nodes becomes the major obstacle for efficient learning as the number of workers a…

Cited by 71SourcePDFScholar
2020

Privacy-Aware UAV Flights through Self-Configuring Motion Planning

ICRA 2020poster

During flights, an unmanned aerial vehicle (UAV) may not be allowed to move across certain areas due to soft constraints such as privacy restrictions. Current methods on self-adaption focus mostly on motion planning such that the trajectory does not trespass predetermined restricted areas. When the…

Cited by 16SourceScholar
2019

Adv-BNN: Improved Adversarial Defense through Robust Bayesian Neural Network

ICLR 2019poster

We present a new algorithm to train a robust neural network against adversarial attacks. Our algorithm is motivated by the following two ideas. First, although recent work has demonstrated that fusing randomness can improve the robustness of neural networks (Liu 2017), we noticed that adding noise…

2018

Learning from Group Comparisons: Exploiting Higher Order Interactions

NeurIPS 2018poster

We study the problem of learning from group comparisons, with applications in predicting outcomes of sports and online games. Most of the previous works in this area focus on learning individual effects---they assume each player has an underlying score, and the ''ability'' of the team is modeled by…

Cited by 28SourcePDFScholar
2017

Attend in Groups: A Weakly-Supervised Deep Learning Framework for Learning From Web Data

CVPR 2017poster

Large-scale datasets have driven the rapid development of deep neural networks for visual recognition. However, annotating a massive dataset is expensive and time-consuming. Web images and their labels are, in comparison, much easier to obtain, but direct training on such automatially harvested imag…

Cited by 103PDFScholar
2017

Sequential Person Recognition in Photo Albums With a Recurrent Network

CVPR 2017poster

Recognizing the identities of people in everyday photos is still a very challenging problem for machine vision, due to issues such as non-frontal faces, changes in clothing, location, lighting. Recent studies have shown that rich relational information between people in the same photo can help in re…

Cited by 32PDFScholar