← Search

Guang Li

28 accepted papers

2026

ASMIL: Attention-Stabilized Multiple Instance Learning for Whole-Slide Imaging

ICLR 2026poster

Attention-based multiple instance learning (MIL) has emerged as a powerful framework for whole slide image (WSI) diagnosis, leveraging attention to aggregate instance-level features into bag-level predictions. Despite this success, we find that such methods exhibit a new failure mode: unstable atte…

Cited by 0SourcecodeScholar
2026

EVLF: Early Vision-Language Fusion for Generative Dataset Distillation

CVPR 2026

Dataset distillation (DD) aims to synthesize compact training sets that enable models to achieve high accuracy with significantly fewer samples. Recent diffusion-based DD methods commonly introduce semantic guidance through late-stage cross-attention, where textual prompts tend to dominate the gener

Cited by 0SourcecodeScholar
2026

EnhanceERASOR: Two-Stage Static 3D Point Cloud Mapping in Dynamic Scenes

ICRA 2026poster

A clean map of the surrounding environment is essential for autonomous driving systems to ensure reliable localization and safe path planning. However, the existence of dynamic objects introduces ghost traces into the map, significantly degrading its quality. To address this issue, we propose Enhanc…

Cited by 0Scholar
2026

Fuzzy Fusion Control Strategy With Efficient Deep Deterministic Policy Gradient for Robotic Peg-in-Hole Assembly

RA-L 2026

The robotic peg-in-hole assembly task remains challenging. Traditional force control methods struggle with complex parameter identification and contact state analysis, while deep reinforcement learning(DRL) suffers from low efficiency and poor adaptability. To address these shortcomings and to capit

Cited by 0SourceScholar
2026

Learning Goal-Directed Rolling: Spherical Robot Point-to-Point Control Through Reinforcement Learning

RA-L 2026

Point-to-point navigation is an important ability for spherical robots. Traditional methods usually use a planner and a tracker for short-range target control. However, this hierarchical method suffers from a mismatch issue. In this work, we propose an end-to-end controller based on reinforcement le

Cited by 0SourceScholar
2026

MeanFuser: Fast One-Step Multi-Modal Trajectory Generation and Adaptive Reconstruction via MeanFlow for End-to-End Autonomous Driving

CVPR 2026

Generative models have shown great potential in trajectory planning. Recent studies demonstrate that anchor-guided generative models are effective in modeling the uncertainty of driving behaviors and improving overall performance. However, these methods rely on discrete anchor vocabularies that must

Cited by 0SourcecodeScholar
2026

Otter: Mitigating Background Distractions of Wide-Angle Few-Shot Action Recognition with Enhanced RWKV

AAAI 2026technical

Wide-angle videos in few-shot action recognition (FSAR) effectively express actions within specific scenarios. However, without a global understanding of both subjects and background, recognizing actions in such samples remains challenging because of the background distractions. Receptance Weighted

Cited by 0SourcePDFScholar
2026

PerlAD: Towards Enhanced Closed-Loop End-to-End Autonomous Driving With Pseudo-Simulation-Based Reinforcement Learning

RA-L 2026

End-to-end autonomous driving policies based on Imitation Learning (IL) often struggle in closed-loop execution due to the misalignment between inadequate open-loop training objectives and real driving requirements. While Reinforcement Learning (RL) offers a solution by directly optimizing driving g

Cited by 1SourceScholar
2026

SimScale: Learning to Drive via Real-World Simulation at Scale

CVPR 2026

Achieving fully autonomous driving systems requires learning rational decisions in a wide span of scenarios, including safety-critical and out-of-distribution ones. However, such cases are underrepresented in real-world corpus collected by human experts. To complement for the lack of data diversity,

Cited by 0SourcecodeScholar
2026

VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

CVPR 2026

The significance of cross-view 3D geometric modeling capabilities for autonomous driving is self-evident, yet existing Vision-Language Models (VLMs) inherently lack this capability, resulting in their mediocre performance. While some promising approaches attempt to mitigate this by constructing Q&A

Cited by 0SourcecodeScholar
2025

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

EMNLP 2025

Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce AgentThink , a pioneering unified framewor

2025

An Online System Identification Algorithm for Spherical Robot Using the Koopman Theory

RA-L 2025

This letter proposes a novel linear online identification framework for the spherical robot to address the modeling difficulties posed by nonlinearity and time-varying characteristics. Firstly, the Koopman theory is applied to the spherical robot to build a linear model to approximate the nonlineari

Cited by 4SourceScholar
2025

Continual Self-supervised Learning Considering Medical Domain Knowledge in Chest CT Images

ICASSP 2025accepted

We propose a novel continual self-supervised learning method (CSSL) considering medical domain knowledge in chest CT images. Our approach addresses the challenge of sequential learning by effectively capturing the relationship between previously learned knowledge and new information at different sta…

Cited by 0SourceScholar
2025

Dataset Distillation via Vision-Language Category Prototype

ICCV 2025poster

Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational consumption. However, previous DD methods mainly focus on distilling information from i…

2025

Generative Dataset Distillation Based on Self-knowledge Distillation

ICASSP 2025accepted

Dataset distillation is an effective technique for reducing the cost and complexity of model training while maintaining performance by compressing large datasets into smaller, more efficient versions. In this paper, we present a novel generative dataset distillation method that can improve the accur…

Cited by 0SourceScholar
2025

Kinematic Model and Trajectory Tracking Algorithm for High-Speed Spherical Robots

IROS 2025

This paper proposes a new turning theory for spherical robots, which better describes the turning mechanism of spherical robots under turning constraints, using a pendulum-driven spherical robot as an example. Compared to the previous turning theory, the new theory shows greater alignment with real-

Cited by 0SourceScholar
2025

Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence

AAAI 2025technical

In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but dire…

2024

Chained Flexible Capsule Endoscope: Unraveling the Conundrum of Size Limitations and Functional Integration for Gastrointestinal Transitivity

ICRA 2024poster

Capsule endoscopes, predominantly serving diagnostic functions, provide lucid internal imagery but are devoid of surgical or therapeutic capabilities. Consequently, despite lesion detection, physicians frequently resort to traditional endoscopic or open surgical procedures for treatment, resulting i…

Cited by 1SourceScholar
2024

SOMTP: A Self-Supervised Learning-Based Optimizer for MPC-Based Safe Trajectory Planning Problems in Robotics

RA-L 2024

Model Predictive Control (MPC)-based trajectory planning has been widely used in robotics, and incorporating Control Barrier Function (CBF) constraints into MPC can greatly improve its obstacle avoidance efficiency. Unfortunately, traditional optimizers are resource-consuming and slow to solve such

Cited by 4SourceScholar
2023

Bimodal Fusion Network for Basic Taste Sensation Recognition from Electroencephalography and Electromyography

ICASSP 2023accepted

Taste sensation can be objectively measured using electroencephalography (EEG) or electromyography (EMG). How-ever, it is still challenging to effectively utilize the complementary information from EEG and EMG signals in taste sensation recognition. This paper proposes a bimodal fusion network (Bi-F…

Cited by 0SourceScholar
2022

A Robust Reference Path Selection Method for Path Planning Algorithm

RA-L 2022

In this letter, a general robust reference path selection method (RPSM) that can be integrated into current existing motion planning algorithms is proposed to improve the mobile performance of autonomous patrol robots. The proposed RPSM maintains a dynamic array of path candidates that contains newl

Cited by 14SourceScholar
2022

Direction and Trajectory Tracking Control for Nonholonomic Spherical Robot by Combining Sliding Mode Controller and Model Prediction Controller

RA-L 2022

A spherical robot is a nonlinear, nonholonomic, and unstable system which increases the difficulty of the direction and trajectory tracking problem. In this study, we propose a new direction controller Hierarchical Terminal Sliding Mode Controller (HTSMC), an instruction planning controller called M

Cited by 33SourceScholar
2022

Multi-Terrain Velocity Control of the Spherical Robot by Online Obtaining the Uncertainties in the Dynamics

RA-L 2022

One controller cannot work on multiple and unknown terrains in the velocity control of the spherical robot, because the dynamic models of the robot vary on different terrains, and unmodeled dynamics and uncertainties exist in estimated dynamic models. Based on the above problem, a new velocity contr

Cited by 24SourceScholar
2022

Self-Knowledge Distillation based Self-Supervised Learning for Covid-19 Detection from Chest X-Ray Images

ICASSP 2022accepted

The global outbreak of the Coronavirus 2019 (COVID-19) has overloaded worldwide healthcare systems. Computer-aided diagnosis for COVID-19 fast detection and patient triage is becoming critical. This paper proposes a novel self-knowledge distillation based self-supervised learning method for COVID-19…

Cited by 0SourceScholar
2021

Fuzzy PID Controller Based on Yaw Angle Prediction of a Spherical Robot

IROS 2021poster

In this paper, a fuzzy PID controller based on yaw angle prediction is applied to design an attitude controller for a spherical rolling robot. The robot consists of a 2-DOF pendulum located inside a spherical shell with freedom to rotate about the transversal and longitudinal axis. The proposed cont…

Cited by 24SourceScholar