← Search

Yuke Li

26 accepted papers

2026

SMM Transformer: Leveraging Spiking Neural Networks for Multimodal Tasks

ICML 2026poster

Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and a mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SN…

Cited by 0SourceScholar
2026

Spike-HTR: Spiking Neural Transformer for Handwritten Text Recognition

ICML 2026poster

Offline handwritten text recognition (HTR) is blank-dominated: task-relevant evidence lies in sparse ink strokes, yet mainstream recognizers still expend dense spatial compute and full-length width-axis token mixing across the canvas. Spiking neural networks (SNNs) promise activity-proportional comp…

Cited by 0SourceScholar
2026

StyleDistillation: A New Insight of Image Style Enables Personalized Aesthetic Manipulation

ICML 2026poster

Text-guided stylized image generation has yielded promising advances by leveraging the powerful capabilities of text-to-image diffusion models. However, the inherent coupling of style and content information within the reference image presents a significant challenge. To address this, we propose Sty…

Cited by 0SourceScholar
2025

BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR

ICASSP 2025accepted

Recently, the Mixture of Expert (MoE) architecture, such as LR-MoE, is often used to alleviate the impact of language confusion on the multilingual ASR (MASR) task. However, it still faces language confusion issues, especially in mismatched domain scenarios. In this paper, we decouple language confu…

Cited by 0SourceScholar
2025

Bridging the Modality Gap for Speech-image Retrieval with Text Supervision

ICASSP 2025accepted

In recent years, while the performance of speech-image retrieval has improved significantly, it still lags behind that of image-text retrieval. Leveraging the text modality to enhance speech-image retrieval remains a promising research direction. In this paper, we propose to leverage text supervisio…

Cited by 0SourceScholar
2025

Compact R-X-Y Stage and Dual-Finger Micromanipulator under Inverted Optical Microscope for Microassembly

IROS 2025

Microassembly plays an important role in fabricating complex structures with small basic components in industrial and biomedical fields. Inverted optical microscope could provide high-quality image feedback for microassembly with its continuously improving resolution. However, a compact stage capabl

Cited by 0SourceScholar
2025

Enhanced Rolling Motion of Magnetic Microparticles by Turning Interface Lubrication

IROS 2025

Micro-nano robots must break the symmetry of the flow field to generate net displacement in the low Reynolds number environment. The spherical micro-robots utilize the frictional forces generated through interaction with the surface. We designed a magnetic microroller robot powered by the rotating A

Cited by 0SourceScholar
2025

Magnetically Actuated Steerable Catheter with Redundant DoF for Cardiovascular Interventions

IROS 2025

A magnetically controlled catheter system is proposed to enhance the precision and safety of vascular interventions by reducing procedure time and radiation exposure. The system can also function as a support channel for guidewire deployment. A novel navigation approach is introduced, employing an e

Cited by 0SourceScholar
2025

On-Chip Dynamic Mechanical Characterization: from Cells to Nucleus

IROS 2025

Traditional single-cell mechanical characterization techniques (e.g., atomic force microscopy) often face limitations in throughput, require invasive labeling, or fail to replicate physiological microenvironments, impeding their clinical utility for rapid cancer cell analysis. To address these limit

Cited by 0SourceScholar
2025

You Think, You ACT: The New Task of Arbitrary Text to Motion Generation

ICCV 2025poster

Text to Motion aims to generate human motions from texts. Existing settings rely on limited Action Texts that include action labels (e.g., "walk, bend"), which limits flexibility and practicability in scenarios difficult to describe directly. This paper extends limited Action Texts to arbitrary ones…

2024

Differentiable Resolution Compression and Alignment for Efficient Video Classification and Retrieval

ICASSP 2024accepted

Optimizing video inference efficiency has become increasingly important with the growing demand for video analysis in various fields. Some existing methods achieve high efficiency by explicit discard of spatial or temporal information, which poses challenges in fast-changing and fine-grained scenari…

Cited by 0SourceScholar
2024

HPL-ViT: A Unified Perception Framework for Heterogeneous Parallel LiDARs in V2V

ICRA 2024poster

To develop the next generation of intelligent LiDARs, we propose a novel framework of parallel LiDARs and construct a hardware prototype in our experimental platform, DAWN (Digital Artificial World for Natural). It emphasizes the tight integration of physical and digital space in LiDAR systems, with…

Cited by 7SourceScholar
2024

HaltingVT: Adaptive Token Halting Transformer for Efficient Video Recognition

ICASSP 2024accepted

Action recognition in videos poses a challenge due to its high computational cost, especially for Joint Space-Time video transformers (Joint VT). Despite their effectiveness, the excessive number of tokens in such architectures significantly limits their efficiency. In this paper, we propose Halting…

Cited by 0SourceScholar
2024

LLCP: Learning Latent Causal Processes for Reasoning-based Video Question Answer

ICLR 2024poster

Current approaches to Video Question Answering (VideoQA) primarily focus on cross-modality matching, which is limited by the requirement for extensive data annotations and the insufficient capacity for causal reasoning (e.g. attributing accidents). To address these challenges, we introduce a causal…

Cited by 2SourcePDFScholar
2024

Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action Recognition

ICML 2024poster

Few-shot action recognition aims at quickly adapting a pre-trained model to the novel data with a distribution shift using only a limited number of samples. Key challenges include how to identify and leverage the transferable knowledge learned by the pre-trained model. We therefore propose CDTD, or…

Cited by 2SourcePDFScholar
2024

Split to Merge: Unifying Separated Modalities for Unsupervised Domain Adaptation

CVPR 2024poster

Large vision-language models (VLMs) like CLIP have demonstrated good zero-shot learning performance in the unsupervised domain adaptation task. Yet most transfer approaches for VLMs focus on either the language or visual branches overlooking the nuanced interplay between both modalities. In this wor…

2024

Textual Grounding for Open-vocabulary Visual Information Extraction in Layout-diversified Documents

ECCV 2024poster

"Current methodologies have achieved notable success in the closed-set visual information extraction (VIE) task, while the exploration into open-vocabulary settings is comparatively underdeveloped, which is practical for individual users in terms of inferring information across documents of diverse…

Cited by 1SourcePDFScholar
2023

Part-Aware Transformer for Generalizable Person Re-identification

ICCV 2023poster

Domain generalization person re-identification (DG ReID) aims to train a model on source domains and generalize well on unseen domains. Vision Transformer usually yields better generalization ability than common CNN networks under distribution shifts. However, Transformer-based ReID models inevitabl…

Cited by 77PDFcodeScholar
2023

Programable On-Chip Fabrication of Magnetic Soft Micro-Robot

IROS 2023poster

In the last decade, researchers have been trying to develop many microrobots that mimic the extraordinary abilities of bionts in complex environments. How to fabricate the biomimetic microrobot with satisfying deformability and complex shapes to realize desired precise motion is the key issue. In th…

Cited by 0SourceScholar
2022

Controlled Fabrication of Micro-Chain Robot Using Magnetically Guided Arraying Microfluidic Devices

IROS 2022poster

The magnetic microrobot has become a promising approach in many biomedical applications due to its small volume, flexible motion, and untethered micromachines. The micro-chain robot is one of the most popular magnetic microrobots. However, the uncontrollable magnetic moment direction and quantity of…

Cited by 0SourceScholar
2022

ELMA: Energy-Based Learning for Multi-Agent Activity Forecasting

AAAI 2022technical

This paper describes an energy-based learning method that predicts the activities of multiple agents simultaneously. It aims to forecast both upcoming actions and paths of all agents in a scene based on their past activities, which can be jointly formulated by a probabilistic model over time. Learni…

Cited by 6SourcePDFScholar
2022

On-Chip Automatic Trapping and Rotating for Zebrafish Embryo Injection

RA-L 2022

Zebrafish embryo injection is often required in biomedical research using zebrafish. In the injecting operation, trapping and rotating the zebrafish embryo to achieve a proper posture is essential for the high success rate. We proposed an on-chip platform capable of efficient and automatic trapping

Cited by 8SourceScholar