← Search

Xing Li

28 accepted papers

2026

Adversarial Latent Embedding Repair for LLM Continual Learning

ICML 2026poster

Research on continual learning for LLMs seeks to acquire new skills without catastrophic forgetting of established prior knowledge. However, domain-specific fine-tuning still triggers severe, long-tailed forgetting issues even under narrow updates, particularly when the pre-training data is inaccess…

Cited by 0SourceScholar
2026

Beyond Speedup - Utilizing KV Cache for Sampling and Reasoning

ICLR 2026poster

KV caches, typically used only to speed up autoregressive decoding, encode contextual information that can be reused for downstream tasks at no extra cost. We propose treating the KV cache as a lightweight representation, eliminating the need to recompute or store full hidden states. Despite being w…

Cited by 0SourceScholar
2026

How Powerful are LLMs in Generating Program Specifications?

ICML 2026poster

Formal verification provides strong guarantees of software correctness, but its adoption is limited by the high cost of writing precise formal specifications. While recent large language models (LLMs) have demonstrated impressive capabilities in theorem proving and verified code generation, how powe…

Cited by 0SourceScholar
2026

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

ICLR 2026poster

Test-time scaling has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs) by allocating additional computational resources during inference. However, this paradigm is inherently inefficient due to the generation of redundant and repetitive reasonin…

Cited by 0SourcecodeScholar
2026

TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time Scaling

ICLR 2026poster

Large Reasoning Models (LRMs) demonstrate exceptional capability in tackling complex mathematical, logical, and coding tasks by leveraging extended Chain-of-Thought (CoT) reasoning. Test-time scaling methods—such as prolonging CoT with explicit token-level exploration—can push LRMs’ accuracy boundar…

Cited by 0SourceScholar
2026

Why Attention Patterns Exist: A Unifying Temporal Perspective Analysis

ICLR 2026poster

Attention patterns play a crucial role in both training and inference of large language models (LLMs). Prior works have identified individual patterns—such as retrieval heads, sink heads, and diagonal traces—but these observations remain fragmented and lack a unifying explanation. To bridge this gap…

Cited by 0SourcecodeScholar
2025

Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM Inference

NeurIPS 2025poster

Key-Value (KV) cache eviction---which retains the KV pairs of the most important tokens while discarding less important ones---is a critical technique for optimizing both memory usage and inference latency in large language models (LLMs). However, existing approaches often rely on simple heuristics-…

Cited by 0SourceScholar
2025

AttentionPredictor: Temporal Patterns Matter for KV Cache Compression

NeurIPS 2025poster

With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention s…

Cited by 0SourcecodeScholar
2025

Circuit Transformer: A Transformer That Preserves Logical Equivalence

ICLR 2025poster

Implementing Boolean functions with circuits consisting of logic gates is fundamental in digital computer design. However, the implemented circuit must be exactly equivalent, which hinders generative neural approaches on this task due to their occasionally wrong predictions. In this study, we introd…

2025

Joint-Wise Distributed Perception Graph Convolutional Network for Skeleton-Based Action Recognition

ICASSP 2025accepted

Recent studies have achieved remarkable results for action recognition with human skeletal data by utilizing graph convolutional models. Traditional approaches typically aggregate local spatio-temporal information bottom-up to form a single spatio-temporal global understanding. However, this method…

Cited by 0SourceScholar
2025

KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition

ICASSP 2025accepted

Point cloud sequence-based 3D action recognition has achieved impressive performance and efficiency. However, existing point cloud sequence modeling methods cannot adequately balance the precision of limb micro-movements with the integrity of posture macro-structure, leading to the loss of crucial i…

Cited by 0SourceScholar
2025

KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference

ICML 2025poster

KV cache quantization can improve Large Language Models (LLMs) inference throughput and latency in long contexts and large batch-size scenarios while preserving LLMs effectiveness. However, current methods have three unsolved issues: overlooking layer-wise sensitivity to KV cache quantization, high…

2025

Meta-Learning for Finger Vein Recognition in Internet of Things Smart Home Security

ICASSP 2025accepted

Recently, convolutional neural networks for finger vein recognition have gained attention, but their application in IoT smart home security is underexplored. Existing methods typically require networks to identify all categories in a dataset, leading to high parameter demands, which is inefficient g…

Cited by 0SourceScholar
2025

Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport

NeurIPS 2025poster

Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i.e., retrospective reconstruc…

Cited by 0SourcecodeScholar
2024

A Circuit Domain Generalization Framework for Efficient Logic Synthesis in Chip Design

ICML 2024spotlight

Logic Synthesis (LS) plays a vital role in chip design. A key task in LS is to simplify circuits---modeled by directed acyclic graphs (DAGs)---with functionality-equivalent transformations. To tackle this task, many LS heuristics apply transformations to subgraphs---rooted at each node on an input D…

2024

Adaptive Neural Network Synchronous Tracking Control for Teleoperation Robots Under Event-Triggered Mechanism

RA-L 2024

This letter proposes an adaptive neural network synchronous tracking control strategy that can be suitable for event-triggered mechanism in response to the modeling uncertainties and communication delays in bilateral teleoperation systems. Through introducing the event-triggered mechanism with the a

Cited by 8SourceScholar
2024

STREAMVC: Real-Time Low-Latency Voice Conversion

ICASSP 2024accepted

We present StreamVC, a streaming voice conversion solution that preserves the content and prosody of any source speech while matching the voice timbre from any target speech. Unlike previous approaches, StreamVC produces the resulting waveform at low latency from the input signal even on a mobile pl…

Cited by 0SourceScholar
2024

Towards Next-Generation Logic Synthesis: A Scalable Neural Circuit Generation Framework

NeurIPS 2024poster

Logic Synthesis (LS) aims to generate an optimized logic circuit satisfying a given functionality, which generally consists of circuit translation and optimization. It is a challenging and fundamental combinatorial optimization problem in integrated circuit design. Traditional LS approaches rely on…

Cited by 5SourcePDFScholar
2023

Augmentation Enables One-Shot Generalization in Learning from Demonstration for Contact-Rich Manipulation

IROS 2023poster

We introduce a Learning from Demonstration (LID) approach for contact-rich manipulation tasks, i.e., tasks in which the manipulandum's motion is constrained by contact with the environment. Our approach is motivated by the insight that even a large number of demonstrations will often not contain suf…

Cited by 4SourceScholar
2023

MST-Q: Micro Suction Tape Quadruped Robot With High Payload Capacity

RA-L 2023

Payload capacity is a crucial factor for climbing robots, as it directly affects their ability to carry and transport heavy loads during various climbing tasks. However, many dry adhesion-based legged robots prioritize foot design from a bionic perspective to accomplish various climbing tasks while

Cited by 14SourceScholar
2023

Masked Representation Learning for Domain Generalized Stereo Matching

CVPR 2023poster

Recently, many deep stereo matching methods have begun to focus on cross-domain performance, achieving impressive achievements. However, these methods did not deal with the significant volatility of generalization performance among different training epochs. Inspired by masked representation learnin…

Cited by 29SourcePDFScholar
2020

Driving Style Encoder: Situational Reward Adaptation for General-Purpose Planning in Automated Driving

ICRA 2020poster

General-purpose planning algorithms for automated driving combine mission, behavior, and local motion planning. Such planning algorithms map features of the environment and driving kinematics into complex reward functions. To achieve this, planning experts often rely on linear reward functions. The…

Cited by 12SourceScholar
2020

Planning on the fast lane: Learning to interact using attention mechanisms in path integral inverse reinforcement learning

IROS 2020poster

General-purpose trajectory planning algorithms for automated driving utilize complex reward functions to perform a combined optimization of strategic, behavioral, and kinematic features. The specification and tuning of a single reward function is a tedious task and does not generalize over a large s…

Cited by 12SourceScholar
2017

An FPGA prototype of dual link algorithm for MIMO interference network

ICASSP 2017accepted

This paper presents an FPGA-based prototype of the dual link algorithm that maximize the achievable weighted sum rate for MIMO interference network. The iterative algorithm is fast monotone convergent but it must be completed quickly through pilot signaling. Therefore we propose an FPGA-based implem…

Cited by 0SourceScholar