← Search

Yuchen Li

29 accepted papers

2026

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

AAAI 2026technical

The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Models (LLMs). However, this architectural enhancement comes at a steep cost: despite minimal increases in computational loa

Cited by 0SourcePDFScholar
2026

Efficient Thought Space Exploration Through Strategic Intervention

AAAI 2026technical

While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs through exhaustive sampling. Through analyzing decoding trajectories, we observe that most next-token predictions align well with the golden

Cited by 0SourcePDFScholar
2026

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

ICML 2026poster

Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limitations: (i) self-attention lacks an explicit inductive bias for locality, leading to redundant modeling of sequence-int…

Cited by 0SourceScholar
2026

SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs

ICLR 2026poster

Humans can imagine and manipulate visual images mentally, a capability known as \textit{spatial visualization}. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through spatial visualization remains insufficiently evaluated…

Cited by 0SourcecodeScholar
2026

SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks

ICML 2026poster

Vision-Language-Action (VLA) models have become a central paradigm for embodied intelligence. However, most existing approaches are built on large-scale Transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scena…

Cited by 0SourceScholar
2026

Taming Imperfect Process Verifiers: A Sampling Perspective on Backtracking

ICLR 2026poster

Test-time algorithms that combine the *generative* power of language models with *process verifiers* that assess the quality of partial generations offer a promising lever for eliciting new reasoning capabilities, but the algorithmic design space and computational scaling properties of such approach…

Cited by 0SourceScholar
2025

Enhancing Retrieval-Augmented Generation via Evidence Tree Search

ACL 2025long

Retrieval-Augmented Generation (RAG) is widely used to enhance Large Language Models (LLMs) by grounding responses in external knowledge. However, in real-world applications, retrievers often return lengthy documents with redundant or irrelevant content, confusing downstream readers. While evidence…

Cited by 0SourcePDFScholar
2025

Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion Model

ICASSP 2025accepted

Electroencephalography (EEG) is a time-series signal containing semantic information that can be used to determine human brain activities. Artifacts within EEG data can interfere with the intrinsic distribution of this semantic information, so removing artifacts is crucial for improving EEG analysis…

Cited by 0SourceScholar
2025

GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation

NeurIPS 2025poster

We introduce GS2E (Gaussian Splatting to Event Generation), a large-scale synthetic event dataset designed for high-fidelity event vision tasks, captured from real-world sparse multi-view RGB images. Existing event datasets are often synthesized from dense RGB videos, which typically suffer from lim…

Cited by 0SourceScholar
2025

MLC-NC: Long-Tailed Multi-Label Image Classification Through the Lens of Neural Collapse

AAAI 2025technical

Long-tailed (LT) data distribution is common in multi-label image classification (MLC) and can significantly impact the performance of classification models. One reason is the challenge of learning unbiased instance representations (i.e. features) for imbalanced datasets. Additionally, the co-occurr…

Cited by 0SourcePDFScholar
2025

On the Query Complexity of Verifier-Assisted Language Generation

ICML 2025poster

Recently, a plethora of works have proposed inference-time algorithms (e.g. best-of-n), which incorporate verifiers to assist the generation process. Their quality-efficiency trade-offs have been empirically benchmarked on a variety of constrained generation tasks, but the algorithmic design landsca…

Cited by 1SourcePDFScholar
2025

SOLA-GCL: Subgraph-Oriented Learnable Augmentation Method for Graph Contrastive Learning

AAAI 2025technical

Graph contrastive learning has emerged as a powerful technique for learning graph representations that are robust and discriminative. However, traditional approaches often neglect the critical role of subgraph structures, particularly the intra-subgraph characteristics and inter-subgraph relationshi…

Cited by 0SourcePDFScholar
2024

Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms

ICML 2024poster

We analyze inexact Riemannian gradient descent (RGD) where Riemannian gradients and retractions are inexactly (and cheaply) computed. Our focus is on understanding when inexact RGD converges and what is the complexity in the general nonconvex and constrained setting. We answer these questions in a g…

Cited by 0SourcePDFScholar
2024

GS2P: A Generative Pre-trained Learning to Rank Model with Over-parameterization for Web-Scale Search (Extended Abstract)

IJCAI 2024poster

While Learning to Rank (LTR) is widely employed in web searches to prioritize pertinent webpages from the retrieved contents based on input queries, traditional LTR models stumble over two principal stumbling blocks leading to subpar performance: 1) the lack of well-annotated query-webpage pairs wit…

Cited by 7SourcePDFScholar
2024

HyperPrism: An Adaptive Non-linear Aggregation Framework for Distributed Machine Learning over Non-IID Data and Time-varying Communication Links

NeurIPS 2024poster

While Distributed Machine Learning (DML) has been widely used to achieve decent performance, it is still challenging to take full advantage of data and devices distributed at multiple vantage points to adapt and learn, especially it is non-trivial to address dynamic and divergence challenges based o…

Cited by 0SourcePDFScholar
2024

MPGraf: a Modular and Pre-trained Graphformer for Learning to Rank at Web-scale (Extended Abstract)

IJCAI 2024poster

Both Transformer and Graph Neural Networks (GNNs) have been used in learning to rank (LTR), however, they adhere to two distinct yet complementary problem formulations, i.e., ranking score regression based on query-webpage pairs and link prediction within query-webpage bipartite graphs, respectively…

Cited by 0SourcePDFScholar
2024

Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines

ICML 2024poster

Autoregressive language models are the currently dominant paradigm for text generation, however they have some fundamental limitations that cannot be remedied by scale---for example inherently sequential and unidirectional generation. While alternate classes of models have been explored, we have lim…

2023

An Adaptive DFE Using Light-Pattern-Protection Algorithm in 12 NM CMOS Technology

ICASSP 2023accepted

The sign-sign least-mean-squares (SSLMS) algorithm has been widely used in decision feedback equalizer (DFE) adaptation. However, the convergence direction of DFE tap coefficients in the training process is closely related to the data flow. In the case of extreme data flow, the coefficients may conv…

Cited by 0SourceScholar
2023

How Do Transformers Learn Topic Structure: Towards a Mechanistic Understanding

ICML 2023poster

While the successes of transformers across many domains are indisputable, accurate understanding of the learning mechanics is still largely lacking. Their capabilities have been probed on benchmarks which include a variety of structured and reasoning tasks---but mathematical understanding is lagging…

2023

Transformers are uninterpretable with myopic methods: a case study with bounded Dyck grammars

NeurIPS 2023poster

Transformer interpretability aims to understand the algorithm implemented by a learned Transformer by examining various aspects of the model, such as the weight matrices or the attention patterns. In this work, through a combination of theoretical results and carefully controlled experiments on synt…

Cited by 23SourcePDFScholar
2022

3D CoMPaT: Composition of Materials on Parts of 3D Things

ECCV 2022poster

"We present 3D CoMPaT, a richly annotated large-scale dataset of more than 7.19 million rendered compositions of Materials on Parts of 7262 unique 3D Models; 990 compositions per model on average. 3D CoMPaT covers 43 shape categories, 235 unique part names, and 167 unique material classes that can b…

Cited by 16SourcePDFScholar
2022

CamMap: Extrinsic Calibration of Non-Overlapping Cameras Based on SLAM Map Alignment

RA-L 2022

Multiple cameras have emerged as a promising technology for robots and vehicles due to their broad fields of view (FoV) and high resolution. However, there are often limited or no overlapping FoVs among cameras, bringing challenges to estimating extrinsic camera parameters. To overcome this problem,

Cited by 12SourceScholar
2022

Contrasting the landscape of contrastive and non-contrastive learning

AISTATS 2022poster

A lot of recent advances in unsupervised feature learning are based on designing features which are invariant under semantic data augmentations. A common way to do this is contrastive learning, which uses positive and negative samples. Some recent works however have shown promising results for non-c…

2022

PointNeXt: Revisiting PointNet++ with Improved Training and Scaling Strategies

NeurIPS 2022accept

PointNet++ is one of the most influential neural architectures for point cloud understanding. Although the accuracy of PointNet++ has been largely surpassed by recent networks such as PointMLP and Point Transformer, we find that a large portion of the performance gain is due to improved training str…

2020

mROBerTO 2.0 - An Autonomous Millirobot With Enhanced Locomotion for Swarm Robotics

RA-L 2020

Numerous millirobots were developed in the past decade for autonomous swarm systems that aim to utilize large numbers of these units in space-constrained environments. However, the size limitation of these robots has often resulted in their reduced computational, sensing, and locomotion capabilities

Cited by 13SourceScholar