← Search

Yao Luo

5 accepted papers

2025

FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

ICLR 2025oral

Large language models (LLMs) encounter computational challenges during long-sequence inference, especially in the attention pre-filling phase, where the complexity grows quadratically with the prompt length. Previous efforts to mitigate these challenges have relied on fixed sparse attention patterns…

2025

Learning-Based Slip Detection and Fine Control Using the Tactile Sensor for Robot Stable Grasping

RA-L 2025

Slip detection and control is critical to achieving stable grasping in robotics. However, accurate and robust slip detection and control remains a challenging task. This letter proposes a learning framework with contrastive learning and feature alignment to improve the accuracy of end-to-end slip de

Cited by 3SourceScholar
2025

Model Merging in Pre-training of Large Language Models

NeurIPS 2025poster

Model merging has emerged as a promising technique for enhancing large language models, though its application in large-scale pre-training remains relatively unexplored. In this paper, we present a comprehensive investigation of model merging techniques during the pre-training process. Through exten…

Cited by 0SourceScholar
2025

Why Does the Effective Context Length of LLMs Fall Short?

ICLR 2025poster

Advancements in distributed training and efficient attention mechanisms have significantly expanded the context window sizes of large language models (LLMs). However, recent work reveals that the effective context lengths of open-source LLMs often fall short, typically not exceeding half of their tr…

Cited by 60SourcePDFScholar
2023

SVMV: Spatiotemporal Variance-Supervised Motion Volume for Video Frame Interpolation

ICASSP 2023accepted

High-performance video frame interpolation is challenging for complex scenes with diverse motion and occlusion characteristics. Existing methods, deploying off-the-shelf flow estimators to acquire initial characterizations refined by multiple subsequent models, often require heavy network architectu…

Cited by 0SourceScholar