← Search

Jiesong Liu

4 accepted papers

2025

A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language Models

ACL 2025long

Large Language Models (LLMs) are cutting-edge generative AI models built on transformer architecture, which tend to be highly memory-intensive when performing real-time inference. Various strategies have been developed to enhance the end-to-end inference speed for LLMs, one of which is speculative d…

Cited by 0SourcePDFScholar
2025

Fourier Token Merging: Understanding and Capitalizing Frequency Domain for Efficient Image Generation

NeurIPS 2025poster

Image generation requires intensive computations and faces challenges due to long latency. Exploiting redundancy in the input images and intermediate representations throughout the neural network pipeline is an effective way to accelerate image generation. Token merging (ToMe) exploits simil…

Cited by 0SourcecodeScholar
2024

UQ-Guided Hyperparameter Optimization for Iterative Learners

NeurIPS 2024poster

Hyperparameter Optimization (HPO) plays a pivotal role in unleashing the potential of iterative machine learning models. This paper addresses a crucial aspect that has largely been overlooked in HPO: the impact of uncertainty in ML model training. The paper introduces the concept of uncertainty-awar…

Cited by 2SourcePDFScholar
2022

TREC: Transient Redundancy Elimination-based Convolution

NeurIPS 2022accept

The intensive computations in convolutional neural networks (CNNs) pose challenges for resource-constrained devices; eliminating redundant computations from convolution is essential. This paper gives a principled method to detect and avoid transient redundancy, a type of redundancy existing in input…

Cited by 5SourcePDFScholar