← Search

Seungcheol Park

4 accepted papers

2025

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information

IJCAI 2025

How can we accelerate large language models (LLMs) without sacrificing accuracy? The slow inference speed of LLMs hinders us to benefit from their remarkable performance in diverse applications. This is mainly because numerous sublayers are stacked together in LLMs. Sublayer pruning compresses and e

2025

Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models

ACL 2025long

How can we quantize large language models while preserving accuracy? Quantization is essential for deploying large language models (LLMs) efficiently. Binary-coding quantization (BCQ) and uniform quantization (UQ) are promising quantization schemes that have strong expressiveness and optimizability,…

2024

Accurate Retraining-free Pruning for Pretrained Encoder-based Language Models

ICLR 2024poster

Given a pretrained encoder-based language model, how can we accurately compress it without retraining? Retraining-free structured pruning algorithms are crucial in pretrained language model compression due to their significantly reduced pruning cost and capability to prune large language models. How…

2019

Curved-Voxel Clustering for Accurate Segmentation of 3D LiDAR Point Clouds with Real-Time Performance

IROS 2019poster

Given 3D LiDAR point clouds, how can we segment them fast and accurately? Fast and accurate segmentation of 3D LiDAR points is an important issue in mobile robotics with various applications in classification, tracking, SLAM, etc. Despite its importance, existing methods do not provide both speed an…

Cited by 39SourceScholar