← Search

Sangtae Ha

4 accepted papers

2024

LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs

IJCAI 2024poster

Fine-tuning pre-trained large language models (LLMs) with limited hardware presents challenges due to GPU memory constraints. Various distributed fine-tuning methods have been proposed to alleviate memory constraints on GPU. However, determining the most effective method for achieving rapid fine-tun…

2023

ASPEN: Breaking Operator Barriers for Efficient Parallelization of Deep Neural Networks

NeurIPS 2023poster

Modern Deep Neural Network (DNN) frameworks use tensor operators as the main building blocks of DNNs. However, we observe that operator-based construction of DNNs incurs significant drawbacks in parallelism in the form of synchronization barriers. Synchronization barriers of operators confine the sc…

2022

CPrune: Compiler-Informed Model Pruning for Efficient Target-Aware DNN Execution

ECCV 2022poster

"Mobile devices run deep learning models for various purposes, such as image classification and speech recognition. Due to the resource constraints of mobile devices, researchers have focused on either making a lightweight deep neural network (DNN) model using model pruning or generating an efficien…

2022

TVConv: Efficient Translation Variant Convolution for Layout-Aware Visual Processing

CVPR 2022poster

As convolution has empowered many smart applications, dynamic convolution further equips it with the ability to adapt to diverse inputs. However, the static and dynamic convolutions are either layout-agnostic or computation-heavy, making it inappropriate for layout-specific applications, e.g., face…

Cited by 38PDFcodeScholar