← Search

Seongjin Shin

4 accepted papers

2025

Peri-LN: Revisiting Normalization Layer in the Transformer Architecture

ICML 2025poster

Selecting a layer normalization (LN) strategy that stabilizes training and speeds convergence in Transformers remains difficult, even for today’s large language models (LLM). We present a comprehensive analytical foundation for understanding how different LN strategies influence training dynamics in…

Cited by 0SourcePDFScholar
2024

Prometheus: Inducing Fine-Grained Evaluation Capability in Language Models

ICLR 2024poster

Recently, GPT-4 has become the de facto evaluator for long-form text generated by large language models (LLMs). However, for practitioners and researchers with large and custom evaluation tasks, GPT-4 is unreliable due to its closed-source nature, uncontrolled versioning, and prohibitive costs. In t…

2022

On the Effect of Pretraining Corpora on In-context Learning by a Large-scale Language Model

NAACL 2022long

Many recent studies on large-scale language models have reported successful in-context zero- and few-shot learning ability. However, the in-depth analysis of when in-context learning occurs is still lacking. For example, it is unknown how in-context learning performance changes as the training corpu…

2021

Two-Stage Textual Knowledge Distillation for End-to-End Spoken Language Understanding

ICASSP 2021accepted

End-to-end approaches open a new way for more accurate and efficient spoken language understanding (SLU) systems by alleviating the drawbacks of traditional pipeline systems. Previous works exploit textual information for an SLU model via pre-training with automatic speech recognition or finetuning…

Cited by 0SourceScholar